DeepSeek

DeepSeek on Hugging Face Daily Papers: 23 papers, 16 in the top 3 of their day, 10 paper of the day.

  1. DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression 180 upvotes, #1 of 2026-09-18
  2. A Programming Paradigm for Spatiotemporal Composability 13 upvotes, #17 of 2026-08-27
  3. DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation 33 upvotes, #7 of 2026-07-08
  4. DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference 37 upvotes, #5 of 2026-02-26
  5. DeepSeek-OCR 2: Visual Causal Flow 53 upvotes, #4 of 2026-01-29
  6. mHC: Manifold-Constrained Hyper-Connections 227 upvotes, #1 of 2026-01-01
  7. DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models 194 upvotes, #1 of 2025-12-03
  8. DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning 66 upvotes, #2 of 2025-12-01
  9. DeepSeek-OCR: Contexts Optical Compression 63 upvotes, #5 of 2025-10-22
  10. Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures 55 upvotes, #2 of 2025-05-15
  11. Inference-Time Scaling for Generalist Reward Modeling 50 upvotes, #5 of 2025-04-04
  12. Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention 134 upvotes, #1 of 2025-02-18
  13. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning 271 upvotes, #1 of 2025-01-23
  14. Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation 27 upvotes, #8 of 2024-10-18
  15. DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search 48 upvotes, #1 of 2024-08-16
  16. Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models 33 upvotes, #2 of 2024-07-05
  17. DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence 53 upvotes, #1 of 2024-06-19
  18. DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data 27 upvotes, #2 of 2024-05-24
  19. DeepSeek-VL: Towards Real-World Vision-Language Understanding 33 upvotes, #3 of 2024-03-11
  20. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models 150 upvotes, #1 of 2024-02-06
  21. DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence 74 upvotes, #1 of 2024-01-26
  22. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models 65 upvotes, #2 of 2024-01-12
  23. DeepSeek LLM: Scaling Open-Source Language Models with Longtermism 56 upvotes, #1 of 2024-01-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.