Yu Cheng

Yu Cheng on Hugging Face Daily Papers: 27 papers, 8 in the top 3 of their day, 1,025 upvotes.

  1. Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling 154 upvotes, #1 of 2026-05-15
  2. The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models 114 upvotes, #1 of 2025-05-29
  3. SATORI-R1: Incentivizing Multimodal Reasoning with Spatial Grounding and Verifiable Rewards 2 upvotes, #62 of 2025-05-28
  4. FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow 14 upvotes, #19 of 2025-05-26
  5. Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models 58 upvotes, #2 of 2025-05-23
  6. OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning 39 upvotes, #4 of 2025-05-16
  7. Learning to Reason under Off-Policy Guidance 77 upvotes, #1 of 2025-04-22
  8. TransMamba: Flexibly Switching between Transformer and Mamba 16 upvotes, #6 of 2025-04-07
  9. A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond 39 upvotes, #3 of 2025-03-31
  10. From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration 9 upvotes, #14 of 2025-03-24
  11. Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts 7 upvotes, #19 of 2025-03-10
  12. Liger: Linearizing Large Language Models to Gated Recurrent Structures 15 upvotes, #9 of 2025-03-04
  13. Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment 23 upvotes, #8 of 2025-02-25
  14. MoM: Linear Sequence Modeling with Mixture-of-Memories 31 upvotes, #5 of 2025-02-20
  15. LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid 23 upvotes, #10 of 2025-02-13
  16. UltraIF: Advancing Instruction Following from the Wild 20 upvotes, #9 of 2025-02-07
  17. Process Reinforcement through Implicit Rewards 53 upvotes, #3 of 2025-02-04
  18. Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback 50 upvotes, #5 of 2025-01-23
  19. Scaling Laws for Floating Point Quantization Training 24 upvotes, #6 of 2025-01-07
  20. PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models 13 upvotes, #13 of 2025-01-07
  21. Diving into Self-Evolving Training for Multimodal Reasoning 37 upvotes, #3 of 2024-12-24
  22. Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation 42 upvotes, #5 of 2024-10-10
  23. Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models 5 upvotes, #11 of 2024-10-09
  24. CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling 17 upvotes, #12 of 2024-10-04
  25. ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM 10 upvotes, #13 of 2024-08-23
  26. Direct Preference Knowledge Distillation for Large Language Models 21 upvotes, #3 of 2024-07-01
  27. Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning 10 upvotes, #11 of 2024-02-20

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.