Yu Cheng
Yu Cheng on Hugging Face Daily Papers: 27 papers, 8 in the top 3 of their day, 1,025 upvotes.
- Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling 154 upvotes, #1 of 2026-05-15
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models 114 upvotes, #1 of 2025-05-29
- SATORI-R1: Incentivizing Multimodal Reasoning with Spatial Grounding and Verifiable Rewards 2 upvotes, #62 of 2025-05-28
- FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow 14 upvotes, #19 of 2025-05-26
- Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models 58 upvotes, #2 of 2025-05-23
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning 39 upvotes, #4 of 2025-05-16
- Learning to Reason under Off-Policy Guidance 77 upvotes, #1 of 2025-04-22
- TransMamba: Flexibly Switching between Transformer and Mamba 16 upvotes, #6 of 2025-04-07
- A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, and Beyond 39 upvotes, #3 of 2025-03-31
- From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration 9 upvotes, #14 of 2025-03-24
- Linear-MoE: Linear Sequence Modeling Meets Mixture-of-Experts 7 upvotes, #19 of 2025-03-10
- Liger: Linearizing Large Language Models to Gated Recurrent Structures 15 upvotes, #9 of 2025-03-04
- Make LoRA Great Again: Boosting LoRA with Adaptive Singular Values and Mixture-of-Experts Optimization Alignment 23 upvotes, #8 of 2025-02-25
- MoM: Linear Sequence Modeling with Mixture-of-Memories 31 upvotes, #5 of 2025-02-20
- LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid 23 upvotes, #10 of 2025-02-13
- UltraIF: Advancing Instruction Following from the Wild 20 upvotes, #9 of 2025-02-07
- Process Reinforcement through Implicit Rewards 53 upvotes, #3 of 2025-02-04
- Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback 50 upvotes, #5 of 2025-01-23
- Scaling Laws for Floating Point Quantization Training 24 upvotes, #6 of 2025-01-07
- PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models 13 upvotes, #13 of 2025-01-07
- Diving into Self-Evolving Training for Multimodal Reasoning 37 upvotes, #3 of 2024-12-24
- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation 42 upvotes, #5 of 2024-10-10
- Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models 5 upvotes, #11 of 2024-10-09
- CLIP-MoE: Towards Building Mixture of Experts for CLIP with Diversified Multiplet Upcycling 17 upvotes, #12 of 2024-10-04
- ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM 10 upvotes, #13 of 2024-08-23
- Direct Preference Knowledge Distillation for Large Language Models 21 upvotes, #3 of 2024-07-01
- Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning 10 upvotes, #11 of 2024-02-20
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.