Yingfa Chen

Yingfa Chen on Hugging Face Daily Papers: 11 papers, 1 in the top 3 of their day, 118 upvotes.

  1. Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It 16 upvotes, #17 of 2026-06-10
  2. DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices 3 upvotes, #43 of 2026-05-12
  3. MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modeling 6 upvotes, #27 of 2026-02-13
  4. Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts 10 upvotes, #22 of 2026-01-30
  5. StateX: Enhancing RNN Recall via Post-training State Expansion 2 upvotes, #37 of 2025-09-29
  6. BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity 8 upvotes, #13 of 2025-07-14
  7. Cost-Optimal Grouped-Query Attention for Long-Context LLMs 5 upvotes, #16 of 2025-03-13
  8. Sparsing Law: Towards Large Language Models with Greater Activation Sparsity 10 upvotes, #15 of 2024-11-05
  9. Stuffed Mamba: State Collapse and State Capacity of RNN-Based Long-Context Modeling 2 upvotes, #47 of 2024-10-10
  10. Configurable Foundation Models: Building LLMs from a Modular Perspective 26 upvotes, #2 of 2024-09-09
  11. Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models 14 upvotes, #12 of 2024-06-25

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.