Daily Papers of 2025-02-11

  1. Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling 128 upvotes, #1 of 2025-02-11
  2. SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators 84 upvotes, #2 of 2025-02-11
  3. Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning 55 upvotes, #3 of 2025-02-11
  4. Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning 31 upvotes, #4 of 2025-02-11
  5. The Curse of Depth in Large Language Models 27 upvotes, #5 of 2025-02-11
  6. LM2: Large Memory Models 26 upvotes, #6 of 2025-02-11
  7. Matryoshka Quantization 24 upvotes, #7 of 2025-02-11
  8. CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging 21 upvotes, #8 of 2025-02-11
  9. Show-o Turbo: Towards Accelerated Unified Multimodal Understanding and Generation 20 upvotes, #9 of 2025-02-11
  10. ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates 18 upvotes, #10 of 2025-02-11
  11. MetaChain: A Fully-Automated and Zero-Code Framework for LLM Agents 15 upvotes, #11 of 2025-02-11
  12. Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding 14 upvotes, #12 of 2025-02-11
  13. Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT 12 upvotes, #13 of 2025-02-11
  14. The Hidden Life of Tokens: Reducing Hallucination of Large Vision-Language Models via Visual Information Steering 11 upvotes, #14 of 2025-02-11
  15. EVEv2: Improved Baselines for Encoder-Free Vision-Language Models 11 upvotes, #14 of 2025-02-11
  16. History-Guided Video Diffusion 10 upvotes, #16 of 2025-02-11
  17. Dual Caption Preference Optimization for Diffusion Models 9 upvotes, #17 of 2025-02-11
  18. CustomVideoX: 3D Reference Attention Driven Dynamic Adaptation for Zero-Shot Customized Video Diffusion Transformers 9 upvotes, #17 of 2025-02-11
  19. Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile 7 upvotes, #19 of 2025-02-11
  20. DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization 6 upvotes, #20 of 2025-02-11
  21. APE: Faster and Longer Context-Augmented Generation via Adaptive Parallel Encoding 5 upvotes, #21 of 2025-02-11
  22. Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE 4 upvotes, #22 of 2025-02-11
  23. Steel-LLM:From Scratch to Open Source -- A Personal Journey in Building a Chinese-Centric LLM 4 upvotes, #22 of 2025-02-11
  24. Towards Internet-Scale Training For Agents 4 upvotes, #22 of 2025-02-11
  25. Embodied Red Teaming for Auditing Robotic Foundation Models 1 upvotes, #25 of 2025-02-11
  26. Forbidden Science: Dual-Use AI Challenge Benchmark and Scientific Refusal Tests 1 upvotes, #25 of 2025-02-11

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.