Daily Papers of 2025-04-02

  1. Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation 71 upvotes, #1 of 2025-04-02
  2. JudgeLRM: Large Reasoning Models as a Judge 55 upvotes, #2 of 2025-04-02
  3. Multi-Token Attention 39 upvotes, #3 of 2025-04-02
  4. Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1 36 upvotes, #4 of 2025-04-02
  5. Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources 33 upvotes, #5 of 2025-04-02
  6. CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis 31 upvotes, #6 of 2025-04-02
  7. GeometryCrafter: Consistent Geometry Estimation for Open-world Videos with Diffusion Priors 28 upvotes, #7 of 2025-04-02
  8. Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models 26 upvotes, #8 of 2025-04-02
  9. Z1: Efficient Test-time Scaling with Code 25 upvotes, #9 of 2025-04-02
  10. Scaling Language-Free Visual Representation Learning 24 upvotes, #10 of 2025-04-02
  11. Command A: An Enterprise-Ready Large Language Model 23 upvotes, #11 of 2025-04-02
  12. Recitation over Reasoning: How Cutting-Edge Language Models Can Fail on Elementary School-Level Reasoning Problems? 21 upvotes, #12 of 2025-04-02
  13. Towards Trustworthy GUI Agents: A Survey 20 upvotes, #13 of 2025-04-02
  14. Agent S2: A Compositional Generalist-Specialist Framework for Computer Use Agents 19 upvotes, #14 of 2025-04-02
  15. OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts 18 upvotes, #15 of 2025-04-02
  16. Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models 17 upvotes, #16 of 2025-04-02
  17. MixerMDM: Learnable Composition of Human Motion Diffusion Models 17 upvotes, #16 of 2025-04-02
  18. Efficient LLaMA-3.2-Vision by Trimming Cross-attended Visual Features 15 upvotes, #18 of 2025-04-02
  19. When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning 15 upvotes, #18 of 2025-04-02
  20. YourBench: Easy Custom Evaluation Sets for Everyone 15 upvotes, #18 of 2025-04-02
  21. AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization 11 upvotes, #21 of 2025-04-02
  22. Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead 10 upvotes, #22 of 2025-04-02
  23. m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning with Large Language Models 10 upvotes, #22 of 2025-04-02
  24. Reasoning-SQL: Reinforcement Learning with SQL Tailored Partial Rewards for Reasoning-Enhanced Text-to-SQL 7 upvotes, #24 of 2025-04-02
  25. Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs 7 upvotes, #24 of 2025-04-02
  26. Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base 6 upvotes, #26 of 2025-04-02
  27. ManipTrans: Efficient Dexterous Bimanual Manipulation Transfer via Residual Learning 3 upvotes, #27 of 2025-04-02
  28. DiET-GS: Diffusion Prior and Event Stream-Assisted Motion Deblurring 3D Gaussian Splatting 3 upvotes, #27 of 2025-04-02
  29. MB-ORES: A Multi-Branch Object Reasoner for Visual Grounding in Remote Sensing 2 upvotes, #29 of 2025-04-02

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.