Daily Papers of 2025-08-08

  1. On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification 144 upvotes, #1 of 2025-08-08
  2. R-Zero: Self-Evolving Reasoning LLM from Zero Data 107 upvotes, #2 of 2025-08-08
  3. Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation 67 upvotes, #3 of 2025-08-08
  4. DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning 61 upvotes, #4 of 2025-08-08
  5. Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity 29 upvotes, #5 of 2025-08-08
  6. Are We on the Right Way for Assessing Document Retrieval-Augmented Generation? 24 upvotes, #6 of 2025-08-08
  7. Are Today's LLMs Ready to Explain Well-Being Concepts? 23 upvotes, #7 of 2025-08-08
  8. Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models 16 upvotes, #8 of 2025-08-08
  9. Marco-Voice Technical Report 15 upvotes, #9 of 2025-08-08
  10. CoAct-1: Computer-using Agents with Coding as Actions 13 upvotes, #10 of 2025-08-08
  11. Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability 10 upvotes, #11 of 2025-08-08
  12. Evaluating, Synthesizing, and Enhancing for Customer Support Conversation 9 upvotes, #12 of 2025-08-08
  13. MOSEv2: A More Challenging Dataset for Video Object Segmentation in Complex Scenes 9 upvotes, #12 of 2025-08-08
  14. InfiAlign: A Scalable and Sample-Efficient Framework for Aligning LLMs to Enhance Reasoning Capabilities 8 upvotes, #14 of 2025-08-08
  15. StrandDesigner: Towards Practical Strand Generation with Sketch Guidance 6 upvotes, #15 of 2025-08-08
  16. Steering One-Step Diffusion Model with Fidelity-Rich Decoder for Fast Image Compression 5 upvotes, #16 of 2025-08-08
  17. Attention Basin: Why Contextual Position Matters in Large Language Models 4 upvotes, #17 of 2025-08-08
  18. Visual Document Understanding and Question Answering: A Multi-Agent Collaboration Framework with Test-Time Scaling 3 upvotes, #18 of 2025-08-08
  19. Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decode 3 upvotes, #18 of 2025-08-08
  20. Learning to Reason for Factuality 3 upvotes, #18 of 2025-08-08
  21. I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking 2 upvotes, #21 of 2025-08-08
  22. I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations 2 upvotes, #21 of 2025-08-08
  23. RPCANet++: Deep Interpretable Robust PCA for Sparse Object Segmentation 1 upvotes, #23 of 2025-08-08
  24. Hop, Skip, and Overthink: Diagnosing Why Reasoning Models Fumble during Multi-Hop Analysis 1 upvotes, #23 of 2025-08-08
  25. REINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech Translation 1 upvotes, #23 of 2025-08-08
  26. PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction 1 upvotes, #23 of 2025-08-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.