Daily Papers of 2025-08-08
- On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification 144 upvotes, #1 of 2025-08-08
- R-Zero: Self-Evolving Reasoning LLM from Zero Data 107 upvotes, #2 of 2025-08-08
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation 67 upvotes, #3 of 2025-08-08
- DeepPHY: Benchmarking Agentic VLMs on Physical Reasoning 61 upvotes, #4 of 2025-08-08
- Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity 29 upvotes, #5 of 2025-08-08
- Are We on the Right Way for Assessing Document Retrieval-Augmented Generation? 24 upvotes, #6 of 2025-08-08
- Are Today's LLMs Ready to Explain Well-Being Concepts? 23 upvotes, #7 of 2025-08-08
- Don't Overthink It: A Survey of Efficient R1-style Large Reasoning Models 16 upvotes, #8 of 2025-08-08
- Marco-Voice Technical Report 15 upvotes, #9 of 2025-08-08
- CoAct-1: Computer-using Agents with Coding as Actions 13 upvotes, #10 of 2025-08-08
- Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability 10 upvotes, #11 of 2025-08-08
- Evaluating, Synthesizing, and Enhancing for Customer Support Conversation 9 upvotes, #12 of 2025-08-08
- MOSEv2: A More Challenging Dataset for Video Object Segmentation in Complex Scenes 9 upvotes, #12 of 2025-08-08
- InfiAlign: A Scalable and Sample-Efficient Framework for Aligning LLMs to Enhance Reasoning Capabilities 8 upvotes, #14 of 2025-08-08
- StrandDesigner: Towards Practical Strand Generation with Sketch Guidance 6 upvotes, #15 of 2025-08-08
- Steering One-Step Diffusion Model with Fidelity-Rich Decoder for Fast Image Compression 5 upvotes, #16 of 2025-08-08
- Attention Basin: Why Contextual Position Matters in Large Language Models 4 upvotes, #17 of 2025-08-08
- Visual Document Understanding and Question Answering: A Multi-Agent Collaboration Framework with Test-Time Scaling 3 upvotes, #18 of 2025-08-08
- Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decode 3 upvotes, #18 of 2025-08-08
- Learning to Reason for Factuality 3 upvotes, #18 of 2025-08-08
- I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking 2 upvotes, #21 of 2025-08-08
- I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations 2 upvotes, #21 of 2025-08-08
- RPCANet++: Deep Interpretable Robust PCA for Sparse Object Segmentation 1 upvotes, #23 of 2025-08-08
- Hop, Skip, and Overthink: Diagnosing Why Reasoning Models Fumble during Multi-Hop Analysis 1 upvotes, #23 of 2025-08-08
- REINA: Regularized Entropy Information-Based Loss for Efficient Simultaneous Speech Translation 1 upvotes, #23 of 2025-08-08
- PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction 1 upvotes, #23 of 2025-08-08
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.