Daily Papers of 2025-08-12
- ReasonRank: Empowering Passage Ranking with Strong Reasoning Ability 112 upvotes, #1 of 2025-08-12
- WideSearch: Benchmarking Agentic Broad Info-Seeking 102 upvotes, #2 of 2025-08-12
- A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems 80 upvotes, #3 of 2025-08-12
- Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation 57 upvotes, #4 of 2025-08-12
- SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens 45 upvotes, #5 of 2025-08-12
- Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning 38 upvotes, #6 of 2025-08-12
- Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization 37 upvotes, #7 of 2025-08-12
- MolmoAct: Action Reasoning Models that can Reason in Space 37 upvotes, #7 of 2025-08-12
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent 36 upvotes, #9 of 2025-08-12
- UserBench: An Interactive Gym Environment for User-Centric Agents 29 upvotes, #10 of 2025-08-12
- Reinforcement Learning in Vision: A Survey 27 upvotes, #11 of 2025-08-12
- Grove MoE: Towards Efficient and Superior MoE LLMs with Adjugate Experts 23 upvotes, #12 of 2025-08-12
- Spectrum Projection Score: Aligning Retrieved Summaries with Reader Models in Retrieval-Augmented Generation 20 upvotes, #13 of 2025-08-12
- OmniEAR: Benchmarking Agent Reasoning in Embodied Tasks 18 upvotes, #14 of 2025-08-12
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future 15 upvotes, #15 of 2025-08-12
- Less Is More: Training-Free Sparse Attention with Global Locality for Efficient Reasoning 13 upvotes, #16 of 2025-08-12
- MoBE: Mixture-of-Basis-Experts for Compressing MoE-based LLMs 10 upvotes, #17 of 2025-08-12
- Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation 10 upvotes, #17 of 2025-08-12
- Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control 9 upvotes, #19 of 2025-08-12
- GLiClass: Generalist Lightweight Model for Sequence Classification Tasks 8 upvotes, #20 of 2025-08-12
- Compressing Chain-of-Thought in LLMs via Step Entropy 7 upvotes, #21 of 2025-08-12
- VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding 7 upvotes, #21 of 2025-08-12
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents 6 upvotes, #23 of 2025-08-12
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs 5 upvotes, #24 of 2025-08-12
- Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences 4 upvotes, #25 of 2025-08-12
- Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System 4 upvotes, #25 of 2025-08-12
- Anatomy of a Machine Learning Ecosystem: 2 Million Models on Hugging Face 3 upvotes, #27 of 2025-08-12
- When Good Sounds Go Adversarial: Jailbreaking Audio-Language Models with Benign Inputs 2 upvotes, #28 of 2025-08-12
- TextQuests: How Good are LLMs at Text-Based Video Games? 1 upvotes, #29 of 2025-08-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.