Daily Papers of 2025-08-07

  1. Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens 204 upvotes, #1 of 2025-08-07
  2. VeriGUI: Verifiable Long-Chain GUI Dataset 137 upvotes, #2 of 2025-08-07
  3. Efficient Agents: Building Effective Agents While Reducing Cost 79 upvotes, #3 of 2025-08-07
  4. Agent Lightning: Train ANY AI Agents with Reinforcement Learning 54 upvotes, #4 of 2025-08-07
  5. Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning 48 upvotes, #5 of 2025-08-07
  6. SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience 46 upvotes, #6 of 2025-08-07
  7. Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success 34 upvotes, #7 of 2025-08-07
  8. Sotopia-RL: Reward Design for Social Intelligence 23 upvotes, #8 of 2025-08-07
  9. CoTox: Chain-of-Thought-Based Molecular Toxicity Reasoning and Prediction 22 upvotes, #9 of 2025-08-07
  10. Web-CogReasoner: Towards Knowledge-Induced Cognitive Reasoning for Web Agents 20 upvotes, #10 of 2025-08-07
  11. LaTCoder: Converting Webpage Design to Code with Layout-as-Thought 19 upvotes, #11 of 2025-08-07
  12. HPSv3: Towards Wide-Spectrum Human Preference Score 18 upvotes, #12 of 2025-08-07
  13. Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis 17 upvotes, #13 of 2025-08-07
  14. DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework 13 upvotes, #14 of 2025-08-07
  15. Position: The Current AI Conference Model is Unsustainable! Diagnosing the Crisis of Centralized AI Conference 12 upvotes, #15 of 2025-08-07
  16. LeanK: Learnable K Cache Channel Pruning for Efficient Decoding 11 upvotes, #16 of 2025-08-07
  17. Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management 10 upvotes, #17 of 2025-08-07
  18. IAUNet: Instance-Aware U-Net 8 upvotes, #18 of 2025-08-07
  19. HarmonyGuard: Toward Safety and Utility in Web Agents via Adaptive Policy Enhancement and Dual-Objective Optimization 8 upvotes, #18 of 2025-08-07
  20. StepFun-Formalizer: Unlocking the Autoformalization Potential of LLMs through Knowledge-Reasoning Fusion 8 upvotes, #18 of 2025-08-07
  21. OpenMed NER: Open-Source, Domain-Adapted State-of-the-Art Transformers for Biomedical NER Across 12 Public Datasets 7 upvotes, #21 of 2025-08-07
  22. MiDashengLM: Efficient Audio Understanding with General Audio Captions 7 upvotes, #21 of 2025-08-07
  23. RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization 6 upvotes, #23 of 2025-08-07
  24. DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose Prior 6 upvotes, #23 of 2025-08-07
  25. EVOC2RUST: A Skeleton-guided Framework for Project-Level C-to-Rust Translation 6 upvotes, #23 of 2025-08-07
  26. Reasoning Language Models for Root Cause Analysis in 5G Wireless Networks 4 upvotes, #26 of 2025-08-07
  27. FACTORY: A Challenging Human-Verified Prompt Set for Long-Form Factuality 4 upvotes, #26 of 2025-08-07
  28. A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding 4 upvotes, #26 of 2025-08-07
  29. Sel3DCraft: Interactive Visual Prompts for User-Friendly Text-to-3D Generation 3 upvotes, #29 of 2025-08-07
  30. DiffSemanticFusion: Semantic Raster BEV Fusion for Autonomous Driving via Online HD Map Diffusion 3 upvotes, #29 of 2025-08-07
  31. C3D-AD: Toward Continual 3D Anomaly Detection via Kernel Attention with Learnable Advisor 2 upvotes, #31 of 2025-08-07
  32. Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following 2 upvotes, #31 of 2025-08-07
  33. IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards 2 upvotes, #31 of 2025-08-07
  34. The Cow of Rembrandt - Analyzing Artistic Prompt Interpretation in Text-to-Image Models 1 upvotes, #34 of 2025-08-07
  35. CM^3: Calibrating Multimodal Recommendation 1 upvotes, #34 of 2025-08-07
  36. SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering 1 upvotes, #34 of 2025-08-07
  37. Data and AI governance: Promoting equity, ethics, and fairness in large language models 1 upvotes, #34 of 2025-08-07
  38. MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine 1 upvotes, #38 of 2025-08-07

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.