Daily Papers of 2025-10-02

  1. DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search 123 upvotes, #1 of 2025-10-02
  2. GEM: A Gym for Agentic LLMs 79 upvotes, #2 of 2025-10-02
  3. SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights 73 upvotes, #3 of 2025-10-02
  4. VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators 60 upvotes, #4 of 2025-10-02
  5. Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation 43 upvotes, #5 of 2025-10-02
  6. PIPer: On-Device Environment Setup via Online Reinforcement Learning 31 upvotes, #6 of 2025-10-02
  7. Code2Video: A Code-centric Paradigm for Educational Video Generation 29 upvotes, #7 of 2025-10-02
  8. ACON: Optimizing Context Compression for Long-horizon LLM Agents 28 upvotes, #8 of 2025-10-02
  9. It Takes Two: Your GRPO Is Secretly DPO 28 upvotes, #8 of 2025-10-02
  10. BroRL: Scaling Reinforcement Learning via Broadened Exploration 16 upvotes, #10 of 2025-10-02
  11. Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution 15 upvotes, #11 of 2025-10-02
  12. EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing 15 upvotes, #11 of 2025-10-02
  13. Why Can't Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls 15 upvotes, #11 of 2025-10-02
  14. BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses 15 upvotes, #11 of 2025-10-02
  15. QUASAR: Quantum Assembly Code Generation Using Tool-Augmented LLMs via Agentic RL 11 upvotes, #15 of 2025-10-02
  16. Beyond Log Likelihood: Probability-Based Objectives for Supervised Fine-Tuning across the Model Capability Continuum 8 upvotes, #16 of 2025-10-02
  17. On Predictability of Reinforcement Learning Dynamics for Large Language Models 8 upvotes, #16 of 2025-10-02
  18. Making, not Taking, the Best of N 8 upvotes, #16 of 2025-10-02
  19. MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources 6 upvotes, #19 of 2025-10-02
  20. GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness 6 upvotes, #19 of 2025-10-02
  21. Infusing Theory of Mind into Socially Intelligent LLM Agents 5 upvotes, #21 of 2025-10-02
  22. Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned 5 upvotes, #21 of 2025-10-02
  23. Pay-Per-Search Models are Abstention Models 5 upvotes, #21 of 2025-10-02
  24. An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications 3 upvotes, #24 of 2025-10-02
  25. VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs 3 upvotes, #24 of 2025-10-02
  26. BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs 3 upvotes, #24 of 2025-10-02
  27. JoyAgent-JDGenie: Technical Report on the GAIA 3 upvotes, #24 of 2025-10-02
  28. Eliciting Secret Knowledge from Language Models 3 upvotes, #24 of 2025-10-02
  29. Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures 2 upvotes, #29 of 2025-10-02
  30. Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models 2 upvotes, #29 of 2025-10-02
  31. Boolean Satisfiability via Imitation Learning 2 upvotes, #29 of 2025-10-02
  32. TGPO: Temporal Grounded Policy Optimization for Signal Temporal Logic Tasks 2 upvotes, #29 of 2025-10-02
  33. BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration 2 upvotes, #29 of 2025-10-02
  34. In-Place Feedback: A New Paradigm for Guiding LLMs in Multi-Turn Reasoning 2 upvotes, #29 of 2025-10-02
  35. CurES: From Gradient Analysis to Efficient Curriculum Learning for Reasoning LLMs 2 upvotes, #29 of 2025-10-02
  36. ReSWD: ReSTIR'd, not shaken. Combining Reservoir Sampling and Sliced Wasserstein Distance for Variance Reduction 2 upvotes, #29 of 2025-10-02

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.