Daily Papers of 2025-10-02
- DeepSearch: Overcome the Bottleneck of Reinforcement Learning with Verifiable Rewards via Monte Carlo Tree Search 123 upvotes, #1 of 2025-10-02
- GEM: A Gym for Agentic LLMs 79 upvotes, #2 of 2025-10-02
- SINQ: Sinkhorn-Normalized Quantization for Calibration-Free Low-Precision LLM Weights 73 upvotes, #3 of 2025-10-02
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators 60 upvotes, #4 of 2025-10-02
- Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation 43 upvotes, #5 of 2025-10-02
- PIPer: On-Device Environment Setup via Online Reinforcement Learning 31 upvotes, #6 of 2025-10-02
- Code2Video: A Code-centric Paradigm for Educational Video Generation 29 upvotes, #7 of 2025-10-02
- ACON: Optimizing Context Compression for Long-horizon LLM Agents 28 upvotes, #8 of 2025-10-02
- It Takes Two: Your GRPO Is Secretly DPO 28 upvotes, #8 of 2025-10-02
- BroRL: Scaling Reinforcement Learning via Broadened Exploration 16 upvotes, #10 of 2025-10-02
- Flash-Searcher: Fast and Effective Web Agents via DAG-Based Parallel Execution 15 upvotes, #11 of 2025-10-02
- EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing 15 upvotes, #11 of 2025-10-02
- Why Can't Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls 15 upvotes, #11 of 2025-10-02
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses 15 upvotes, #11 of 2025-10-02
- QUASAR: Quantum Assembly Code Generation Using Tool-Augmented LLMs via Agentic RL 11 upvotes, #15 of 2025-10-02
- Beyond Log Likelihood: Probability-Based Objectives for Supervised Fine-Tuning across the Model Capability Continuum 8 upvotes, #16 of 2025-10-02
- On Predictability of Reinforcement Learning Dynamics for Large Language Models 8 upvotes, #16 of 2025-10-02
- Making, not Taking, the Best of N 8 upvotes, #16 of 2025-10-02
- MixtureVitae: Open Web-Scale Pretraining Dataset With High Quality Instruction and Reasoning Data Built from Permissive-First Text Sources 6 upvotes, #19 of 2025-10-02
- GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness 6 upvotes, #19 of 2025-10-02
- Infusing Theory of Mind into Socially Intelligent LLM Agents 5 upvotes, #21 of 2025-10-02
- Training Vision-Language Process Reward Models for Test-Time Scaling in Multimodal Reasoning: Key Insights and Lessons Learned 5 upvotes, #21 of 2025-10-02
- Pay-Per-Search Models are Abstention Models 5 upvotes, #21 of 2025-10-02
- An Empirical Study of Testing Practices in Open Source AI Agent Frameworks and Agentic Applications 3 upvotes, #24 of 2025-10-02
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs 3 upvotes, #24 of 2025-10-02
- BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs 3 upvotes, #24 of 2025-10-02
- JoyAgent-JDGenie: Technical Report on the GAIA 3 upvotes, #24 of 2025-10-02
- Eliciting Secret Knowledge from Language Models 3 upvotes, #24 of 2025-10-02
- Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures 2 upvotes, #29 of 2025-10-02
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models 2 upvotes, #29 of 2025-10-02
- Boolean Satisfiability via Imitation Learning 2 upvotes, #29 of 2025-10-02
- TGPO: Temporal Grounded Policy Optimization for Signal Temporal Logic Tasks 2 upvotes, #29 of 2025-10-02
- BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration 2 upvotes, #29 of 2025-10-02
- In-Place Feedback: A New Paradigm for Guiding LLMs in Multi-Turn Reasoning 2 upvotes, #29 of 2025-10-02
- CurES: From Gradient Analysis to Efficient Curriculum Learning for Reasoning LLMs 2 upvotes, #29 of 2025-10-02
- ReSWD: ReSTIR'd, not shaken. Combining Reservoir Sampling and Sliced Wasserstein Distance for Variance Reduction 2 upvotes, #29 of 2025-10-02
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.