Daily Papers of 2025-09-03

  1. The Landscape of Agentic Reinforcement Learning for LLMs: A Survey 177 upvotes, #1 of 2025-09-03
  2. UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning 112 upvotes, #2 of 2025-09-03
  3. SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning 80 upvotes, #3 of 2025-09-03
  4. LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model 76 upvotes, #4 of 2025-09-03
  5. VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use 64 upvotes, #5 of 2025-09-03
  6. Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic 54 upvotes, #6 of 2025-09-03
  7. ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding 53 upvotes, #7 of 2025-09-03
  8. POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion 46 upvotes, #8 of 2025-09-03
  9. Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling 41 upvotes, #9 of 2025-09-03
  10. Baichuan-M2: Scaling Medical Capability with Large Verifier System 36 upvotes, #10 of 2025-09-03
  11. Kwai Keye-VL 1.5 Technical Report 33 upvotes, #11 of 2025-09-03
  12. OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning 28 upvotes, #12 of 2025-09-03
  13. Jointly Reinforcing Diversity and Quality in Language Model Generations 25 upvotes, #13 of 2025-09-03
  14. Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR 24 upvotes, #14 of 2025-09-03
  15. Benchmarking Optimizers for Large Language Model Pretraining 23 upvotes, #15 of 2025-09-03
  16. GenCompositor: Generative Video Compositing with Diffusion Transformer 23 upvotes, #15 of 2025-09-03
  17. FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games 19 upvotes, #17 of 2025-09-03
  18. DCPO: Dynamic Clipping Policy Optimization 19 upvotes, #17 of 2025-09-03
  19. DynaGuard: A Dynamic Guardrail Model With User-Defined Policies 18 upvotes, #19 of 2025-09-03
  20. On the Theoretical Limitations of Embedding-Based Retrieval 17 upvotes, #20 of 2025-09-03
  21. Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation 14 upvotes, #21 of 2025-09-03
  22. Universal Deep Research: Bring Your Own Model and Strategy 12 upvotes, #22 of 2025-09-03
  23. M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision 11 upvotes, #23 of 2025-09-03
  24. Fantastic Pretraining Optimizers and Where to Find Them 11 upvotes, #23 of 2025-09-03
  25. The Gold Medals in an Empty Room: Diagnosing Metalinguistic Reasoning in LLMs with Camlang 10 upvotes, #25 of 2025-09-03
  26. SQL-of-Thought: Multi-agentic Text-to-SQL with Guided Error Correction 7 upvotes, #26 of 2025-09-03
  27. MobiAgent: A Systematic Framework for Customizable Mobile Agents 6 upvotes, #27 of 2025-09-03
  28. ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association 6 upvotes, #27 of 2025-09-03
  29. Metis: Training Large Language Models with Advanced Low-Bit Quantization 5 upvotes, #29 of 2025-09-03
  30. Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing 5 upvotes, #29 of 2025-09-03
  31. Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs 3 upvotes, #31 of 2025-09-03
  32. AMBEDKAR-A Multi-level Bias Elimination through a Decoding Approach with Knowledge Augmentation for Robust Constitutional Alignment of Language Models 3 upvotes, #31 of 2025-09-03
  33. Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices 3 upvotes, #31 of 2025-09-03
  34. FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models 2 upvotes, #34 of 2025-09-03
  35. Stairway to Fairness: Connecting Group and Individual Fairness 2 upvotes, #34 of 2025-09-03
  36. Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views 2 upvotes, #34 of 2025-09-03
  37. Improving Large Vision and Language Models by Learning from a Panel of Peers 2 upvotes, #34 of 2025-09-03
  38. MedDINOv3: How to adapt vision foundation models for medical image segmentation? 2 upvotes, #34 of 2025-09-03
  39. C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Object Detection 1 upvotes, #39 of 2025-09-03

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.