Daily Papers of 2025-09-03
- The Landscape of Agentic Reinforcement Learning for LLMs: A Survey 177 upvotes, #1 of 2025-09-03
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning 112 upvotes, #2 of 2025-09-03
- SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning 80 upvotes, #3 of 2025-09-03
- LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model 76 upvotes, #4 of 2025-09-03
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use 64 upvotes, #5 of 2025-09-03
- Reasoning Vectors: Transferring Chain-of-Thought Capabilities via Task Arithmetic 54 upvotes, #6 of 2025-09-03
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding 53 upvotes, #7 of 2025-09-03
- POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion 46 upvotes, #8 of 2025-09-03
- Gated Associative Memory: A Parallel O(N) Architecture for Efficient Sequence Modeling 41 upvotes, #9 of 2025-09-03
- Baichuan-M2: Scaling Medical Capability with Large Verifier System 36 upvotes, #10 of 2025-09-03
- Kwai Keye-VL 1.5 Technical Report 33 upvotes, #11 of 2025-09-03
- OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning 28 upvotes, #12 of 2025-09-03
- Jointly Reinforcing Diversity and Quality in Language Model Generations 25 upvotes, #13 of 2025-09-03
- Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR 24 upvotes, #14 of 2025-09-03
- Benchmarking Optimizers for Large Language Model Pretraining 23 upvotes, #15 of 2025-09-03
- GenCompositor: Generative Video Compositing with Diffusion Transformer 23 upvotes, #15 of 2025-09-03
- FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games 19 upvotes, #17 of 2025-09-03
- DCPO: Dynamic Clipping Policy Optimization 19 upvotes, #17 of 2025-09-03
- DynaGuard: A Dynamic Guardrail Model With User-Defined Policies 18 upvotes, #19 of 2025-09-03
- On the Theoretical Limitations of Embedding-Based Retrieval 17 upvotes, #20 of 2025-09-03
- Attributes as Textual Genes: Leveraging LLMs as Genetic Algorithm Simulators for Conditional Synthetic Data Generation 14 upvotes, #21 of 2025-09-03
- Universal Deep Research: Bring Your Own Model and Strategy 12 upvotes, #22 of 2025-09-03
- M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision 11 upvotes, #23 of 2025-09-03
- Fantastic Pretraining Optimizers and Where to Find Them 11 upvotes, #23 of 2025-09-03
- The Gold Medals in an Empty Room: Diagnosing Metalinguistic Reasoning in LLMs with Camlang 10 upvotes, #25 of 2025-09-03
- SQL-of-Thought: Multi-agentic Text-to-SQL with Guided Error Correction 7 upvotes, #26 of 2025-09-03
- MobiAgent: A Systematic Framework for Customizable Mobile Agents 6 upvotes, #27 of 2025-09-03
- ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association 6 upvotes, #27 of 2025-09-03
- Metis: Training Large Language Models with Advanced Low-Bit Quantization 5 upvotes, #29 of 2025-09-03
- Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing 5 upvotes, #29 of 2025-09-03
- Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs 3 upvotes, #31 of 2025-09-03
- AMBEDKAR-A Multi-level Bias Elimination through a Decoding Approach with Knowledge Augmentation for Robust Constitutional Alignment of Language Models 3 upvotes, #31 of 2025-09-03
- Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices 3 upvotes, #31 of 2025-09-03
- FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models 2 upvotes, #34 of 2025-09-03
- Stairway to Fairness: Connecting Group and Individual Fairness 2 upvotes, #34 of 2025-09-03
- Towards More Diverse and Challenging Pre-training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled Views 2 upvotes, #34 of 2025-09-03
- Improving Large Vision and Language Models by Learning from a Panel of Peers 2 upvotes, #34 of 2025-09-03
- MedDINOv3: How to adapt vision foundation models for medical image segmentation? 2 upvotes, #34 of 2025-09-03
- C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Object Detection 1 upvotes, #39 of 2025-09-03
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.