Daily Papers of 2025-09-29
- LongLive: Real-time Interactive Long Video Generation 170 upvotes, #1 of 2025-09-29
- EPO: Entropy-regularized Policy Optimization for LLM Agents Reinforcement Learning 124 upvotes, #2 of 2025-09-29
- Quantile Advantage Estimation for Entropy-Safe Reasoning 113 upvotes, #3 of 2025-09-29
- MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing 100 upvotes, #4 of 2025-09-29
- Variational Reasoning for Language Models 66 upvotes, #5 of 2025-09-29
- Language Models Can Learn from Verbal Feedback Without Scalar Rewards 64 upvotes, #6 of 2025-09-29
- ReviewScore: Misinformed Peer Review Detection with Large Language Models 62 upvotes, #7 of 2025-09-29
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training 38 upvotes, #8 of 2025-09-29
- No Prompt Left Behind: Exploiting Zero-Variance Prompts in LLM Reinforcement Learning via Entropy-Guided Advantage Shaping 37 upvotes, #9 of 2025-09-29
- PromptCoT 2.0: Scaling Prompt Synthesis for Large Language Model Reasoning 31 upvotes, #10 of 2025-09-29
- CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning 31 upvotes, #10 of 2025-09-29
- MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning 30 upvotes, #12 of 2025-09-29
- Fine-tuning Done Right in Model Editing 27 upvotes, #13 of 2025-09-29
- UltraHorizon: Benchmarking Agent Capabilities in Ultra Long-Horizon Scenarios 23 upvotes, #14 of 2025-09-29
- See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation 23 upvotes, #14 of 2025-09-29
- Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation 21 upvotes, #16 of 2025-09-29
- LucidFlux: Caption-Free Universal Image Restoration via a Large-Scale Diffusion Transformer 21 upvotes, #16 of 2025-09-29
- VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing 21 upvotes, #16 of 2025-09-29
- COSPADI: Compressing LLMs via Calibration-Guided Sparse Dictionary Learning 20 upvotes, #19 of 2025-09-29
- WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning 18 upvotes, #20 of 2025-09-29
- Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training 16 upvotes, #21 of 2025-09-29
- Think-on-Graph 3.0: Efficient and Adaptive LLM Reasoning on Heterogeneous Graphs via Multi-Agent Dual-Evolving Context Retrieval 16 upvotes, #21 of 2025-09-29
- SPARK: Synergistic Policy And Reward Co-Evolving Framework 16 upvotes, #21 of 2025-09-29
- TUN3D: Towards Real-World Scene Understanding from Unposed Images 14 upvotes, #24 of 2025-09-29
- UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models 12 upvotes, #25 of 2025-09-29
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning 10 upvotes, #26 of 2025-09-29
- WoW: Towards a World omniscient World model Through Embodied Interaction 9 upvotes, #27 of 2025-09-29
- Real-Time Object Detection Meets DINOv3 8 upvotes, #28 of 2025-09-29
- D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents 8 upvotes, #28 of 2025-09-29
- X-Streamer: Unified Human World Modeling with Audiovisual Interaction 6 upvotes, #30 of 2025-09-29
- ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models 5 upvotes, #31 of 2025-09-29
- FlashEdit: Decoupling Speed, Structure, and Semantics for Precise Image Editing 5 upvotes, #31 of 2025-09-29
- The role of synthetic data in Multilingual, Multi-cultural AI systems: Lessons from Indic Languages 3 upvotes, #33 of 2025-09-29
- RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards 3 upvotes, #33 of 2025-09-29
- Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation 3 upvotes, #33 of 2025-09-29
- RefAM: Attention Magnets for Zero-Shot Referral Segmentation 3 upvotes, #33 of 2025-09-29
- CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition 2 upvotes, #37 of 2025-09-29
- X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning 2 upvotes, #37 of 2025-09-29
- HiGS: History-Guided Sampling for Plug-and-Play Enhancement of Diffusion Models 2 upvotes, #37 of 2025-09-29
- StateX: Enhancing RNN Recall via Post-training State Expansion 2 upvotes, #37 of 2025-09-29
- Scale-Wise VAR is Secretly Discrete Diffusion 2 upvotes, #37 of 2025-09-29
- Finding 3D Positions of Distant Objects from Noisy Camera Movement and Semantic Segmentation Sequences 1 upvotes, #42 of 2025-09-29
- CAD-Tokenizer: Towards Text-based CAD Prototyping via Modality-Specific Tokenization 1 upvotes, #42 of 2025-09-29
- Instruction-Following Evaluation in Function Calling for Large Language Models 3 upvotes, #44 of 2025-09-29
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.