Daily Papers of 2025-10-03
- LongCodeZip: Compress Long Context for Code Language Models 102 upvotes, #1 of 2025-10-03
- Self-Forcing++: Towards Minute-Scale High-Quality Video Generation 86 upvotes, #2 of 2025-10-03
- ExGRPO: Learning to Reason from Experience 72 upvotes, #3 of 2025-10-03
- StealthAttack: Robust 3D Gaussian Splatting Poisoning via Density-Guided Illusions 56 upvotes, #4 of 2025-10-03
- StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets? 47 upvotes, #5 of 2025-10-03
- F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data 40 upvotes, #6 of 2025-10-03
- Interactive Training: Feedback-Driven Neural Network Optimization 38 upvotes, #7 of 2025-10-03
- RLP: Reinforcement as a Pretraining Objective 34 upvotes, #8 of 2025-10-03
- ModernVBERT: Towards Smaller Visual Document Retrievers 29 upvotes, #9 of 2025-10-03
- Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks 28 upvotes, #10 of 2025-10-03
- Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation 26 upvotes, #11 of 2025-10-03
- TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments 24 upvotes, #12 of 2025-10-03
- CLUE: Non-parametric Verification from Experience via Hidden-State Clustering 22 upvotes, #13 of 2025-10-03
- The Unreasonable Effectiveness of Scaling Agents for Computer Use 22 upvotes, #13 of 2025-10-03
- The Rogue Scalpel: Activation Steering Compromises LLM Safety 21 upvotes, #15 of 2025-10-03
- VOGUE: Guiding Exploration with Visual Uncertainty Improves Multimodal Reasoning 19 upvotes, #16 of 2025-10-03
- Learning to Reason for Hallucination Span Detection 18 upvotes, #17 of 2025-10-03
- A Rigorous Benchmark with Multidimensional Evaluation for Deep Research Agents: From Answers to Reports 18 upvotes, #17 of 2025-10-03
- Aristotle: IMO-level Automated Theorem Proving 16 upvotes, #19 of 2025-10-03
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning 16 upvotes, #19 of 2025-10-03
- Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction 12 upvotes, #21 of 2025-10-03
- DragFlow: Unleashing DiT Priors with Region Based Supervision for Drag Editing 12 upvotes, #21 of 2025-10-03
- Fine-Grained Detection of Context-Grounded Hallucinations Using LLMs 10 upvotes, #23 of 2025-10-03
- Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow 9 upvotes, #24 of 2025-10-03
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models 9 upvotes, #24 of 2025-10-03
- VideoNSA: Native Sparse Attention Scales Video Understanding 9 upvotes, #24 of 2025-10-03
- Go with Your Gut: Scaling Confidence for Autoregressive Image Generation 8 upvotes, #27 of 2025-10-03
- Automated Structured Radiology Report Generation with Rich Clinical Context 7 upvotes, #28 of 2025-10-03
- VLA-R1: Enhancing Reasoning in Vision-Language-Action Models 7 upvotes, #28 of 2025-10-03
- RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems 7 upvotes, #28 of 2025-10-03
- Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends 6 upvotes, #31 of 2025-10-03
- VIRTUE: Visual-Interactive Text-Image Universal Embedder 6 upvotes, #31 of 2025-10-03
- Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness 6 upvotes, #31 of 2025-10-03
- Transformers Discover Molecular Structure Without Graph Priors 6 upvotes, #31 of 2025-10-03
- TimeSeriesScientist: A General-Purpose AI Agent for Time Series Analysis 5 upvotes, #35 of 2025-10-03
- Optimal Control Meets Flow Matching: A Principled Route to Multi-Subject Fidelity 5 upvotes, #35 of 2025-10-03
- FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting 4 upvotes, #37 of 2025-10-03
- One-Token Rollout: Guiding Supervised Fine-Tuning of LLMs with Policy Gradient 4 upvotes, #37 of 2025-10-03
- Rethinking Thinking Tokens: LLMs as Improvement Operators 4 upvotes, #37 of 2025-10-03
- Generalized Parallel Scaling with Interdependent Generations 4 upvotes, #37 of 2025-10-03
- SKYLENAGE Technical Report: Mathematical Reasoning and Contest-Innovation Benchmarks for Multi-Level Math Evaluation 3 upvotes, #41 of 2025-10-03
- Rethinking the shape convention of an MLP 3 upvotes, #41 of 2025-10-03
- Parallel Scaling Law: Unveiling Reasoning Generalization through A Cross-Linguistic Perspective 3 upvotes, #41 of 2025-10-03
- Drawing Conclusions from Draws: Rethinking Preference Semantics in Arena-Style LLM Evaluation 3 upvotes, #41 of 2025-10-03
- Controlled Generation for Private Synthetic Text 2 upvotes, #45 of 2025-10-03
- Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval 2 upvotes, #45 of 2025-10-03
- IoT-MCP: Bridging LLMs and IoT Systems Through Model Context Protocol 2 upvotes, #45 of 2025-10-03
- MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs 2 upvotes, #45 of 2025-10-03
- SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval 1 upvotes, #49 of 2025-10-03
- Spectral Scaling Laws in Language Models: How Effectively Do Feed-Forward Networks Use Their Latent Space? 1 upvotes, #49 of 2025-10-03
- AReUReDi: Annealed Rectified Updates for Refining Discrete Flows with Multi-Objective Guidance 1 upvotes, #51 of 2025-10-03
- Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression 2 upvotes, #51 of 2025-10-03
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.