Daily Papers of 2025-10-03

  1. LongCodeZip: Compress Long Context for Code Language Models 102 upvotes, #1 of 2025-10-03
  2. Self-Forcing++: Towards Minute-Scale High-Quality Video Generation 86 upvotes, #2 of 2025-10-03
  3. ExGRPO: Learning to Reason from Experience 72 upvotes, #3 of 2025-10-03
  4. StealthAttack: Robust 3D Gaussian Splatting Poisoning via Density-Guided Illusions 56 upvotes, #4 of 2025-10-03
  5. StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets? 47 upvotes, #5 of 2025-10-03
  6. F2LLM Technical Report: Matching SOTA Embedding Performance with 6 Million Open-Source Data 40 upvotes, #6 of 2025-10-03
  7. Interactive Training: Feedback-Driven Neural Network Optimization 38 upvotes, #7 of 2025-10-03
  8. RLP: Reinforcement as a Pretraining Objective 34 upvotes, #8 of 2025-10-03
  9. ModernVBERT: Towards Smaller Visual Document Retrievers 29 upvotes, #9 of 2025-10-03
  10. Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacks 28 upvotes, #10 of 2025-10-03
  11. Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation 26 upvotes, #11 of 2025-10-03
  12. TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments 24 upvotes, #12 of 2025-10-03
  13. CLUE: Non-parametric Verification from Experience via Hidden-State Clustering 22 upvotes, #13 of 2025-10-03
  14. The Unreasonable Effectiveness of Scaling Agents for Computer Use 22 upvotes, #13 of 2025-10-03
  15. The Rogue Scalpel: Activation Steering Compromises LLM Safety 21 upvotes, #15 of 2025-10-03
  16. VOGUE: Guiding Exploration with Visual Uncertainty Improves Multimodal Reasoning 19 upvotes, #16 of 2025-10-03
  17. Learning to Reason for Hallucination Span Detection 18 upvotes, #17 of 2025-10-03
  18. A Rigorous Benchmark with Multidimensional Evaluation for Deep Research Agents: From Answers to Reports 18 upvotes, #17 of 2025-10-03
  19. Aristotle: IMO-level Automated Theorem Proving 16 upvotes, #19 of 2025-10-03
  20. RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning 16 upvotes, #19 of 2025-10-03
  21. Sparse Query Attention (SQA): A Computationally Efficient Attention Mechanism with Query Heads Reduction 12 upvotes, #21 of 2025-10-03
  22. DragFlow: Unleashing DiT Priors with Region Based Supervision for Drag Editing 12 upvotes, #21 of 2025-10-03
  23. Fine-Grained Detection of Context-Grounded Hallucinations Using LLMs 10 upvotes, #23 of 2025-10-03
  24. Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow 9 upvotes, #24 of 2025-10-03
  25. Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models 9 upvotes, #24 of 2025-10-03
  26. VideoNSA: Native Sparse Attention Scales Video Understanding 9 upvotes, #24 of 2025-10-03
  27. Go with Your Gut: Scaling Confidence for Autoregressive Image Generation 8 upvotes, #27 of 2025-10-03
  28. Automated Structured Radiology Report Generation with Rich Clinical Context 7 upvotes, #28 of 2025-10-03
  29. VLA-R1: Enhancing Reasoning in Vision-Language-Action Models 7 upvotes, #28 of 2025-10-03
  30. RLAD: Training LLMs to Discover Abstractions for Solving Reasoning Problems 7 upvotes, #28 of 2025-10-03
  31. Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends 6 upvotes, #31 of 2025-10-03
  32. VIRTUE: Visual-Interactive Text-Image Universal Embedder 6 upvotes, #31 of 2025-10-03
  33. Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness 6 upvotes, #31 of 2025-10-03
  34. Transformers Discover Molecular Structure Without Graph Priors 6 upvotes, #31 of 2025-10-03
  35. TimeSeriesScientist: A General-Purpose AI Agent for Time Series Analysis 5 upvotes, #35 of 2025-10-03
  36. Optimal Control Meets Flow Matching: A Principled Route to Multi-Subject Fidelity 5 upvotes, #35 of 2025-10-03
  37. FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting 4 upvotes, #37 of 2025-10-03
  38. One-Token Rollout: Guiding Supervised Fine-Tuning of LLMs with Policy Gradient 4 upvotes, #37 of 2025-10-03
  39. Rethinking Thinking Tokens: LLMs as Improvement Operators 4 upvotes, #37 of 2025-10-03
  40. Generalized Parallel Scaling with Interdependent Generations 4 upvotes, #37 of 2025-10-03
  41. SKYLENAGE Technical Report: Mathematical Reasoning and Contest-Innovation Benchmarks for Multi-Level Math Evaluation 3 upvotes, #41 of 2025-10-03
  42. Rethinking the shape convention of an MLP 3 upvotes, #41 of 2025-10-03
  43. Parallel Scaling Law: Unveiling Reasoning Generalization through A Cross-Linguistic Perspective 3 upvotes, #41 of 2025-10-03
  44. Drawing Conclusions from Draws: Rethinking Preference Semantics in Arena-Style LLM Evaluation 3 upvotes, #41 of 2025-10-03
  45. Controlled Generation for Private Synthetic Text 2 upvotes, #45 of 2025-10-03
  46. Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval 2 upvotes, #45 of 2025-10-03
  47. IoT-MCP: Bridging LLMs and IoT Systems Through Model Context Protocol 2 upvotes, #45 of 2025-10-03
  48. MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs 2 upvotes, #45 of 2025-10-03
  49. SQUARE: Semantic Query-Augmented Fusion and Efficient Batch Reranking for Training-free Zero-Shot Composed Image Retrieval 1 upvotes, #49 of 2025-10-03
  50. Spectral Scaling Laws in Language Models: How Effectively Do Feed-Forward Networks Use Their Latent Space? 1 upvotes, #49 of 2025-10-03
  51. AReUReDi: Annealed Rectified Updates for Refining Discrete Flows with Multi-Objective Guidance 1 upvotes, #51 of 2025-10-03
  52. Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression 2 upvotes, #51 of 2025-10-03

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.