Daily Papers of 2025-09-30

  1. SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention 109 upvotes, #1 of 2025-09-30
  2. StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs 62 upvotes, #2 of 2025-09-30
  3. Multiplayer Nash Preference Optimization 60 upvotes, #3 of 2025-09-30
  4. OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing 49 upvotes, #4 of 2025-09-30
  5. Beyond the Exploration-Exploitation Trade-off: A Hidden State Approach for LLM Reasoning in RLVR 47 upvotes, #5 of 2025-09-30
  6. RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark 44 upvotes, #6 of 2025-09-30
  7. Democratizing AI scientists using ToolUniverse 38 upvotes, #7 of 2025-09-30
  8. SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer 38 upvotes, #7 of 2025-09-30
  9. When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance 37 upvotes, #9 of 2025-09-30
  10. Sequential Diffusion Language Models 36 upvotes, #10 of 2025-09-30
  11. Visual Jigsaw Post-Training Improves MLLMs 34 upvotes, #11 of 2025-09-30
  12. SparseD: Sparse Attention for Diffusion Language Models 29 upvotes, #12 of 2025-09-30
  13. Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards 27 upvotes, #13 of 2025-09-30
  14. Towards Personalized Deep Research: Benchmarks and Evaluations 27 upvotes, #13 of 2025-09-30
  15. GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts 27 upvotes, #13 of 2025-09-30
  16. EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling 26 upvotes, #16 of 2025-09-30
  17. EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering 26 upvotes, #16 of 2025-09-30
  18. VideoScore2: Think before You Score in Generative Video Evaluation 22 upvotes, #18 of 2025-09-30
  19. Rolling Forcing: Autoregressive Long Video Diffusion in Real Time 21 upvotes, #19 of 2025-09-30
  20. Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning 20 upvotes, #20 of 2025-09-30
  21. HunyuanImage 3.0 Technical Report 20 upvotes, #20 of 2025-09-30
  22. WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning 18 upvotes, #22 of 2025-09-30
  23. From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old Ones 17 upvotes, #23 of 2025-09-30
  24. VGGT-X: When VGGT Meets Dense Novel View Synthesis 17 upvotes, #23 of 2025-09-30
  25. Socratic-Zero : Bootstrapping Reasoning via Data-Free Agent Co-evolution 16 upvotes, #25 of 2025-09-30
  26. Scaling Generalist Data-Analytic Agents 16 upvotes, #25 of 2025-09-30
  27. The Era of Real-World Human Interaction: RL from User Conversations 16 upvotes, #25 of 2025-09-30
  28. Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks 15 upvotes, #28 of 2025-09-30
  29. MMPB: It's Time for Multi-Modal Personalization 14 upvotes, #29 of 2025-09-30
  30. Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding 13 upvotes, #30 of 2025-09-30
  31. Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference Learning 13 upvotes, #30 of 2025-09-30
  32. BRIDGE - Building Reinforcement-Learning Depth-to-Image Data Generation Engine for Monocular Depth Estimation 13 upvotes, #30 of 2025-09-30
  33. MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech 13 upvotes, #30 of 2025-09-30
  34. InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation 12 upvotes, #34 of 2025-09-30
  35. SIRI: Scaling Iterative Reinforcement Learning with Interleaved Compression 12 upvotes, #34 of 2025-09-30
  36. Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time 11 upvotes, #36 of 2025-09-30
  37. Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective 11 upvotes, #36 of 2025-09-30
  38. Where LLM Agents Fail and How They can Learn From Failures 11 upvotes, #36 of 2025-09-30
  39. Pretraining Large Language Models with NVFP4 10 upvotes, #39 of 2025-09-30
  40. LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals 9 upvotes, #40 of 2025-09-30
  41. From Harm to Help: Turning Reasoning In-Context Demos into Assets for Reasoning LMs 9 upvotes, #40 of 2025-09-30
  42. Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation 9 upvotes, #40 of 2025-09-30
  43. GRPO-MA: Multi-Answer Generation in GRPO for Stable and Efficient Chain-of-Thought Training 8 upvotes, #43 of 2025-09-30
  44. DepthLM: Metric Depth From Vision Language Models 8 upvotes, #43 of 2025-09-30
  45. Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment 7 upvotes, #45 of 2025-09-30
  46. Taming Masked Diffusion Language Models via Consistency Trajectory Reinforcement Learning with Fewer Decoding Step 7 upvotes, #45 of 2025-09-30
  47. SCI-Verifier: Scientific Verifier with Thinking 7 upvotes, #45 of 2025-09-30
  48. MultiCrafter: High-Fidelity Multi-Subject Generation via Spatially Disentangled Attention and Identity-Aware Reinforcement Learning 6 upvotes, #48 of 2025-09-30
  49. Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification 6 upvotes, #48 of 2025-09-30
  50. AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play 6 upvotes, #48 of 2025-09-30
  51. Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers 6 upvotes, #48 of 2025-09-30
  52. Hyperspherical Latents Improve Continuous-Token Autoregressive Generation 6 upvotes, #48 of 2025-09-30
  53. DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space 6 upvotes, #48 of 2025-09-30
  54. Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization 5 upvotes, #54 of 2025-09-30
  55. LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning 5 upvotes, #54 of 2025-09-30
  56. The Photographer Eye: Teaching Multimodal Large Language Models to See and Critique like Photographers 4 upvotes, #56 of 2025-09-30
  57. ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents 4 upvotes, #56 of 2025-09-30
  58. MathBode: Frequency-Domain Fingerprints of LLM Mathematical Reasoning 4 upvotes, #56 of 2025-09-30
  59. PARROT: A Benchmark for Evaluating LLMs in Cross-System SQL Translation 4 upvotes, #56 of 2025-09-30
  60. Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning 4 upvotes, #56 of 2025-09-30
  61. Scalable GANs with Transformers 4 upvotes, #56 of 2025-09-30
  62. Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models 4 upvotes, #56 of 2025-09-30
  63. PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images 4 upvotes, #56 of 2025-09-30
  64. UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration 3 upvotes, #64 of 2025-09-30
  65. TENET: Leveraging Tests Beyond Validation for Code Generation 3 upvotes, #64 of 2025-09-30
  66. UniVid: The Open-Source Unified Video Model 3 upvotes, #64 of 2025-09-30
  67. AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models 3 upvotes, #64 of 2025-09-30
  68. IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video? 3 upvotes, #64 of 2025-09-30
  69. ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation 3 upvotes, #64 of 2025-09-30
  70. Charting a Decade of Computational Linguistics in Italy: The CLiC-it Corpus 2 upvotes, #70 of 2025-09-30
  71. Combinatorial Creativity: A New Frontier in Generalization Abilities 2 upvotes, #70 of 2025-09-30
  72. REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model 2 upvotes, #70 of 2025-09-30
  73. ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning 2 upvotes, #70 of 2025-09-30
  74. RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human Mobility 2 upvotes, #70 of 2025-09-30
  75. BPMN Assistant: An LLM-Based Approach to Business Process Modeling 2 upvotes, #70 of 2025-09-30
  76. Learning Goal-Oriented Language-Guided Navigation with Self-Improving Demonstrations at Scale 2 upvotes, #70 of 2025-09-30
  77. Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning 2 upvotes, #70 of 2025-09-30
  78. Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models 1 upvotes, #78 of 2025-09-30
  79. BOE-XSUM: Extreme Summarization in Clear Language of Spanish Legal Decrees and Notifications 1 upvotes, #78 of 2025-09-30
  80. Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns 1 upvotes, #78 of 2025-09-30
  81. TR2-D2: Tree Search Guided Trajectory-Aware Fine-Tuning for Discrete Diffusion 1 upvotes, #78 of 2025-09-30
  82. Advancing Reference-free Evaluation of Video Captions with Factual Analysis 2 upvotes, #82 of 2025-09-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.