Daily Papers of 2025-09-30
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse-Linear Attention 109 upvotes, #1 of 2025-09-30
- StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs 62 upvotes, #2 of 2025-09-30
- Multiplayer Nash Preference Optimization 60 upvotes, #3 of 2025-09-30
- OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing 49 upvotes, #4 of 2025-09-30
- Beyond the Exploration-Exploitation Trade-off: A Hidden State Approach for LLM Reasoning in RLVR 47 upvotes, #5 of 2025-09-30
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark 44 upvotes, #6 of 2025-09-30
- Democratizing AI scientists using ToolUniverse 38 upvotes, #7 of 2025-09-30
- SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer 38 upvotes, #7 of 2025-09-30
- When Does Reasoning Matter? A Controlled Study of Reasoning's Contribution to Model Performance 37 upvotes, #9 of 2025-09-30
- Sequential Diffusion Language Models 36 upvotes, #10 of 2025-09-30
- Visual Jigsaw Post-Training Improves MLLMs 34 upvotes, #11 of 2025-09-30
- SparseD: Sparse Attention for Diffusion Language Models 29 upvotes, #12 of 2025-09-30
- Random Policy Valuation is Enough for LLM Reasoning with Verifiable Rewards 27 upvotes, #13 of 2025-09-30
- Towards Personalized Deep Research: Benchmarks and Evaluations 27 upvotes, #13 of 2025-09-30
- GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts 27 upvotes, #13 of 2025-09-30
- EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling 26 upvotes, #16 of 2025-09-30
- EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering 26 upvotes, #16 of 2025-09-30
- VideoScore2: Think before You Score in Generative Video Evaluation 22 upvotes, #18 of 2025-09-30
- Rolling Forcing: Autoregressive Long Video Diffusion in Real Time 21 upvotes, #19 of 2025-09-30
- Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learning 20 upvotes, #20 of 2025-09-30
- HunyuanImage 3.0 Technical Report 20 upvotes, #20 of 2025-09-30
- WirelessMathLM: Teaching Mathematical Reasoning for LLMs in Wireless Communications with Reinforcement Learning 18 upvotes, #22 of 2025-09-30
- From f(x) and g(x) to f(g(x)): LLMs Learn New Skills in RL by Composing Old Ones 17 upvotes, #23 of 2025-09-30
- VGGT-X: When VGGT Meets Dense Novel View Synthesis 17 upvotes, #23 of 2025-09-30
- Socratic-Zero : Bootstrapping Reasoning via Data-Free Agent Co-evolution 16 upvotes, #25 of 2025-09-30
- Scaling Generalist Data-Analytic Agents 16 upvotes, #25 of 2025-09-30
- The Era of Real-World Human Interaction: RL from User Conversations 16 upvotes, #25 of 2025-09-30
- Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks 15 upvotes, #28 of 2025-09-30
- MMPB: It's Time for Multi-Modal Personalization 14 upvotes, #29 of 2025-09-30
- Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding 13 upvotes, #30 of 2025-09-30
- Toward Effective Tool-Integrated Reasoning via Self-Evolved Preference Learning 13 upvotes, #30 of 2025-09-30
- BRIDGE - Building Reinforcement-Learning Depth-to-Image Data Generation Engine for Monocular Depth Estimation 13 upvotes, #30 of 2025-09-30
- MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech 13 upvotes, #30 of 2025-09-30
- InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation 12 upvotes, #34 of 2025-09-30
- SIRI: Scaling Iterative Reinforcement Learning with Interleaved Compression 12 upvotes, #34 of 2025-09-30
- Dynamic Experts Search: Enhancing Reasoning in Mixture-of-Experts LLMs at Test Time 11 upvotes, #36 of 2025-09-30
- Rethinking Large Language Model Distillation: A Constrained Markov Decision Process Perspective 11 upvotes, #36 of 2025-09-30
- Where LLM Agents Fail and How They can Learn From Failures 11 upvotes, #36 of 2025-09-30
- Pretraining Large Language Models with NVFP4 10 upvotes, #39 of 2025-09-30
- LUMINA: Detecting Hallucinations in RAG System with Context-Knowledge Signals 9 upvotes, #40 of 2025-09-30
- From Harm to Help: Turning Reasoning In-Context Demos into Assets for Reasoning LMs 9 upvotes, #40 of 2025-09-30
- Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation 9 upvotes, #40 of 2025-09-30
- GRPO-MA: Multi-Answer Generation in GRPO for Stable and Efficient Chain-of-Thought Training 8 upvotes, #43 of 2025-09-30
- DepthLM: Metric Depth From Vision Language Models 8 upvotes, #43 of 2025-09-30
- Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment 7 upvotes, #45 of 2025-09-30
- Taming Masked Diffusion Language Models via Consistency Trajectory Reinforcement Learning with Fewer Decoding Step 7 upvotes, #45 of 2025-09-30
- SCI-Verifier: Scientific Verifier with Thinking 7 upvotes, #45 of 2025-09-30
- MultiCrafter: High-Fidelity Multi-Subject Generation via Spatially Disentangled Attention and Identity-Aware Reinforcement Learning 6 upvotes, #48 of 2025-09-30
- Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification 6 upvotes, #48 of 2025-09-30
- AceSearcher: Bootstrapping Reasoning and Search for LLMs via Reinforced Self-Play 6 upvotes, #48 of 2025-09-30
- Rethinking JEPA: Compute-Efficient Video SSL with Frozen Teachers 6 upvotes, #48 of 2025-09-30
- Hyperspherical Latents Improve Continuous-Token Autoregressive Generation 6 upvotes, #48 of 2025-09-30
- DC-Gen: Post-Training Diffusion Acceleration with Deeply Compressed Latent Space 6 upvotes, #48 of 2025-09-30
- Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization 5 upvotes, #54 of 2025-09-30
- LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning 5 upvotes, #54 of 2025-09-30
- The Photographer Eye: Teaching Multimodal Large Language Models to See and Critique like Photographers 4 upvotes, #56 of 2025-09-30
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents 4 upvotes, #56 of 2025-09-30
- MathBode: Frequency-Domain Fingerprints of LLM Mathematical Reasoning 4 upvotes, #56 of 2025-09-30
- PARROT: A Benchmark for Evaluating LLMs in Cross-System SQL Translation 4 upvotes, #56 of 2025-09-30
- Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning 4 upvotes, #56 of 2025-09-30
- Scalable GANs with Transformers 4 upvotes, #56 of 2025-09-30
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models 4 upvotes, #56 of 2025-09-30
- PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images 4 upvotes, #56 of 2025-09-30
- UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration 3 upvotes, #64 of 2025-09-30
- TENET: Leveraging Tests Beyond Validation for Code Generation 3 upvotes, #64 of 2025-09-30
- UniVid: The Open-Source Unified Video Model 3 upvotes, #64 of 2025-09-30
- AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models 3 upvotes, #64 of 2025-09-30
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video? 3 upvotes, #64 of 2025-09-30
- ThermalGen: Style-Disentangled Flow-Based Generative Models for RGB-to-Thermal Image Translation 3 upvotes, #64 of 2025-09-30
- Charting a Decade of Computational Linguistics in Italy: The CLiC-it Corpus 2 upvotes, #70 of 2025-09-30
- Combinatorial Creativity: A New Frontier in Generalization Abilities 2 upvotes, #70 of 2025-09-30
- REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model 2 upvotes, #70 of 2025-09-30
- ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning 2 upvotes, #70 of 2025-09-30
- RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human Mobility 2 upvotes, #70 of 2025-09-30
- BPMN Assistant: An LLM-Based Approach to Business Process Modeling 2 upvotes, #70 of 2025-09-30
- Learning Goal-Oriented Language-Guided Navigation with Self-Improving Demonstrations at Scale 2 upvotes, #70 of 2025-09-30
- Cogito, Ergo Ludo: An Agent that Learns to Play by Reasoning and Planning 2 upvotes, #70 of 2025-09-30
- Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models 1 upvotes, #78 of 2025-09-30
- BOE-XSUM: Extreme Summarization in Clear Language of Spanish Legal Decrees and Notifications 1 upvotes, #78 of 2025-09-30
- Generalized Correctness Models: Learning Calibrated and Model-Agnostic Correctness Predictors from Historical Patterns 1 upvotes, #78 of 2025-09-30
- TR2-D2: Tree Search Guided Trajectory-Aware Fine-Tuning for Discrete Diffusion 1 upvotes, #78 of 2025-09-30
- Advancing Reference-free Evaluation of Video Captions with Factual Analysis 2 upvotes, #82 of 2025-09-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.