Daily Papers of 2025-10-01
- The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain 441 upvotes, #1 of 2025-10-01
- MCPMark: A Benchmark for Stress-Testing Realistic and Comprehensive MCP Use 156 upvotes, #2 of 2025-10-01
- Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play 123 upvotes, #3 of 2025-10-01
- Winning the Pruning Gamble: A Unified Approach to Joint Sample and Token Pruning for Efficient Supervised Fine-Tuning 65 upvotes, #4 of 2025-10-01
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models 56 upvotes, #5 of 2025-10-01
- TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning 47 upvotes, #6 of 2025-10-01
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training 41 upvotes, #7 of 2025-10-01
- DC-VideoGen: Efficient Video Generation with Deep Compression Video Autoencoder 33 upvotes, #8 of 2025-10-01
- OceanGym: A Benchmark Environment for Underwater Embodied Agents 33 upvotes, #8 of 2025-10-01
- Thinking-Free Policy Initialization Makes Distilled Reasoning Models More Effective and Efficient Reasoners 29 upvotes, #10 of 2025-10-01
- Who's Your Judge? On the Detectability of LLM-Generated Judgments 27 upvotes, #11 of 2025-10-01
- Rethinking Reward Models for Multi-Domain Test-Time Scaling 24 upvotes, #12 of 2025-10-01
- Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training 21 upvotes, #13 of 2025-10-01
- DA^2: Depth Anything in Any Direction 20 upvotes, #14 of 2025-10-01
- Muon Outperforms Adam in Tail-End Associative Memory Learning 18 upvotes, #15 of 2025-10-01
- dParallel: Learnable Parallel Decoding for dLLMs 18 upvotes, #15 of 2025-10-01
- IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance 17 upvotes, #17 of 2025-10-01
- MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation 17 upvotes, #17 of 2025-10-01
- VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications 17 upvotes, #17 of 2025-10-01
- DeepScientist: Advancing Frontier-Pushing Scientific Findings Progressively 16 upvotes, #20 of 2025-10-01
- Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs 15 upvotes, #21 of 2025-10-01
- Regression Language Models for Code 15 upvotes, #21 of 2025-10-01
- Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention 13 upvotes, #23 of 2025-10-01
- Mem-α: Learning Memory Construction via Reinforcement Learning 12 upvotes, #24 of 2025-10-01
- Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models 12 upvotes, #24 of 2025-10-01
- TTT3R: 3D Reconstruction as Test-Time Training 11 upvotes, #26 of 2025-10-01
- InfoAgent: Advancing Autonomous Information-Seeking Agents 10 upvotes, #27 of 2025-10-01
- OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always! 10 upvotes, #27 of 2025-10-01
- Humanline: Online Alignment as Perceptual Loss 9 upvotes, #29 of 2025-10-01
- VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes 9 upvotes, #29 of 2025-10-01
- A Cartography of Open Collaboration in Open Source AI: Mapping Practices, Motivations, and Governance in 14 Open Large Language Model Projects 9 upvotes, #29 of 2025-10-01
- Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap 8 upvotes, #32 of 2025-10-01
- Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspective 7 upvotes, #33 of 2025-10-01
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents 7 upvotes, #33 of 2025-10-01
- Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMs 5 upvotes, #35 of 2025-10-01
- jina-reranker-v3: Last but Not Late Interaction for Document Reranking 5 upvotes, #35 of 2025-10-01
- Learning to Reason as Action Abstractions with Scalable Mid-Training RL 5 upvotes, #35 of 2025-10-01
- Context Is What You Need: The Maximum Effective Context Window for Real World Limits of LLMs 4 upvotes, #38 of 2025-10-01
- The Pitfalls of KV Cache Compression 4 upvotes, #38 of 2025-10-01
- Nudging the Boundaries of LLM Reasoning 3 upvotes, #40 of 2025-10-01
- DeepCodeSeek: Real-Time API Retrieval for Context-Aware Code Generation 3 upvotes, #40 of 2025-10-01
- Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception 2 upvotes, #42 of 2025-10-01
- d^2Cache: Accelerating Diffusion-Based LLMs via Dual Adaptive Caching 2 upvotes, #42 of 2025-10-01
- Knowledge Homophily in Large Language Models 2 upvotes, #42 of 2025-10-01
- Specialization after Generalization: Towards Understanding Test-Time Training in Foundation Models 2 upvotes, #42 of 2025-10-01
- Who invented deep residual learning? 2 upvotes, #42 of 2025-10-01
- BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software 2 upvotes, #42 of 2025-10-01
- EntroPE: Entropy-Guided Dynamic Patch Encoder for Time Series Forecasting 2 upvotes, #42 of 2025-10-01
- TAU: A Benchmark for Cultural Sound Understanding Beyond Semantics 2 upvotes, #42 of 2025-10-01
- Convolutional Set Transformer 1 upvotes, #50 of 2025-10-01
- Estimating Time Series Foundation Model Transferability via In-Context Learning 1 upvotes, #50 of 2025-10-01
- CORRECT: COndensed eRror RECognition via knowledge Transfer in multi-agent systems 1 upvotes, #50 of 2025-10-01
- MANI-Pure: Magnitude-Adaptive Noise Injection for Adversarial Purification 1 upvotes, #50 of 2025-10-01
- LayerD: Decomposing Raster Graphic Designs into Layers 1 upvotes, #50 of 2025-10-01
- Swift: An Autoregressive Consistency Model for Efficient Weather Forecasting 1 upvotes, #50 of 2025-10-01
- Probing the Critical Point (CritPt) of AI Reasoning: a Frontier Physics Research Benchmark 1 upvotes, #50 of 2025-10-01
- Video Object Segmentation-Aware Audio Generation 1 upvotes, #50 of 2025-10-01
- GeoRemover: Removing Objects and Their Causal Visual Artifacts 1 upvotes, #58 of 2025-10-01
- LLM Watermark Evasion via Bias Inversion 2 upvotes, #58 of 2025-10-01
- ProfVLM: A Lightweight Video-Language Model for Multi-View Proficiency Estimation 1 upvotes, #58 of 2025-10-01
- Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation 1 upvotes, #58 of 2025-10-01
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.