Daily Papers of 2026-09-15
- Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation 695 upvotes, #1 of 2026-09-15
- Atria Dawn: The Dawn of Agentic Superintelligence 421 upvotes, #2 of 2026-09-15
- ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search 255 upvotes, #3 of 2026-09-15
- Dream-RSI: Recursive Self-Improvement through Evolving Worlds 242 upvotes, #4 of 2026-09-15
- PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models 171 upvotes, #5 of 2026-09-15
- RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments 81 upvotes, #6 of 2026-09-15
- Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction 79 upvotes, #7 of 2026-09-15
- BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender 34 upvotes, #8 of 2026-09-15
- LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows 34 upvotes, #8 of 2026-09-15
- Discovery Foundation Models: Toward Open-Ended Discovery Intelligence 33 upvotes, #10 of 2026-09-15
- AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video 26 upvotes, #11 of 2026-09-15
- How Lossless Is Lossless Speculative Decoding? The Role of Numerical Precision in Orthrus 23 upvotes, #12 of 2026-09-15
- Omni-Streaming Thinking 22 upvotes, #13 of 2026-09-15
- HazardAuditor: From Executable Threats to Safer Computer-Use Agents 18 upvotes, #14 of 2026-09-15
- Kaininja: Extending Native 3D Generators to the Part Level 18 upvotes, #14 of 2026-09-15
- Agent as Policy for Robotic Manipulation 17 upvotes, #16 of 2026-09-15
- LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents 16 upvotes, #17 of 2026-09-15
- MInTRL: Off-policy Intervention can boost On-policy RL 14 upvotes, #18 of 2026-09-15
- Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training 14 upvotes, #18 of 2026-09-15
- Learning to Solve Hard Problems in RL for LLMs by Never Giving Up 13 upvotes, #20 of 2026-09-15
- When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis 10 upvotes, #21 of 2026-09-15
- Building a Production Greek-English Speech Recognizer 9 upvotes, #22 of 2026-09-15
- Enabling Creative Exploration for Vibe Design Agents 9 upvotes, #22 of 2026-09-15
- Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model 8 upvotes, #24 of 2026-09-15
- Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks 8 upvotes, #24 of 2026-09-15
- ModaLens: Measuring Image Sensitivity in Report-Conditioned Medical VLMs 8 upvotes, #24 of 2026-09-15
- Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning 7 upvotes, #27 of 2026-09-15
- Expert-Space Exploration in MoE Reinforcement Learning 5 upvotes, #28 of 2026-09-15
- Root-Cause Attribution Is a Search Problem: Continual Search for Long-Horizon Agent Failures 5 upvotes, #28 of 2026-09-15
- Realtime-Venus: A full-duplex interaction system with asynchronous delegation 3 upvotes, #30 of 2026-09-15
- Thought without systematicity? Evaluating reasoning models on rule induction tasks 3 upvotes, #30 of 2026-09-15
- E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning 3 upvotes, #30 of 2026-09-15
- Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition 3 upvotes, #30 of 2026-09-15
- Learning Sparse Decision Trees via Transformer Variational Auto-Encoders 2 upvotes, #34 of 2026-09-15
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.