Daily Papers of 2026-07-17
- LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget 197 upvotes, #1 of 2026-07-17
- VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding 168 upvotes, #2 of 2026-07-17
- SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning 99 upvotes, #3 of 2026-07-17
- SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration 69 upvotes, #4 of 2026-07-17
- BadWAM: When World-Action Models Dream Right but Act Wrong 53 upvotes, #5 of 2026-07-17
- KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation 42 upvotes, #6 of 2026-07-17
- From Pixels to States: Rethinking Interactive World Models as Game Engines 35 upvotes, #7 of 2026-07-17
- Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes 34 upvotes, #8 of 2026-07-17
- MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation 34 upvotes, #8 of 2026-07-17
- UniVR: Thinking in Visual Space for Unified Visual Reasoning 32 upvotes, #10 of 2026-07-17
- RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination 31 upvotes, #11 of 2026-07-17
- Spectral Rewiring for Exploration, Purification, and Model Merging 25 upvotes, #12 of 2026-07-17
- RoboTTT: Context Scaling for Robot Policies 22 upvotes, #13 of 2026-07-17
- Video = World + Event Stream 21 upvotes, #14 of 2026-07-17
- Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations 20 upvotes, #15 of 2026-07-17
- WanSong v1.0 Technical Report 17 upvotes, #16 of 2026-07-17
- MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators 17 upvotes, #16 of 2026-07-17
- DeepLoop: Depth Scaling for Looped Transformers 16 upvotes, #18 of 2026-07-17
- Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel 11 upvotes, #19 of 2026-07-17
- AsySplat: Efficient Asymmetric 3D Gaussian Splatting for Long-Sequence Scene Modeling 9 upvotes, #20 of 2026-07-17
- Rethinking the Evaluation of Harness Evolution for Agents 9 upvotes, #20 of 2026-07-17
- Token Time Continuous Diffusion for Language Modeling 9 upvotes, #20 of 2026-07-17
- VIABench: A Comprehensive Video Benchmark Collected from Blind Individuals for Visual Impairment Assistance 9 upvotes, #20 of 2026-07-17
- Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models 9 upvotes, #20 of 2026-07-17
- GRASP: GRanularity-Aware Search Policy for Agentic RAG 8 upvotes, #25 of 2026-07-17
- SUFLECA: Scaling Up Feature Learning for CAD-to-image Alignment 8 upvotes, #25 of 2026-07-17
- Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving 7 upvotes, #27 of 2026-07-17
- Hierarchical Denoising For Multi-Step Visual Reasoning 5 upvotes, #28 of 2026-07-17
- On Locality and Length Generalization in Visual Reasoning 4 upvotes, #29 of 2026-07-17
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.