Daily Papers of 2026-07-08
- RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation 92 upvotes, #1 of 2026-07-08
- AlayaWorld: Long-Horizon and Playable Video World Generation 87 upvotes, #2 of 2026-07-08
- Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling 76 upvotes, #3 of 2026-07-08
- RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation 76 upvotes, #3 of 2026-07-08
- Gemma 4 Technical Report 63 upvotes, #5 of 2026-07-08
- Vision as Unified Multimodal Generation 46 upvotes, #6 of 2026-07-08
- DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation 33 upvotes, #7 of 2026-07-08
- LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL 32 upvotes, #8 of 2026-07-08
- SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe 30 upvotes, #9 of 2026-07-08
- Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning 27 upvotes, #10 of 2026-07-08
- Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory 22 upvotes, #11 of 2026-07-08
- From Foundation to Application: Improving VLA Models in Practice 19 upvotes, #12 of 2026-07-08
- TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training 18 upvotes, #13 of 2026-07-08
- PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation 16 upvotes, #14 of 2026-07-08
- JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications 15 upvotes, #15 of 2026-07-08
- Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model 14 upvotes, #16 of 2026-07-08
- MentalThink: Shaping Thoughts in Mental SVG World 14 upvotes, #16 of 2026-07-08
- Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding 12 upvotes, #18 of 2026-07-08
- CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration 11 upvotes, #19 of 2026-07-08
- Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models 10 upvotes, #20 of 2026-07-08
- Attending to Multimodal Generation One Token at a Time 9 upvotes, #21 of 2026-07-08
- CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation 9 upvotes, #21 of 2026-07-08
- TREK: Distill to Explore, Reinforce to Refine 8 upvotes, #23 of 2026-07-08
- SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review 8 upvotes, #23 of 2026-07-08
- 3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance 7 upvotes, #25 of 2026-07-08
- Rank-Then-Act: Reward-Free Control from Frame-Order Progress 7 upvotes, #25 of 2026-07-08
- HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better 7 upvotes, #25 of 2026-07-08
- MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs 6 upvotes, #28 of 2026-07-08
- When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers 6 upvotes, #28 of 2026-07-08
- Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training 6 upvotes, #28 of 2026-07-08
- PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages 6 upvotes, #28 of 2026-07-08
- Bibby AI: An Editor-Native Agentic Platform for Academic Research, Writing, and Publishing 5 upvotes, #32 of 2026-07-08
- SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models 5 upvotes, #32 of 2026-07-08
- Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES 4 upvotes, #34 of 2026-07-08
- Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers 3 upvotes, #35 of 2026-07-08
- RuleChef: Grounding LLM Task Knowledge in Human-Editable Rules 3 upvotes, #35 of 2026-07-08
- Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment 3 upvotes, #35 of 2026-07-08
- SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control 3 upvotes, #35 of 2026-07-08
- Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator 3 upvotes, #35 of 2026-07-08
- VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech 1 upvotes, #40 of 2026-07-08
- SiamJEPA: On the Role of Siamese Student Encoders in JEPA 1 upvotes, #40 of 2026-07-08
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.