Daily Papers of 2026-07-08

  1. RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation 92 upvotes, #1 of 2026-07-08
  2. AlayaWorld: Long-Horizon and Playable Video World Generation 87 upvotes, #2 of 2026-07-08
  3. Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling 76 upvotes, #3 of 2026-07-08
  4. RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation 76 upvotes, #3 of 2026-07-08
  5. Gemma 4 Technical Report 63 upvotes, #5 of 2026-07-08
  6. Vision as Unified Multimodal Generation 46 upvotes, #6 of 2026-07-08
  7. DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation 33 upvotes, #7 of 2026-07-08
  8. LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL 32 upvotes, #8 of 2026-07-08
  9. SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe 30 upvotes, #9 of 2026-07-08
  10. Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning 27 upvotes, #10 of 2026-07-08
  11. Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory 22 upvotes, #11 of 2026-07-08
  12. From Foundation to Application: Improving VLA Models in Practice 19 upvotes, #12 of 2026-07-08
  13. TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training 18 upvotes, #13 of 2026-07-08
  14. PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation 16 upvotes, #14 of 2026-07-08
  15. JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications 15 upvotes, #15 of 2026-07-08
  16. Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion Model 14 upvotes, #16 of 2026-07-08
  17. MentalThink: Shaping Thoughts in Mental SVG World 14 upvotes, #16 of 2026-07-08
  18. Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding 12 upvotes, #18 of 2026-07-08
  19. CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration 11 upvotes, #19 of 2026-07-08
  20. Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models 10 upvotes, #20 of 2026-07-08
  21. Attending to Multimodal Generation One Token at a Time 9 upvotes, #21 of 2026-07-08
  22. CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation 9 upvotes, #21 of 2026-07-08
  23. TREK: Distill to Explore, Reinforce to Refine 8 upvotes, #23 of 2026-07-08
  24. SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review 8 upvotes, #23 of 2026-07-08
  25. 3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance 7 upvotes, #25 of 2026-07-08
  26. Rank-Then-Act: Reward-Free Control from Frame-Order Progress 7 upvotes, #25 of 2026-07-08
  27. HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better 7 upvotes, #25 of 2026-07-08
  28. MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs 6 upvotes, #28 of 2026-07-08
  29. When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffers 6 upvotes, #28 of 2026-07-08
  30. Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training 6 upvotes, #28 of 2026-07-08
  31. PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages 6 upvotes, #28 of 2026-07-08
  32. Bibby AI: An Editor-Native Agentic Platform for Academic Research, Writing, and Publishing 5 upvotes, #32 of 2026-07-08
  33. SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models 5 upvotes, #32 of 2026-07-08
  34. Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES 4 upvotes, #34 of 2026-07-08
  35. Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers 3 upvotes, #35 of 2026-07-08
  36. RuleChef: Grounding LLM Task Knowledge in Human-Editable Rules 3 upvotes, #35 of 2026-07-08
  37. Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment 3 upvotes, #35 of 2026-07-08
  38. SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control 3 upvotes, #35 of 2026-07-08
  39. Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator 3 upvotes, #35 of 2026-07-08
  40. VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech 1 upvotes, #40 of 2026-07-08
  41. SiamJEPA: On the Role of Siamese Student Encoders in JEPA 1 upvotes, #40 of 2026-07-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.