Daily Papers of 2026-07-21

  1. RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model 197 upvotes, #1 of 2026-07-21
  2. TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs 165 upvotes, #2 of 2026-07-21
  3. EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World 92 upvotes, #3 of 2026-07-21
  4. DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment 91 upvotes, #4 of 2026-07-21
  5. SWE-Pruner Pro: The Coder LLM Already Knows What to Prune 77 upvotes, #5 of 2026-07-21
  6. Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning 68 upvotes, #6 of 2026-07-21
  7. HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement 59 upvotes, #7 of 2026-07-21
  8. FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry 50 upvotes, #8 of 2026-07-21
  9. Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence 43 upvotes, #9 of 2026-07-21
  10. GigaChat Audio: Time-aware Large Audio Language Model 37 upvotes, #10 of 2026-07-21
  11. GigaAM Multilingual: Foundation Model for Underrepresented Languages 33 upvotes, #11 of 2026-07-21
  12. Group Entropy-Controlled Policy Optimization 29 upvotes, #12 of 2026-07-21
  13. ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams 28 upvotes, #13 of 2026-07-21
  14. SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction 26 upvotes, #14 of 2026-07-21
  15. DiffGI: Differentiable Geometry Images for High-Fidelity Thin-Shell 3D Generation 20 upvotes, #15 of 2026-07-21
  16. Environment-free Synthetic Data Generation for API-Calling Agents 20 upvotes, #15 of 2026-07-21
  17. NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs 15 upvotes, #17 of 2026-07-21
  18. Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints 15 upvotes, #17 of 2026-07-21
  19. LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks 14 upvotes, #19 of 2026-07-21
  20. HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis 11 upvotes, #20 of 2026-07-21
  21. Distilled Reinforcement Learning for LLM Post-training 9 upvotes, #21 of 2026-07-21
  22. JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models 8 upvotes, #22 of 2026-07-21
  23. Can Multimodal Large Language Models Understand OCT? 8 upvotes, #22 of 2026-07-21
  24. Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL 7 upvotes, #24 of 2026-07-21
  25. DiFA: Inference-Time Forward-Process Alignment for Diffusion Models 7 upvotes, #24 of 2026-07-21
  26. Nonuniformity Principle in Human-AI Coworking 6 upvotes, #26 of 2026-07-21
  27. Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift 6 upvotes, #26 of 2026-07-21
  28. ShotPlan: Cinematic Video Generation with Learnable Planning Token 6 upvotes, #26 of 2026-07-21
  29. FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications 6 upvotes, #26 of 2026-07-21
  30. The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture 5 upvotes, #30 of 2026-07-21
  31. Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go? 5 upvotes, #30 of 2026-07-21
  32. WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting 5 upvotes, #30 of 2026-07-21
  33. Diagnosing and Calibrating Tool-Call Boundary Drift in Multi-Teacher On-Policy Distillation 4 upvotes, #33 of 2026-07-21
  34. OpenLongTail: Generative Scaling of Long-Tail Driving Data 4 upvotes, #33 of 2026-07-21
  35. Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation 4 upvotes, #33 of 2026-07-21
  36. ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video 4 upvotes, #33 of 2026-07-21
  37. UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation 3 upvotes, #37 of 2026-07-21

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.