Daily Papers of 2026-08-27

  1. VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 266 upvotes, #1 of 2026-08-27
  2. VGI-BENCH: Probing Visual Intelligence in Video Generation Models 177 upvotes, #2 of 2026-08-27
  3. VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction 171 upvotes, #3 of 2026-08-27
  4. FrontierChallenge: Evaluating Scientific Workflow Completion 143 upvotes, #4 of 2026-08-27
  5. WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation 141 upvotes, #5 of 2026-08-27
  6. JIT-Agent: Scaling Harness Intelligence via Just-in-Time Harness Evolution 112 upvotes, #6 of 2026-08-27
  7. Code World Model: Coding Agent as World Brain 34 upvotes, #7 of 2026-08-27
  8. Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning 27 upvotes, #8 of 2026-08-27
  9. D^3-MOPD: Adaptive Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation 25 upvotes, #9 of 2026-08-27
  10. Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation 21 upvotes, #10 of 2026-08-27
  11. Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data 21 upvotes, #10 of 2026-08-27
  12. Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds 19 upvotes, #12 of 2026-08-27
  13. StreamPI: Streaming Multimodal Temporal Modeling for Vision-Language-Action Models 19 upvotes, #12 of 2026-08-27
  14. Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios 16 upvotes, #14 of 2026-08-27
  15. V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning 15 upvotes, #15 of 2026-08-27
  16. SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? 14 upvotes, #16 of 2026-08-27
  17. The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents 13 upvotes, #17 of 2026-08-27
  18. A Programming Paradigm for Spatiotemporal Composability 13 upvotes, #17 of 2026-08-27
  19. Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments 12 upvotes, #19 of 2026-08-27
  20. Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation 12 upvotes, #19 of 2026-08-27
  21. Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation in Transformers 10 upvotes, #21 of 2026-08-27
  22. GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding 9 upvotes, #22 of 2026-08-27
  23. MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization 9 upvotes, #22 of 2026-08-27
  24. Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models 5 upvotes, #24 of 2026-08-27
  25. Super Star: Towards Streaming Real-time Interactive Agents for Digital Humans 5 upvotes, #24 of 2026-08-27
  26. Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling 4 upvotes, #26 of 2026-08-27
  27. FIRM-Video: Check Before You Score for Reliable Text-to-Video Reward Modeling 4 upvotes, #26 of 2026-08-27
  28. Skill Issue: Are Skills Language-Invariant in LLMs? 4 upvotes, #26 of 2026-08-27
  29. Prefix Sliding for efficient test-time scaling 4 upvotes, #26 of 2026-08-27
  30. LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale 3 upvotes, #30 of 2026-08-27
  31. RetrievalRouter: Joint Modality and Architecture Selection for Document Retrieval 3 upvotes, #30 of 2026-08-27
  32. A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans 2 upvotes, #32 of 2026-08-27
  33. Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction 0 upvotes, #33 of 2026-08-27

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.