Daily Papers of 2026-05-07

  1. Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation 122 upvotes, #1 of 2026-05-07
  2. RLDX-1 Technical Report 118 upvotes, #2 of 2026-05-07
  3. Stream-T1: Test-Time Scaling for Streaming Video Generation 102 upvotes, #3 of 2026-05-07
  4. OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents 96 upvotes, #4 of 2026-05-07
  5. HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation 71 upvotes, #5 of 2026-05-07
  6. MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction 68 upvotes, #6 of 2026-05-07
  7. Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems 38 upvotes, #7 of 2026-05-07
  8. PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World 35 upvotes, #8 of 2026-05-07
  9. D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models 25 upvotes, #9 of 2026-05-07
  10. CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing 21 upvotes, #10 of 2026-05-07
  11. Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation 17 upvotes, #11 of 2026-05-07
  12. Lightning Unified Video Editing via In-Context Sparse Attention 17 upvotes, #11 of 2026-05-07
  13. XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity 11 upvotes, #13 of 2026-05-07
  14. ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning 7 upvotes, #14 of 2026-05-07
  15. Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback 6 upvotes, #15 of 2026-05-07
  16. APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music 5 upvotes, #16 of 2026-05-07
  17. Diffusion Model as a Generalist Segmentation Learner 4 upvotes, #17 of 2026-05-07
  18. TT4D: A Pipeline and Dataset for Table Tennis 4D Reconstruction From Monocular Videos 4 upvotes, #17 of 2026-05-07
  19. A Foundation Model for Zero-Shot Logical Rule Induction 4 upvotes, #17 of 2026-05-07
  20. MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills 3 upvotes, #20 of 2026-05-07
  21. SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies 3 upvotes, #20 of 2026-05-07
  22. Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Environments 2 upvotes, #22 of 2026-05-07
  23. KinDER: A Physical Reasoning Benchmark for Robot Learning and Planning 2 upvotes, #22 of 2026-05-07
  24. When to Think, When to Speak: Learning Disclosure Policies for LLM Reasoning 2 upvotes, #22 of 2026-05-07
  25. The First Token Knows: Single-Decode Confidence for Hallucination Detection 2 upvotes, #22 of 2026-05-07

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.