Daily Papers of 2025-12-30

  1. Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss 93 upvotes, #1 of 2025-12-30
  2. LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation 64 upvotes, #2 of 2025-12-30
  3. Yume-1.5: A Text-Controlled Interactive World Generation Model 57 upvotes, #3 of 2025-12-30
  4. Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion 46 upvotes, #4 of 2025-12-30
  5. Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation 44 upvotes, #5 of 2025-12-30
  6. Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone 43 upvotes, #6 of 2025-12-30
  7. SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents 38 upvotes, #7 of 2025-12-30
  8. SpotEdit: Selective Region Editing in Diffusion Transformers 37 upvotes, #8 of 2025-12-30
  9. GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models 24 upvotes, #9 of 2025-12-30
  10. Web World Models 22 upvotes, #10 of 2025-12-30
  11. Act2Goal: From World Model To General Goal-conditioned Policy 21 upvotes, #11 of 2025-12-30
  12. DiRL: An Efficient Post-Training Framework for Diffusion Language Models 19 upvotes, #12 of 2025-12-30
  13. Nested Browser-Use Learning for Agentic Information Seeking 17 upvotes, #13 of 2025-12-30
  14. Training AI Co-Scientists Using Rubric Rewards 17 upvotes, #13 of 2025-12-30
  15. Self-Evaluation Unlocks Any-Step Text-to-Image Generation 15 upvotes, #15 of 2025-12-30
  16. OmniAgent: Audio-Guided Active Perception Agent for Omnimodal Audio-Video Understanding 14 upvotes, #16 of 2025-12-30
  17. YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection 13 upvotes, #17 of 2025-12-30
  18. Video-BrowseComp: Benchmarking Agentic Video Research on Open Web 9 upvotes, #18 of 2025-12-30
  19. SurgWorld: Learning Surgical Robot Policies from Videos via World Modeling 9 upvotes, #18 of 2025-12-30
  20. VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs 8 upvotes, #20 of 2025-12-30
  21. Monadic Context Engineering 8 upvotes, #20 of 2025-12-30
  22. An Information Theoretic Perspective on Agentic System Design 7 upvotes, #22 of 2025-12-30
  23. Quantile Rendering: Efficiently Embedding High-dimensional Feature on 3D Gaussian Splatting 6 upvotes, #23 of 2025-12-30
  24. Bridging Your Imagination with Audio-Video Generation via a Unified Director 5 upvotes, #24 of 2025-12-30
  25. ProGuard: Towards Proactive Multimodal Safeguard 5 upvotes, #24 of 2025-12-30
  26. Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation 5 upvotes, #24 of 2025-12-30
  27. Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation 4 upvotes, #27 of 2025-12-30
  28. KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta 3 upvotes, #28 of 2025-12-30
  29. Introducing TrGLUE and SentiTurca: A Comprehensive Benchmark for Turkish General Language Understanding and Sentiment Analysis 2 upvotes, #29 of 2025-12-30
  30. Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks 2 upvotes, #29 of 2025-12-30
  31. Reverse Personalization 1 upvotes, #31 of 2025-12-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.