Daily Papers of 2025-12-30
- Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss 93 upvotes, #1 of 2025-12-30
- LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation 64 upvotes, #2 of 2025-12-30
- Yume-1.5: A Text-Controlled Interactive World Generation Model 57 upvotes, #3 of 2025-12-30
- Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion 46 upvotes, #4 of 2025-12-30
- Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation 44 upvotes, #5 of 2025-12-30
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone 43 upvotes, #6 of 2025-12-30
- SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents 38 upvotes, #7 of 2025-12-30
- SpotEdit: Selective Region Editing in Diffusion Transformers 37 upvotes, #8 of 2025-12-30
- GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models 24 upvotes, #9 of 2025-12-30
- Web World Models 22 upvotes, #10 of 2025-12-30
- Act2Goal: From World Model To General Goal-conditioned Policy 21 upvotes, #11 of 2025-12-30
- DiRL: An Efficient Post-Training Framework for Diffusion Language Models 19 upvotes, #12 of 2025-12-30
- Nested Browser-Use Learning for Agentic Information Seeking 17 upvotes, #13 of 2025-12-30
- Training AI Co-Scientists Using Rubric Rewards 17 upvotes, #13 of 2025-12-30
- Self-Evaluation Unlocks Any-Step Text-to-Image Generation 15 upvotes, #15 of 2025-12-30
- OmniAgent: Audio-Guided Active Perception Agent for Omnimodal Audio-Video Understanding 14 upvotes, #16 of 2025-12-30
- YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection 13 upvotes, #17 of 2025-12-30
- Video-BrowseComp: Benchmarking Agentic Video Research on Open Web 9 upvotes, #18 of 2025-12-30
- SurgWorld: Learning Surgical Robot Policies from Videos via World Modeling 9 upvotes, #18 of 2025-12-30
- VL-LN Bench: Towards Long-horizon Goal-oriented Navigation with Active Dialogs 8 upvotes, #20 of 2025-12-30
- Monadic Context Engineering 8 upvotes, #20 of 2025-12-30
- An Information Theoretic Perspective on Agentic System Design 7 upvotes, #22 of 2025-12-30
- Quantile Rendering: Efficiently Embedding High-dimensional Feature on 3D Gaussian Splatting 6 upvotes, #23 of 2025-12-30
- Bridging Your Imagination with Audio-Video Generation via a Unified Director 5 upvotes, #24 of 2025-12-30
- ProGuard: Towards Proactive Multimodal Safeguard 5 upvotes, #24 of 2025-12-30
- Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation 5 upvotes, #24 of 2025-12-30
- Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation 4 upvotes, #27 of 2025-12-30
- KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta 3 upvotes, #28 of 2025-12-30
- Introducing TrGLUE and SentiTurca: A Comprehensive Benchmark for Turkish General Language Understanding and Sentiment Analysis 2 upvotes, #29 of 2025-12-30
- Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks 2 upvotes, #29 of 2025-12-30
- Reverse Personalization 1 upvotes, #31 of 2025-12-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.