Daily Papers of 2025-12-10

  1. Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance 123 upvotes, #1 of 2025-12-10
  2. Visionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting Platform 75 upvotes, #2 of 2025-12-10
  3. Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality 46 upvotes, #3 of 2025-12-10
  4. OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory 43 upvotes, #4 of 2025-12-10
  5. DeepCode: Open Agentic Coding 29 upvotes, #5 of 2025-12-10
  6. From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs 23 upvotes, #6 of 2025-12-10
  7. Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation 21 upvotes, #7 of 2025-12-10
  8. ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models 19 upvotes, #8 of 2025-12-10
  9. Arbitrage: Efficient Reasoning via Advantage-Aware Speculation 15 upvotes, #9 of 2025-12-10
  10. MIND-V: Hierarchical Video Generation for Long-Horizon Robotic Manipulation with RL-based Physical Alignment 12 upvotes, #10 of 2025-12-10
  11. Boosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training 12 upvotes, #10 of 2025-12-10
  12. COREA: Coarse-to-Fine 3D Representation Alignment Between Relightable 3D Gaussians and SDF via Bidirectional 3D-to-3D Supervision 8 upvotes, #12 of 2025-12-10
  13. See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models 7 upvotes, #13 of 2025-12-10
  14. SUCCESS-GS: Survey of Compactness and Compression for Efficient Static and Dynamic Gaussian Splatting 6 upvotes, #14 of 2025-12-10
  15. TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models 6 upvotes, #14 of 2025-12-10
  16. Efficiently Reconstructing Dynamic Scenes One D4RT at a Time 6 upvotes, #14 of 2025-12-10
  17. Modular Neural Image Signal Processing 5 upvotes, #17 of 2025-12-10
  18. MemLoRA: Distilling Expert Adapters for On-Device Memory Systems 3 upvotes, #18 of 2025-12-10
  19. TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels 3 upvotes, #18 of 2025-12-10
  20. LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning 2 upvotes, #20 of 2025-12-10
  21. Novel Deep Learning Architectures for Classification and Segmentation of Brain Tumors from MRI Images 2 upvotes, #20 of 2025-12-10
  22. Terrain Diffusion: A Diffusion-Based Successor to Perlin Noise in Infinite, Real-Time Terrain Generation 2 upvotes, #20 of 2025-12-10
  23. SAM-Body4D: Training-Free 4D Human Body Mesh Recovery from Videos 2 upvotes, #20 of 2025-12-10
  24. EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce 2 upvotes, #20 of 2025-12-10
  25. Predicting Time-Dependent Flow Over Complex Geometries Using Operator Networks 1 upvotes, #25 of 2025-12-10
  26. SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images 1 upvotes, #25 of 2025-12-10
  27. Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs 1 upvotes, #27 of 2025-12-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.