Daily Papers of 2025-12-10
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance 123 upvotes, #1 of 2025-12-10
- Visionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting Platform 75 upvotes, #2 of 2025-12-10
- Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality 46 upvotes, #3 of 2025-12-10
- OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory 43 upvotes, #4 of 2025-12-10
- DeepCode: Open Agentic Coding 29 upvotes, #5 of 2025-12-10
- From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs 23 upvotes, #6 of 2025-12-10
- Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation 21 upvotes, #7 of 2025-12-10
- ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models 19 upvotes, #8 of 2025-12-10
- Arbitrage: Efficient Reasoning via Advantage-Aware Speculation 15 upvotes, #9 of 2025-12-10
- MIND-V: Hierarchical Video Generation for Long-Horizon Robotic Manipulation with RL-based Physical Alignment 12 upvotes, #10 of 2025-12-10
- Boosting Unsupervised Video Instance Segmentation with Automatic Quality-Guided Self-Training 12 upvotes, #10 of 2025-12-10
- COREA: Coarse-to-Fine 3D Representation Alignment Between Relightable 3D Gaussians and SDF via Bidirectional 3D-to-3D Supervision 8 upvotes, #12 of 2025-12-10
- See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models 7 upvotes, #13 of 2025-12-10
- SUCCESS-GS: Survey of Compactness and Compression for Efficient Static and Dynamic Gaussian Splatting 6 upvotes, #14 of 2025-12-10
- TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models 6 upvotes, #14 of 2025-12-10
- Efficiently Reconstructing Dynamic Scenes One D4RT at a Time 6 upvotes, #14 of 2025-12-10
- Modular Neural Image Signal Processing 5 upvotes, #17 of 2025-12-10
- MemLoRA: Distilling Expert Adapters for On-Device Memory Systems 3 upvotes, #18 of 2025-12-10
- TrackingWorld: World-centric Monocular 3D Tracking of Almost All Pixels 3 upvotes, #18 of 2025-12-10
- LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning 2 upvotes, #20 of 2025-12-10
- Novel Deep Learning Architectures for Classification and Segmentation of Brain Tumors from MRI Images 2 upvotes, #20 of 2025-12-10
- Terrain Diffusion: A Diffusion-Based Successor to Perlin Noise in Infinite, Real-Time Terrain Generation 2 upvotes, #20 of 2025-12-10
- SAM-Body4D: Training-Free 4D Human Body Mesh Recovery from Videos 2 upvotes, #20 of 2025-12-10
- EcomBench: Towards Holistic Evaluation of Foundation Agents in E-commerce 2 upvotes, #20 of 2025-12-10
- Predicting Time-Dependent Flow Over Complex Geometries Using Operator Networks 1 upvotes, #25 of 2025-12-10
- SegEarth-OV3: Exploring SAM 3 for Open-Vocabulary Semantic Segmentation in Remote Sensing Images 1 upvotes, #25 of 2025-12-10
- Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs 1 upvotes, #27 of 2025-12-10
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.