Daily Papers of 2025-12-03
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models 194 upvotes, #1 of 2025-12-03
- ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration 100 upvotes, #2 of 2025-12-03
- Deep Research: A Systematic Survey 62 upvotes, #3 of 2025-12-03
- MultiShotMaster: A Controllable Multi-Shot Video Generation Framework 62 upvotes, #3 of 2025-12-03
- Guided Self-Evolving LLMs with Minimal Human Supervision 48 upvotes, #5 of 2025-12-03
- MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory 47 upvotes, #6 of 2025-12-03
- Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch 45 upvotes, #7 of 2025-12-03
- DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation 43 upvotes, #8 of 2025-12-03
- SimScale: Learning to Drive via Real-World Simulation at Scale 37 upvotes, #9 of 2025-12-03
- SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds 33 upvotes, #10 of 2025-12-03
- InnoGym: Benchmarking the Innovation Potential of AI Agents 33 upvotes, #10 of 2025-12-03
- PixelDiT: Pixel Diffusion Transformers for Image Generation 26 upvotes, #12 of 2025-12-03
- Glance: Accelerating Diffusion Models with 1 Sample 25 upvotes, #13 of 2025-12-03
- WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning 22 upvotes, #14 of 2025-12-03
- ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation 20 upvotes, #15 of 2025-12-03
- Mixture of Horizons in Action Chunking 17 upvotes, #16 of 2025-12-03
- WUSH: Near-Optimal Adaptive Transforms for LLM Quantization 17 upvotes, #16 of 2025-12-03
- Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation 13 upvotes, #18 of 2025-12-03
- GoRL: An Algorithm-Agnostic Framework for Online Reinforcement Learning with Generative Policies 13 upvotes, #18 of 2025-12-03
- The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models 12 upvotes, #20 of 2025-12-03
- CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning 11 upvotes, #21 of 2025-12-03
- MagicQuillV2: Precise and Interactive Image Editing with Layered Visual Cues 11 upvotes, #21 of 2025-12-03
- TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition 9 upvotes, #23 of 2025-12-03
- DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models 9 upvotes, #23 of 2025-12-03
- RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence 9 upvotes, #23 of 2025-12-03
- Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization 6 upvotes, #26 of 2025-12-03
- Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation 6 upvotes, #26 of 2025-12-03
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead 5 upvotes, #28 of 2025-12-03
- PAI-Bench: A Comprehensive Benchmark For Physical AI 5 upvotes, #28 of 2025-12-03
- Ovis-Image Technical Report 4 upvotes, #30 of 2025-12-03
- YingVideo-MV: Music-Driven Multi-Stage Video Generation 4 upvotes, #30 of 2025-12-03
- Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents 3 upvotes, #32 of 2025-12-03
- C^2DLM: Causal Concept-Guided Diffusion Large Language Models 3 upvotes, #32 of 2025-12-03
- BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation 3 upvotes, #32 of 2025-12-03
- FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention 3 upvotes, #32 of 2025-12-03
- Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion 3 upvotes, #32 of 2025-12-03
- GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning 3 upvotes, #32 of 2025-12-03
- UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits 3 upvotes, #32 of 2025-12-03
- Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench 2 upvotes, #39 of 2025-12-03
- In-Context Sync-LoRA for Portrait Video Editing 2 upvotes, #39 of 2025-12-03
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization 1 upvotes, #41 of 2025-12-03
- Masks Can Be Distracting: On Context Comprehension in Diffusion Language Models 1 upvotes, #41 of 2025-12-03
- Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimation 1 upvotes, #41 of 2025-12-03
- Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions 1 upvotes, #41 of 2025-12-03
- Artemis: Structured Visual Reasoning for Perception Policy Learning 1 upvotes, #41 of 2025-12-03
- Understanding and Harnessing Sparsity in Unified Multimodal Models 1 upvotes, #41 of 2025-12-03
- BOOM: Beyond Only One Modality KIT's Multimodal Multilingual Lecture Companion 1 upvotes, #41 of 2025-12-03
- Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click 1 upvotes, #48 of 2025-12-03
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.