Daily Papers of 2025-12-03

  1. DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models 194 upvotes, #1 of 2025-12-03
  2. ToolOrchestra: Elevating Intelligence via Efficient Model and Tool Orchestration 100 upvotes, #2 of 2025-12-03
  3. Deep Research: A Systematic Survey 62 upvotes, #3 of 2025-12-03
  4. MultiShotMaster: A Controllable Multi-Shot Video Generation Framework 62 upvotes, #3 of 2025-12-03
  5. Guided Self-Evolving LLMs with Minimal Human Supervision 48 upvotes, #5 of 2025-12-03
  6. MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory 47 upvotes, #6 of 2025-12-03
  7. Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch 45 upvotes, #7 of 2025-12-03
  8. DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation 43 upvotes, #8 of 2025-12-03
  9. SimScale: Learning to Drive via Real-World Simulation at Scale 37 upvotes, #9 of 2025-12-03
  10. SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds 33 upvotes, #10 of 2025-12-03
  11. InnoGym: Benchmarking the Innovation Potential of AI Agents 33 upvotes, #10 of 2025-12-03
  12. PixelDiT: Pixel Diffusion Transformers for Image Generation 26 upvotes, #12 of 2025-12-03
  13. Glance: Accelerating Diffusion Models with 1 Sample 25 upvotes, #13 of 2025-12-03
  14. WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning 22 upvotes, #14 of 2025-12-03
  15. ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation 20 upvotes, #15 of 2025-12-03
  16. Mixture of Horizons in Action Chunking 17 upvotes, #16 of 2025-12-03
  17. WUSH: Near-Optimal Adaptive Transforms for LLM Quantization 17 upvotes, #16 of 2025-12-03
  18. Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation 13 upvotes, #18 of 2025-12-03
  19. GoRL: An Algorithm-Agnostic Framework for Online Reinforcement Learning with Generative Policies 13 upvotes, #18 of 2025-12-03
  20. The Curious Case of Analogies: Investigating Analogical Reasoning in Large Language Models 12 upvotes, #20 of 2025-12-03
  21. CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning 11 upvotes, #21 of 2025-12-03
  22. MagicQuillV2: Precise and Interactive Image Editing with Layered Visual Cues 11 upvotes, #21 of 2025-12-03
  23. TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition 9 upvotes, #23 of 2025-12-03
  24. DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models 9 upvotes, #23 of 2025-12-03
  25. RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence 9 upvotes, #23 of 2025-12-03
  26. Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization 6 upvotes, #26 of 2025-12-03
  27. Video4Spatial: Towards Visuospatial Intelligence with Context-Guided Video Generation 6 upvotes, #26 of 2025-12-03
  28. SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead 5 upvotes, #28 of 2025-12-03
  29. PAI-Bench: A Comprehensive Benchmark For Physical AI 5 upvotes, #28 of 2025-12-03
  30. Ovis-Image Technical Report 4 upvotes, #30 of 2025-12-03
  31. YingVideo-MV: Music-Driven Multi-Stage Video Generation 4 upvotes, #30 of 2025-12-03
  32. Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents 3 upvotes, #32 of 2025-12-03
  33. C^2DLM: Causal Concept-Guided Diffusion Large Language Models 3 upvotes, #32 of 2025-12-03
  34. BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation 3 upvotes, #32 of 2025-12-03
  35. FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention 3 upvotes, #32 of 2025-12-03
  36. Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion 3 upvotes, #32 of 2025-12-03
  37. GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning 3 upvotes, #32 of 2025-12-03
  38. UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits 3 upvotes, #32 of 2025-12-03
  39. Benchmarking Scientific Understanding and Reasoning for Video Generation using VideoScience-Bench 2 upvotes, #39 of 2025-12-03
  40. In-Context Sync-LoRA for Portrait Video Editing 2 upvotes, #39 of 2025-12-03
  41. CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization 1 upvotes, #41 of 2025-12-03
  42. Masks Can Be Distracting: On Context Comprehension in Diffusion Language Models 1 upvotes, #41 of 2025-12-03
  43. Shoe Style-Invariant and Ground-Aware Learning for Dense Foot Contact Estimation 1 upvotes, #41 of 2025-12-03
  44. Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions 1 upvotes, #41 of 2025-12-03
  45. Artemis: Structured Visual Reasoning for Perception Policy Learning 1 upvotes, #41 of 2025-12-03
  46. Understanding and Harnessing Sparsity in Unified Multimodal Models 1 upvotes, #41 of 2025-12-03
  47. BOOM: Beyond Only One Modality KIT's Multimodal Multilingual Lecture Companion 1 upvotes, #41 of 2025-12-03
  48. Click2Graph: Interactive Panoptic Video Scene Graphs from a Single Click 1 upvotes, #48 of 2025-12-03

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.