Daily Papers of 2025-12-11

  1. StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation 69 upvotes, #1 of 2025-12-11
  2. OmniPSD: Layered PSD Generation with Diffusion Transformer 45 upvotes, #2 of 2025-12-11
  3. BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain 37 upvotes, #3 of 2025-12-11
  4. Composing Concepts from Images and Videos via Concept-prompt Binding 27 upvotes, #4 of 2025-12-11
  5. MotionEdit: Benchmarking and Learning Motion-Centric Image Editing 25 upvotes, #5 of 2025-12-11
  6. InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models 17 upvotes, #6 of 2025-12-11
  7. Rethinking Chain-of-Thought Reasoning for Videos 16 upvotes, #7 of 2025-12-11
  8. Towards a Science of Scaling Agent Systems 12 upvotes, #8 of 2025-12-11
  9. WonderZoom: Multi-Scale 3D World Generation 11 upvotes, #9 of 2025-12-11
  10. HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models 11 upvotes, #9 of 2025-12-11
  11. UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving 10 upvotes, #11 of 2025-12-11
  12. Fast-Decoding Diffusion Language Models via Progress-Aware Confidence Schedules 9 upvotes, #12 of 2025-12-11
  13. EtCon: Edit-then-Consolidate for Reliable Knowledge Editing 7 upvotes, #13 of 2025-12-11
  14. TED-4DGS: Temporally Activated and Embedding-based Deformation for 4DGS Compression 7 upvotes, #13 of 2025-12-11
  15. Learning Unmasking Policies for Diffusion Language Models 7 upvotes, #13 of 2025-12-11
  16. Reinventing Clinical Dialogue: Agentic Paradigms for LLM Enabled Healthcare Communication 3 upvotes, #16 of 2025-12-11
  17. VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory 3 upvotes, #16 of 2025-12-11
  18. IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting 3 upvotes, #16 of 2025-12-11
  19. Beyond Unified Models: A Service-Oriented Approach to Low Latency, Context Aware Phonemization for Real Time TTS 2 upvotes, #19 of 2025-12-11
  20. Smart Timing for Mining: A Deep Learning Framework for Bitcoin Hardware ROI Prediction 1 upvotes, #20 of 2025-12-11
  21. Pay Less Attention to Function Words for Free Robustness of Vision-Language Models 1 upvotes, #20 of 2025-12-11
  22. GimbalDiffusion: Gravity-Aware Camera Control for Video Generation 1 upvotes, #20 of 2025-12-11

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.