Daily Papers of 2025-12-01

  1. Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer 163 upvotes, #1 of 2025-12-01
  2. DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning 66 upvotes, #2 of 2025-12-01
  3. REASONEDIT: Towards Reasoning-Enhanced Image Editing Models 45 upvotes, #3 of 2025-12-01
  4. Vision Bridge Transformer at Scale 43 upvotes, #4 of 2025-12-01
  5. AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement 41 upvotes, #5 of 2025-12-01
  6. Geometrically-Constrained Agent for Spatial Reasoning 38 upvotes, #6 of 2025-12-01
  7. Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models 29 upvotes, #7 of 2025-12-01
  8. Architecture Decoupling Is Not All You Need For Unified Multimodal Model 28 upvotes, #8 of 2025-12-01
  9. DiP: Taming Diffusion Models in Pixel Space 25 upvotes, #9 of 2025-12-01
  10. CaptionQA: Is Your Caption as Useful as the Image Itself? 25 upvotes, #9 of 2025-12-01
  11. Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield 22 upvotes, #11 of 2025-12-01
  12. DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action 21 upvotes, #12 of 2025-12-01
  13. Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models 21 upvotes, #12 of 2025-12-01
  14. Adversarial Flow Models 20 upvotes, #14 of 2025-12-01
  15. RefineBench: Evaluating Refinement Capability of Language Models via Checklists 12 upvotes, #15 of 2025-12-01
  16. Captain Safari: A World Engine 9 upvotes, #16 of 2025-12-01
  17. World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models 8 upvotes, #17 of 2025-12-01
  18. The Collapse of Patches 6 upvotes, #18 of 2025-12-01
  19. SO-Bench: A Structural Output Evaluation of Multimodal LLMs 5 upvotes, #19 of 2025-12-01
  20. OralGPT-Omni: A Versatile Dental Multimodal Large Language Model 5 upvotes, #19 of 2025-12-01
  21. Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information 4 upvotes, #21 of 2025-12-01
  22. Test-time scaling of diffusions with flow maps 4 upvotes, #21 of 2025-12-01
  23. OmniRefiner: Reinforcement-Guided Local Diffusion Refinement 3 upvotes, #23 of 2025-12-01
  24. From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images 3 upvotes, #23 of 2025-12-01
  25. YOLO Meets Mixture-of-Experts: Adaptive Expert Routing for Robust Object Detection 2 upvotes, #25 of 2025-12-01
  26. Layer-Aware Video Composition via Split-then-Merge 2 upvotes, #25 of 2025-12-01
  27. FedRE: A Representation Entanglement Framework for Model-Heterogeneous Federated Learning 1 upvotes, #27 of 2025-12-01
  28. Fast3Dcache: Training-free 3D Geometry Synthesis Acceleration 1 upvotes, #27 of 2025-12-01
  29. Recognition of Abnormal Events in Surveillance Videos using Weakly Supervised Dual-Encoder Models 1 upvotes, #29 of 2025-12-01
  30. Find the Leak, Fix the Split: Cluster-Based Method to Prevent Leakage in Video-Derived Datasets 1 upvotes, #29 of 2025-12-01
  31. MRI Super-Resolution with Deep Learning: A Comprehensive Survey 1 upvotes, #29 of 2025-12-01
  32. Xmodel-2.5: 1.3B Data-Efficient Reasoning SLM 1 upvotes, #29 of 2025-12-01

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.