Daily Papers of 2025-12-01
- Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer 163 upvotes, #1 of 2025-12-01
- DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning 66 upvotes, #2 of 2025-12-01
- REASONEDIT: Towards Reasoning-Enhanced Image Editing Models 45 upvotes, #3 of 2025-12-01
- Vision Bridge Transformer at Scale 43 upvotes, #4 of 2025-12-01
- AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity Refinement 41 upvotes, #5 of 2025-12-01
- Geometrically-Constrained Agent for Spatial Reasoning 38 upvotes, #6 of 2025-12-01
- Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models 29 upvotes, #7 of 2025-12-01
- Architecture Decoupling Is Not All You Need For Unified Multimodal Model 28 upvotes, #8 of 2025-12-01
- DiP: Taming Diffusion Models in Pixel Space 25 upvotes, #9 of 2025-12-01
- CaptionQA: Is Your Caption as Useful as the Image Itself? 25 upvotes, #9 of 2025-12-01
- Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield 22 upvotes, #11 of 2025-12-01
- DualVLA: Building a Generalizable Embodied Agent via Partial Decoupling of Reasoning and Action 21 upvotes, #12 of 2025-12-01
- Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models 21 upvotes, #12 of 2025-12-01
- Adversarial Flow Models 20 upvotes, #14 of 2025-12-01
- RefineBench: Evaluating Refinement Capability of Language Models via Checklists 12 upvotes, #15 of 2025-12-01
- Captain Safari: A World Engine 9 upvotes, #16 of 2025-12-01
- World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models 8 upvotes, #17 of 2025-12-01
- The Collapse of Patches 6 upvotes, #18 of 2025-12-01
- SO-Bench: A Structural Output Evaluation of Multimodal LLMs 5 upvotes, #19 of 2025-12-01
- OralGPT-Omni: A Versatile Dental Multimodal Large Language Model 5 upvotes, #19 of 2025-12-01
- Focused Chain-of-Thought: Efficient LLM Reasoning via Structured Input Information 4 upvotes, #21 of 2025-12-01
- Test-time scaling of diffusions with flow maps 4 upvotes, #21 of 2025-12-01
- OmniRefiner: Reinforcement-Guided Local Diffusion Refinement 3 upvotes, #23 of 2025-12-01
- From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images 3 upvotes, #23 of 2025-12-01
- YOLO Meets Mixture-of-Experts: Adaptive Expert Routing for Robust Object Detection 2 upvotes, #25 of 2025-12-01
- Layer-Aware Video Composition via Split-then-Merge 2 upvotes, #25 of 2025-12-01
- FedRE: A Representation Entanglement Framework for Model-Heterogeneous Federated Learning 1 upvotes, #27 of 2025-12-01
- Fast3Dcache: Training-free 3D Geometry Synthesis Acceleration 1 upvotes, #27 of 2025-12-01
- Recognition of Abnormal Events in Surveillance Videos using Weakly Supervised Dual-Encoder Models 1 upvotes, #29 of 2025-12-01
- Find the Leak, Fix the Split: Cluster-Based Method to Prevent Leakage in Video-Derived Datasets 1 upvotes, #29 of 2025-12-01
- MRI Super-Resolution with Deep Learning: A Comprehensive Survey 1 upvotes, #29 of 2025-12-01
- Xmodel-2.5: 1.3B Data-Efficient Reasoning SLM 1 upvotes, #29 of 2025-12-01
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.