Daily Papers of 2025-12-08
- TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows 69 upvotes, #1 of 2025-12-08
- EditThinker: Unlocking Iterative Reasoning for Any Image Editor 36 upvotes, #2 of 2025-12-08
- From Imitation to Discrimination: Toward A Generalized Curriculum Advantage Mechanism Enhancing Cross-Domain Reasoning Tasks 27 upvotes, #3 of 2025-12-08
- EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture 25 upvotes, #4 of 2025-12-08
- RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards 24 upvotes, #5 of 2025-12-08
- PaCo-RL: Advancing Reinforcement Learning for Consistent Image Generation with Pairwise Reward Modeling 24 upvotes, #5 of 2025-12-08
- SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling 23 upvotes, #7 of 2025-12-08
- SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations 19 upvotes, #8 of 2025-12-08
- Self-Improving VLM Judges Without Human Annotations 18 upvotes, #9 of 2025-12-08
- Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning 16 upvotes, #10 of 2025-12-08
- Joint 3D Geometry Reconstruction and Motion Generation for 4D Synthesis from a Single Image 15 upvotes, #11 of 2025-12-08
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence 13 upvotes, #12 of 2025-12-08
- World Models That Know When They Don't Know: Controllable Video Generation with Calibrated Uncertainty 10 upvotes, #13 of 2025-12-08
- ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning 9 upvotes, #14 of 2025-12-08
- AI & Human Co-Improvement for Safer Co-Superintelligence 8 upvotes, #15 of 2025-12-08
- M3DR: Towards Universal Multilingual Multimodal Document Retrieval 7 upvotes, #16 of 2025-12-08
- Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding 6 upvotes, #17 of 2025-12-08
- From Segments to Scenes: Temporal Understanding in Autonomous Driving via Vision-Language Model 4 upvotes, #18 of 2025-12-08
- ProPhy: Progressive Physical Alignment for Dynamic World Simulation 4 upvotes, #18 of 2025-12-08
- Colon-X: Advancing Intelligent Colonoscopy from Multimodal Understanding to Clinical Reasoning 3 upvotes, #20 of 2025-12-08
- TimesNet-Gen: Deep Learning-based Site Specific Strong Motion Generation 2 upvotes, #21 of 2025-12-08
- SQ-format: A Unified Sparse-Quantized Hardware-friendly Data Format for LLMs 2 upvotes, #21 of 2025-12-08
- From FLOPs to Footprints: The Resource Cost of Artificial Intelligence 1 upvotes, #23 of 2025-12-08
- Taxonomy-Adaptive Moderation Model with Robust Guardrails for Large Language Models 1 upvotes, #24 of 2025-12-08
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.