Daily Papers of 2025-12-05
- Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length 166 upvotes, #1 of 2025-12-05
- DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle 147 upvotes, #2 of 2025-12-05
- Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction 71 upvotes, #3 of 2025-12-05
- PaperDebugger: A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing 58 upvotes, #4 of 2025-12-05
- ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning 45 upvotes, #5 of 2025-12-05
- Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion 40 upvotes, #6 of 2025-12-05
- Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation 39 upvotes, #7 of 2025-12-05
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling 35 upvotes, #8 of 2025-12-05
- Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression 30 upvotes, #9 of 2025-12-05
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds 19 upvotes, #10 of 2025-12-05
- 4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer 18 upvotes, #11 of 2025-12-05
- Splannequin: Freezing Monocular Mannequin-Challenge Footage with Dual-Detection Splatting 17 upvotes, #12 of 2025-12-05
- UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers 15 upvotes, #13 of 2025-12-05
- NeuralRemaster: Phase-Preserving Diffusion for Structure-Aligned Generation 15 upvotes, #13 of 2025-12-05
- Reflection Removal through Efficient Adaptation of Diffusion Transformers 14 upvotes, #15 of 2025-12-05
- TV2TV: A Unified Framework for Interleaved Language and Video Generation 14 upvotes, #15 of 2025-12-05
- SignRoundV2: Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs 12 upvotes, #17 of 2025-12-05
- On GRPO Collapse in Search-R1: The Lazy Likelihood-Displacement Death Spiral 11 upvotes, #18 of 2025-12-05
- Model-Based and Sample-Efficient AI-Assisted Math Discovery in Sphere Packing 11 upvotes, #18 of 2025-12-05
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation 11 upvotes, #18 of 2025-12-05
- LATTICE: Democratize High-Fidelity 3D Generation at Scale 9 upvotes, #21 of 2025-12-05
- Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment 9 upvotes, #21 of 2025-12-05
- SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization 8 upvotes, #23 of 2025-12-05
- FMA-Net++: Motion- and Exposure-Aware Real-World Joint Video Super-Resolution and Deblurring 8 upvotes, #23 of 2025-12-05
- Generative Neural Video Compression via Video Diffusion Prior 8 upvotes, #23 of 2025-12-05
- Some Modalities are More Equal Than Others: Decoding and Architecting Multimodal Integration in MLLMs 7 upvotes, #26 of 2025-12-05
- Aligned but Stereotypical? The Hidden Influence of System Prompts on Social Bias in LVLM-Based Text-to-Image Models 7 upvotes, #26 of 2025-12-05
- BulletTime: Decoupled Control of Time and Camera Pose for Video Generation 7 upvotes, #26 of 2025-12-05
- When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models 5 upvotes, #29 of 2025-12-05
- EgoLCD: Egocentric Video Generation with Long Context Diffusion 5 upvotes, #29 of 2025-12-05
- Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos 4 upvotes, #31 of 2025-12-05
- Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates 4 upvotes, #31 of 2025-12-05
- ShadowDraw: From Any Object to Shadow-Drawing Compositional Art 3 upvotes, #33 of 2025-12-05
- GaussianBlender: Instant Stylization of 3D Gaussians with Disentangled Latent Spaces 2 upvotes, #34 of 2025-12-05
- REFLEX: Self-Refining Explainable Fact-Checking via Disentangling Truth into Style and Substance 1 upvotes, #35 of 2025-12-05
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models 1 upvotes, #35 of 2025-12-05
- QKAN-LSTM: Quantum-inspired Kolmogorov-Arnold Long Short-term Memory 1 upvotes, #35 of 2025-12-05
- A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models 2 upvotes, #38 of 2025-12-05
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.