Daily Papers of 2025-12-05

  1. Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length 166 upvotes, #1 of 2025-12-05
  2. DAComp: Benchmarking Data Agents across the Full Data Intelligence Lifecycle 147 upvotes, #2 of 2025-12-05
  3. Nex-N1: Agentic Models Trained via a Unified Ecosystem for Large-Scale Environment Construction 71 upvotes, #3 of 2025-12-05
  4. PaperDebugger: A Plugin-Based Multi-Agent System for In-Editor Academic Writing, Review, and Editing 58 upvotes, #4 of 2025-12-05
  5. ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning 45 upvotes, #5 of 2025-12-05
  6. Semantics Lead the Way: Harmonizing Semantic and Texture Modeling with Asynchronous Latent Diffusion 40 upvotes, #6 of 2025-12-05
  7. Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation 39 upvotes, #7 of 2025-12-05
  8. DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling 35 upvotes, #8 of 2025-12-05
  9. Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression 30 upvotes, #9 of 2025-12-05
  10. SIMA 2: A Generalist Embodied Agent for Virtual Worlds 19 upvotes, #10 of 2025-12-05
  11. 4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer 18 upvotes, #11 of 2025-12-05
  12. Splannequin: Freezing Monocular Mannequin-Challenge Footage with Dual-Detection Splatting 17 upvotes, #12 of 2025-12-05
  13. UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers 15 upvotes, #13 of 2025-12-05
  14. NeuralRemaster: Phase-Preserving Diffusion for Structure-Aligned Generation 15 upvotes, #13 of 2025-12-05
  15. Reflection Removal through Efficient Adaptation of Diffusion Transformers 14 upvotes, #15 of 2025-12-05
  16. TV2TV: A Unified Framework for Interleaved Language and Video Generation 14 upvotes, #15 of 2025-12-05
  17. SignRoundV2: Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMs 12 upvotes, #17 of 2025-12-05
  18. On GRPO Collapse in Search-R1: The Lazy Likelihood-Displacement Death Spiral 11 upvotes, #18 of 2025-12-05
  19. Model-Based and Sample-Efficient AI-Assisted Math Discovery in Sphere Packing 11 upvotes, #18 of 2025-12-05
  20. DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation 11 upvotes, #18 of 2025-12-05
  21. LATTICE: Democratize High-Fidelity 3D Generation at Scale 9 upvotes, #21 of 2025-12-05
  22. Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment 9 upvotes, #21 of 2025-12-05
  23. SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization 8 upvotes, #23 of 2025-12-05
  24. FMA-Net++: Motion- and Exposure-Aware Real-World Joint Video Super-Resolution and Deblurring 8 upvotes, #23 of 2025-12-05
  25. Generative Neural Video Compression via Video Diffusion Prior 8 upvotes, #23 of 2025-12-05
  26. Some Modalities are More Equal Than Others: Decoding and Architecting Multimodal Integration in MLLMs 7 upvotes, #26 of 2025-12-05
  27. Aligned but Stereotypical? The Hidden Influence of System Prompts on Social Bias in LVLM-Based Text-to-Image Models 7 upvotes, #26 of 2025-12-05
  28. BulletTime: Decoupled Control of Time and Camera Pose for Video Generation 7 upvotes, #26 of 2025-12-05
  29. When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models 5 upvotes, #29 of 2025-12-05
  30. EgoLCD: Egocentric Video Generation with Long Context Diffusion 5 upvotes, #29 of 2025-12-05
  31. Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos 4 upvotes, #31 of 2025-12-05
  32. Mitigating Catastrophic Forgetting in Target Language Adaptation of LLMs via Source-Shielded Updates 4 upvotes, #31 of 2025-12-05
  33. ShadowDraw: From Any Object to Shadow-Drawing Compositional Art 3 upvotes, #33 of 2025-12-05
  34. GaussianBlender: Instant Stylization of 3D Gaussians with Disentangled Latent Spaces 2 upvotes, #34 of 2025-12-05
  35. REFLEX: Self-Refining Explainable Fact-Checking via Disentangling Truth into Style and Substance 1 upvotes, #35 of 2025-12-05
  36. Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models 1 upvotes, #35 of 2025-12-05
  37. QKAN-LSTM: Quantum-inspired Kolmogorov-Arnold Long Short-term Memory 1 upvotes, #35 of 2025-12-05
  38. A Theoretical Framework for Auxiliary-Loss-Free Load Balancing of Sparse Mixture-of-Experts in Large-Scale AI Models 2 upvotes, #38 of 2025-12-05

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.