Daily Papers of 2025-04-03

  1. MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization 78 upvotes, #1 of 2025-04-03
  2. Improved Visual-Spatial Reasoning via R1-Zero-Like Training 60 upvotes, #2 of 2025-04-03
  3. DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance 60 upvotes, #2 of 2025-04-03
  4. AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction 57 upvotes, #4 of 2025-04-03
  5. Towards Physically Plausible Video Generation via VLM Planning 38 upvotes, #5 of 2025-04-03
  6. ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations 38 upvotes, #5 of 2025-04-03
  7. Understanding R1-Zero-Like Training: A Critical Perspective 36 upvotes, #7 of 2025-04-03
  8. VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step 36 upvotes, #7 of 2025-04-03
  9. PaperBench: Evaluating AI's Ability to Replicate AI Research 34 upvotes, #9 of 2025-04-03
  10. Articulated Kinematics Distillation from Video Diffusion Models 22 upvotes, #10 of 2025-04-03
  11. ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement 20 upvotes, #11 of 2025-04-03
  12. Boost Your Own Human Image Generation Model via Direct Preference Optimization with AI Feedback 19 upvotes, #12 of 2025-04-03
  13. Safeguarding Vision-Language Models: Mitigating Vulnerabilities to Gaussian Noise in Perturbation-based Attacks 13 upvotes, #13 of 2025-04-03
  14. MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis 12 upvotes, #14 of 2025-04-03
  15. DASH: Detection and Assessment of Systematic Hallucinations of VLMs 12 upvotes, #14 of 2025-04-03
  16. Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models 10 upvotes, #16 of 2025-04-03
  17. LSNet: See Large, Focus Small 8 upvotes, #17 of 2025-04-03
  18. Adaptive Layer-skipping in Pre-trained LLMs 6 upvotes, #18 of 2025-04-03
  19. VerifiAgent: a Unified Verification Agent in Language Model Reasoning 6 upvotes, #18 of 2025-04-03
  20. Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations 3 upvotes, #20 of 2025-04-03
  21. Target-Aware Video Diffusion Models 3 upvotes, #20 of 2025-04-03
  22. Medical large language models are easily distracted 3 upvotes, #20 of 2025-04-03

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.