Daily Papers of 2025-04-03
- MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization 78 upvotes, #1 of 2025-04-03
- Improved Visual-Spatial Reasoning via R1-Zero-Like Training 60 upvotes, #2 of 2025-04-03
- DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance 60 upvotes, #2 of 2025-04-03
- AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction 57 upvotes, #4 of 2025-04-03
- Towards Physically Plausible Video Generation via VLM Planning 38 upvotes, #5 of 2025-04-03
- ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations 38 upvotes, #5 of 2025-04-03
- Understanding R1-Zero-Like Training: A Critical Perspective 36 upvotes, #7 of 2025-04-03
- VideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One Step 36 upvotes, #7 of 2025-04-03
- PaperBench: Evaluating AI's Ability to Replicate AI Research 34 upvotes, #9 of 2025-04-03
- Articulated Kinematics Distillation from Video Diffusion Models 22 upvotes, #10 of 2025-04-03
- ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement 20 upvotes, #11 of 2025-04-03
- Boost Your Own Human Image Generation Model via Direct Preference Optimization with AI Feedback 19 upvotes, #12 of 2025-04-03
- Safeguarding Vision-Language Models: Mitigating Vulnerabilities to Gaussian Noise in Perturbation-based Attacks 13 upvotes, #13 of 2025-04-03
- MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis 12 upvotes, #14 of 2025-04-03
- DASH: Detection and Assessment of Systematic Hallucinations of VLMs 12 upvotes, #14 of 2025-04-03
- Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models 10 upvotes, #16 of 2025-04-03
- LSNet: See Large, Focus Small 8 upvotes, #17 of 2025-04-03
- Adaptive Layer-skipping in Pre-trained LLMs 6 upvotes, #18 of 2025-04-03
- VerifiAgent: a Unified Verification Agent in Language Model Reasoning 6 upvotes, #18 of 2025-04-03
- Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations 3 upvotes, #20 of 2025-04-03
- Target-Aware Video Diffusion Models 3 upvotes, #20 of 2025-04-03
- Medical large language models are easily distracted 3 upvotes, #20 of 2025-04-03
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.