Daily Papers of 2025-03-21
- One-Step Residual Shifting Diffusion for Image Super-Resolution via Distillation 89 upvotes, #1 of 2025-03-21
- Survey on Evaluation of LLM-based Agents 78 upvotes, #2 of 2025-03-21
- Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models 62 upvotes, #3 of 2025-03-21
- Reinforcement Learning for Reasoning in Small LLMs: What Works and What Doesn't 45 upvotes, #4 of 2025-03-21
- Why Do Multi-Agent LLM Systems Fail? 41 upvotes, #5 of 2025-03-21
- Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning 41 upvotes, #5 of 2025-03-21
- Unleashing Vecset Diffusion Model for Fast Shape Generation 40 upvotes, #7 of 2025-03-21
- Inside-Out: Hidden Factual Knowledge in LLMs 39 upvotes, #8 of 2025-03-21
- Scale-wise Distillation of Diffusion Models 35 upvotes, #9 of 2025-03-21
- JARVIS-VLA: Post-Training Large-Scale Vision Language Models to Play Visual Games with Keyboards and Mouse 34 upvotes, #10 of 2025-03-21
- InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity 32 upvotes, #11 of 2025-03-21
- DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers 27 upvotes, #12 of 2025-03-21
- LHM: Large Animatable Human Reconstruction Model from a Single Image in Seconds 25 upvotes, #13 of 2025-03-21
- Fin-R1: A Large Language Model for Financial Reasoning through Reinforcement Learning 25 upvotes, #13 of 2025-03-21
- Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models 23 upvotes, #15 of 2025-03-21
- SynCity: Training-Free Generation of 3D Worlds 23 upvotes, #15 of 2025-03-21
- MathFusion: Enhancing Mathematic Problem-solving of LLM through Instruction Fusion 22 upvotes, #17 of 2025-03-21
- CaKE: Circuit-aware Editing Enables Generalizable Knowledge Learners 15 upvotes, #18 of 2025-03-21
- Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts 14 upvotes, #19 of 2025-03-21
- M3: 3D-Spatial MultiModal Memory 14 upvotes, #19 of 2025-03-21
- 1000+ FPS 4D Gaussian Splatting for Dynamic Scene Rendering 14 upvotes, #19 of 2025-03-21
- Tokenize Image as a Set 13 upvotes, #22 of 2025-03-21
- MotionStreamer: Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent Space 12 upvotes, #23 of 2025-03-21
- Ultra-Resolution Adaptation with Ease 12 upvotes, #23 of 2025-03-21
- XAttention: Block Sparse Attention with Antidiagonal Scoring 12 upvotes, #23 of 2025-03-21
- Sonata: Self-Supervised Learning of Reliable Point Representations 11 upvotes, #26 of 2025-03-21
- Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion 10 upvotes, #27 of 2025-03-21
- BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity? 9 upvotes, #28 of 2025-03-21
- CLS-RL: Image Classification with Rule-Based Reinforcement Learning 9 upvotes, #28 of 2025-03-21
- NuiScene: Exploring Efficient Generation of Unbounded Outdoor Scenes 9 upvotes, #28 of 2025-03-21
- Agents Play Thousands of 3D Video Games 8 upvotes, #31 of 2025-03-21
- Where do Large Vision-Language Models Look at when Answering Questions? 8 upvotes, #31 of 2025-03-21
- SALT: Singular Value Adaptation with Low-Rank Transformation 8 upvotes, #31 of 2025-03-21
- MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance 8 upvotes, #31 of 2025-03-21
- Uni-3DAR: Unified 3D Generation and Understanding via Autoregression on Compressed Spatial Tokens 7 upvotes, #35 of 2025-03-21
- Towards Unified Latent Space for 3D Molecular Latent Diffusion Modeling 6 upvotes, #36 of 2025-03-21
- Improving Autoregressive Image Generation through Coarse-to-Fine Token Prediction 6 upvotes, #36 of 2025-03-21
- See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality Bias 5 upvotes, #38 of 2025-03-21
- Make Your Training Flexible: Towards Deployment-Efficient Video Models 5 upvotes, #38 of 2025-03-21
- Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning 4 upvotes, #40 of 2025-03-21
- UVE: Are MLLMs Unified Evaluators for AI-Generated Videos? 4 upvotes, #40 of 2025-03-21
- MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization 4 upvotes, #40 of 2025-03-21
- TikZero: Zero-Shot Text-Guided Graphics Program Synthesis 3 upvotes, #43 of 2025-03-21
- Why Personalizing Deep Learning-Based Code Completion Tools Matters 3 upvotes, #43 of 2025-03-21
- GASP: Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving 3 upvotes, #43 of 2025-03-21
- VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling 3 upvotes, #43 of 2025-03-21
- Deceptive Humor: A Synthetic Multilingual Benchmark Dataset for Bridging Fabricated Claims with Humorous Content 3 upvotes, #43 of 2025-03-21
- AIMI: Leveraging Future Knowledge and Personalization in Sparse Event Forecasting for Treatment Adherence 1 upvotes, #48 of 2025-03-21
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.