Daily Papers of 2025-03-27
- Qwen2.5-Omni Technical Report 113 upvotes, #1 of 2025-03-27
- Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy 47 upvotes, #2 of 2025-03-27
- Wan: Open and Advanced Large-Scale Video Generative Models 44 upvotes, #3 of 2025-03-27
- Gemma 3 Technical Report 40 upvotes, #4 of 2025-03-27
- Open Deep Search: Democratizing Search with Open-source Reasoning Agents 39 upvotes, #5 of 2025-03-27
- LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning? 31 upvotes, #6 of 2025-03-27
- Gemini Robotics: Bringing AI into the Physical World 21 upvotes, #7 of 2025-03-27
- Unconditional Priors Matter! Improving Conditional Generation of Fine-Tuned Diffusion Models 21 upvotes, #7 of 2025-03-27
- GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers 15 upvotes, #9 of 2025-03-27
- BizGen: Advancing Article-level Visual Text Rendering for Infographics Generation 13 upvotes, #10 of 2025-03-27
- AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset 10 upvotes, #11 of 2025-03-27
- LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation 9 upvotes, #12 of 2025-03-27
- MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search 9 upvotes, #12 of 2025-03-27
- Attention IoU: Examining Biases in CelebA using Attention Maps 7 upvotes, #14 of 2025-03-27
- ViLBench: A Suite for Vision-Language Process Reward Modeling 7 upvotes, #14 of 2025-03-27
- Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging 7 upvotes, #14 of 2025-03-27
- Image as an IMU: Estimating Camera Motion from a Single Motion-Blurred Image 6 upvotes, #17 of 2025-03-27
- ADS-Edit: A Multimodal Knowledge Editing Dataset for Autonomous Driving Systems 6 upvotes, #17 of 2025-03-27
- Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs 5 upvotes, #19 of 2025-03-27
- Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models 4 upvotes, #20 of 2025-03-27
- Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training 3 upvotes, #21 of 2025-03-27
- Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals 3 upvotes, #21 of 2025-03-27
- DINeMo: Learning Neural Mesh Models with no 3D Annotations 3 upvotes, #21 of 2025-03-27
- UniHDSA: A Unified Relation Prediction Approach for Hierarchical Document Structure Analysis 2 upvotes, #24 of 2025-03-27
- RecTable: Fast Modeling Tabular Data with Rectified Flow 2 upvotes, #24 of 2025-03-27
- RONA: Pragmatically Diverse Image Captioning with Coherence Relations 1 upvotes, #26 of 2025-03-27
- PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images 1 upvotes, #26 of 2025-03-27
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.