Daily Papers of 2025-03-27

  1. Qwen2.5-Omni Technical Report 113 upvotes, #1 of 2025-03-27
  2. Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy 47 upvotes, #2 of 2025-03-27
  3. Wan: Open and Advanced Large-Scale Video Generative Models 44 upvotes, #3 of 2025-03-27
  4. Gemma 3 Technical Report 40 upvotes, #4 of 2025-03-27
  5. Open Deep Search: Democratizing Search with Open-source Reasoning Agents 39 upvotes, #5 of 2025-03-27
  6. LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning? 31 upvotes, #6 of 2025-03-27
  7. Gemini Robotics: Bringing AI into the Physical World 21 upvotes, #7 of 2025-03-27
  8. Unconditional Priors Matter! Improving Conditional Generation of Fine-Tuned Diffusion Models 21 upvotes, #7 of 2025-03-27
  9. GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers 15 upvotes, #9 of 2025-03-27
  10. BizGen: Advancing Article-level Visual Text Rendering for Infographics Generation 13 upvotes, #10 of 2025-03-27
  11. AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset 10 upvotes, #11 of 2025-03-27
  12. LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation 9 upvotes, #12 of 2025-03-27
  13. MCTS-RAG: Enhancing Retrieval-Augmented Generation with Monte Carlo Tree Search 9 upvotes, #12 of 2025-03-27
  14. Attention IoU: Examining Biases in CelebA using Attention Maps 7 upvotes, #14 of 2025-03-27
  15. ViLBench: A Suite for Vision-Language Process Reward Modeling 7 upvotes, #14 of 2025-03-27
  16. Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging 7 upvotes, #14 of 2025-03-27
  17. Image as an IMU: Estimating Camera Motion from a Single Motion-Blurred Image 6 upvotes, #17 of 2025-03-27
  18. ADS-Edit: A Multimodal Knowledge Editing Dataset for Autonomous Driving Systems 6 upvotes, #17 of 2025-03-27
  19. Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs 5 upvotes, #19 of 2025-03-27
  20. Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models 4 upvotes, #20 of 2025-03-27
  21. Trajectory Balance with Asynchrony: Decoupling Exploration and Learning for Fast, Scalable LLM Post-Training 3 upvotes, #21 of 2025-03-27
  22. Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals 3 upvotes, #21 of 2025-03-27
  23. DINeMo: Learning Neural Mesh Models with no 3D Annotations 3 upvotes, #21 of 2025-03-27
  24. UniHDSA: A Unified Relation Prediction Approach for Hierarchical Document Structure Analysis 2 upvotes, #24 of 2025-03-27
  25. RecTable: Fast Modeling Tabular Data with Rectified Flow 2 upvotes, #24 of 2025-03-27
  26. RONA: Pragmatically Diverse Image Captioning with Coherence Relations 1 upvotes, #26 of 2025-03-27
  27. PathoHR: Breast Cancer Survival Prediction on High-Resolution Pathological Images 1 upvotes, #26 of 2025-03-27

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.