Daily Papers of 2025-05-15

  1. BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset 80 upvotes, #1 of 2025-05-15
  2. Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures 55 upvotes, #2 of 2025-05-15
  3. DeCLIP: Decoupled Learning for Open-Vocabulary Dense Perception 42 upvotes, #3 of 2025-05-15
  4. MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning 40 upvotes, #4 of 2025-05-15
  5. LightLab: Controlling Light Sources in Images with Diffusion Models 27 upvotes, #5 of 2025-05-15
  6. Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis 22 upvotes, #6 of 2025-05-15
  7. UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations 13 upvotes, #7 of 2025-05-15
  8. CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image 12 upvotes, #8 of 2025-05-15
  9. SweRank: Software Issue Localization with Code Ranking 8 upvotes, #9 of 2025-05-15
  10. Omni-R1: Do You Really Need Audio to Fine-Tune Your Audio LLM? 8 upvotes, #9 of 2025-05-15
  11. WavReward: Spoken Dialogue Models With Generalist Reward Evaluators 8 upvotes, #9 of 2025-05-15
  12. VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models 4 upvotes, #12 of 2025-05-15
  13. Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA 3 upvotes, #13 of 2025-05-15
  14. DetReIDX: A Stress-Test Dataset for Real-World UAV-Based Person Recognition 2 upvotes, #14 of 2025-05-15
  15. Steepest Descent Density Control for Compact 3D Gaussian Splatting 1 upvotes, #15 of 2025-05-15
  16. Visually Interpretable Subtask Reasoning for Visual Question Answering 1 upvotes, #15 of 2025-05-15
  17. Behind Maya: Building a Multilingual Vision Language Model 1 upvotes, #15 of 2025-05-15

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.