Daily Papers of 2025-09-24

  1. Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR 121 upvotes, #1 of 2025-09-24
  2. Reinforcement Learning on Pre-Training Data 61 upvotes, #2 of 2025-09-24
  3. Do You Need Proprioceptive States in Visuomotor Policies? 49 upvotes, #3 of 2025-09-24
  4. MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe 46 upvotes, #4 of 2025-09-24
  5. SWE-QA: Can Language Models Answer Repository-level Code Questions? 34 upvotes, #5 of 2025-09-24
  6. How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective 28 upvotes, #6 of 2025-09-24
  7. MAPO: Mixed Advantage Policy Optimization 25 upvotes, #7 of 2025-09-24
  8. VolSplat: Rethinking Feed-Forward 3D Gaussian Splatting with Voxel-Aligned Prediction 23 upvotes, #8 of 2025-09-24
  9. What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT 22 upvotes, #9 of 2025-09-24
  10. Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation 21 upvotes, #10 of 2025-09-24
  11. Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation 21 upvotes, #10 of 2025-09-24
  12. Soft Tokens, Hard Truths 15 upvotes, #12 of 2025-09-24
  13. Large Language Models Discriminate Against Speakers of German Dialects 7 upvotes, #13 of 2025-09-24
  14. HyRF: Hybrid Radiance Fields for Memory-efficient and High-quality Novel View Synthesis 7 upvotes, #13 of 2025-09-24
  15. CAR-Flow: Condition-Aware Reparameterization Aligns Source and Target for Better Flow Matching 4 upvotes, #15 of 2025-09-24
  16. OpenGVL - Benchmarking Visual Temporal Progress for Data Curation 3 upvotes, #16 of 2025-09-24
  17. CommonForms: A Large, Diverse Dataset for Form Field Detection 2 upvotes, #17 of 2025-09-24
  18. Better Late Than Never: Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation 2 upvotes, #17 of 2025-09-24
  19. GeoSVR: Taming Sparse Voxels for Geometrically Accurate Surface Reconstruction 2 upvotes, #17 of 2025-09-24
  20. VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction 2 upvotes, #17 of 2025-09-24
  21. RadEval: A framework for radiology text evaluation 1 upvotes, #21 of 2025-09-24
  22. PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies 1 upvotes, #21 of 2025-09-24
  23. Zero-Shot Multi-Spectral Learning: Reimagining a Generalist Multimodal Gemini 2.5 Model for Remote Sensing Applications 1 upvotes, #21 of 2025-09-24
  24. DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture 2 upvotes, #24 of 2025-09-24

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.