Daily Papers of 2025-04-15

  1. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models 239 upvotes, #1 of 2025-04-15
  2. PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters 119 upvotes, #2 of 2025-04-15
  3. Have we unified image generation and understanding yet? An empirical study of GPT-4o's image generation ability 47 upvotes, #3 of 2025-04-15
  4. VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning 42 upvotes, #4 of 2025-04-15
  5. FUSION: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding 38 upvotes, #5 of 2025-04-15
  6. Iterative Self-Training for Code Generation via Reinforced Re-Ranking 34 upvotes, #6 of 2025-04-15
  7. Mavors: Multi-granularity Video Representation for Multimodal Large Language Model 30 upvotes, #7 of 2025-04-15
  8. AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories 27 upvotes, #8 of 2025-04-15
  9. S1-Bench: A Simple Benchmark for Evaluating System 1 Thinking Capability of Large Reasoning Models 21 upvotes, #9 of 2025-04-15
  10. DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training 19 upvotes, #10 of 2025-04-15
  11. Breaking the Data Barrier -- Building GUI Agents Through Task Generalization 17 upvotes, #11 of 2025-04-15
  12. TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning 15 upvotes, #12 of 2025-04-15
  13. SocioVerse: A World Model for Social Simulation Powered by LLM Agents and A Pool of 10 Million Real-World Users 15 upvotes, #12 of 2025-04-15
  14. MIEB: Massive Image Embedding Benchmark 15 upvotes, #12 of 2025-04-15
  15. Executable Functional Abstractions: Inferring Generative Programs for Advanced Math Problems 13 upvotes, #15 of 2025-04-15
  16. VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search 11 upvotes, #16 of 2025-04-15
  17. The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search 10 upvotes, #17 of 2025-04-15
  18. Reasoning Models Can Be Effective Without Thinking 10 upvotes, #17 of 2025-04-15
  19. M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models 10 upvotes, #17 of 2025-04-15
  20. LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models 8 upvotes, #20 of 2025-04-15
  21. How new data permeates LLM knowledge and how to dilute it 7 upvotes, #21 of 2025-04-15
  22. EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety 7 upvotes, #21 of 2025-04-15
  23. 3D CoCa: Contrastive Learners are 3D Captioners 5 upvotes, #23 of 2025-04-15
  24. MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models 4 upvotes, #24 of 2025-04-15
  25. DeepSeek vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization? 4 upvotes, #24 of 2025-04-15
  26. LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models 4 upvotes, #24 of 2025-04-15
  27. MCP Safety Audit: LLMs with the Model Context Protocol Allow Major Security Exploits 3 upvotes, #27 of 2025-04-15
  28. DiffuMural: Restoring Dunhuang Murals with Multi-scale Diffusion 1 upvotes, #28 of 2025-04-15

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.