Daily Papers of 2025-05-23

  1. NovelSeek: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification 110 upvotes, #1 of 2025-05-23
  2. Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models 58 upvotes, #2 of 2025-05-23
  3. Tool-Star: Empowering LLM-Brained Multi-Tool Reasoner via Reinforcement Learning 53 upvotes, #3 of 2025-05-23
  4. Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning 48 upvotes, #4 of 2025-05-23
  5. KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models 41 upvotes, #5 of 2025-05-23
  6. QuickVideo: Real-Time Long Video Understanding with System Algorithm Co-Design 38 upvotes, #6 of 2025-05-23
  7. Scaling Diffusion Transformers Efficiently via μP 31 upvotes, #7 of 2025-05-23
  8. LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning 29 upvotes, #8 of 2025-05-23
  9. AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning 27 upvotes, #9 of 2025-05-23
  10. GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning 26 upvotes, #10 of 2025-05-23
  11. Risk-Averse Reinforcement Learning with Itakura-Saito Loss 24 upvotes, #11 of 2025-05-23
  12. Let LLMs Break Free from Overthinking via Self-Braking Tuning 23 upvotes, #12 of 2025-05-23
  13. Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning 23 upvotes, #12 of 2025-05-23
  14. Understanding Generative AI Capabilities in Everyday Image Editing Tasks 23 upvotes, #12 of 2025-05-23
  15. Fixing Data That Hurts Performance: Cascading LLMs to Relabel Hard Negatives for Robust Information Retrieval 22 upvotes, #15 of 2025-05-23
  16. Training-Free Efficient Video Generation via Dynamic Token Carving 21 upvotes, #16 of 2025-05-23
  17. Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding 20 upvotes, #17 of 2025-05-23
  18. VideoGameQA-Bench: Evaluating Vision-Language Models for Video Game Quality Assurance 19 upvotes, #18 of 2025-05-23
  19. WebAgent-R1: Training Web Agents via End-to-End Multi-Turn Reinforcement Learning 18 upvotes, #19 of 2025-05-23
  20. Backdoor Cleaning without External Guidance in MLLM Fine-tuning 16 upvotes, #20 of 2025-05-23
  21. SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward 14 upvotes, #21 of 2025-05-23
  22. TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning 12 upvotes, #22 of 2025-05-23
  23. GRIT: Teaching MLLMs to Think with Images 12 upvotes, #22 of 2025-05-23
  24. SpatialScore: Towards Unified Evaluation for Multimodal Spatial Understanding 12 upvotes, #22 of 2025-05-23
  25. LaViDa: A Large Diffusion Language Model for Multimodal Understanding 11 upvotes, #25 of 2025-05-23
  26. Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models 11 upvotes, #25 of 2025-05-23
  27. Reinforcement Learning Finetunes Small Subnetworks in Large Language Models 10 upvotes, #27 of 2025-05-23
  28. OViP: Online Vision-Language Preference Learning 9 upvotes, #28 of 2025-05-23
  29. Training-Free Reasoning and Reflection in MLLMs 9 upvotes, #28 of 2025-05-23
  30. Think-RM: Enabling Long-Horizon Reasoning in Generative Reward Models 8 upvotes, #30 of 2025-05-23
  31. AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios 8 upvotes, #30 of 2025-05-23
  32. Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models 8 upvotes, #30 of 2025-05-23
  33. Training Step-Level Reasoning Verifiers with Formal Verification Tools 7 upvotes, #33 of 2025-05-23
  34. SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning 7 upvotes, #33 of 2025-05-23
  35. VLM-R^3: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought 7 upvotes, #33 of 2025-05-23
  36. RoPECraft: Training-Free Motion Transfer with Trajectory-Guided RoPE Optimization on Diffusion Transformers 6 upvotes, #36 of 2025-05-23
  37. MUG-Eval: A Proxy Evaluation Framework for Multilingual Generation Capabilities in Any Language 5 upvotes, #37 of 2025-05-23
  38. Steering Large Language Models for Machine Translation Personalization 5 upvotes, #37 of 2025-05-23
  39. RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding 4 upvotes, #39 of 2025-05-23
  40. Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets 4 upvotes, #39 of 2025-05-23
  41. How Do Large Vision-Language Models See Text in Image? Unveiling the Distinctive Role of OCR Heads 4 upvotes, #39 of 2025-05-23
  42. When Do LLMs Admit Their Mistakes? Understanding the Role of Model Belief in Retraction 4 upvotes, #39 of 2025-05-23
  43. Let Androids Dream of Electric Sheep: A Human-like Image Implication Understanding and Reasoning Framework 4 upvotes, #39 of 2025-05-23
  44. gen2seg: Generative Models Enable Generalizable Instance Segmentation 3 upvotes, #44 of 2025-05-23
  45. SPhyR: Spatial-Physical Reasoning Benchmark on Material Distribution 3 upvotes, #44 of 2025-05-23
  46. Date Fragments: A Hidden Bottleneck of Tokenization for Temporal Reasoning 3 upvotes, #44 of 2025-05-23
  47. SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information 1 upvotes, #47 of 2025-05-23

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.