Daily Papers of 2025-06-02

  1. ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models 118 upvotes, #1 of 2025-06-02
  2. AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time 89 upvotes, #2 of 2025-06-02
  3. Time Blindness: Why Video-Language Models Can't See What Humans Can? 75 upvotes, #3 of 2025-06-02
  4. Large Language Models for Data Synthesis 48 upvotes, #4 of 2025-06-02
  5. HardTests: Synthesizing High-Quality Test Cases for LLM Coding 42 upvotes, #5 of 2025-06-02
  6. Don't Look Only Once: Towards Multimodal Interactive Reasoning with Selective Visual Revisitation 36 upvotes, #6 of 2025-06-02
  7. ViStoryBench: Comprehensive Benchmark Suite for Story Visualization 31 upvotes, #7 of 2025-06-02
  8. DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models 25 upvotes, #8 of 2025-06-02
  9. EXP-Bench: Can AI Conduct AI Research Experiments? 23 upvotes, #9 of 2025-06-02
  10. Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents 22 upvotes, #10 of 2025-06-02
  11. CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated Objects 21 upvotes, #11 of 2025-06-02
  12. Vision Language Models are Biased 20 upvotes, #12 of 2025-06-02
  13. MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning 20 upvotes, #12 of 2025-06-02
  14. EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge 17 upvotes, #14 of 2025-06-02
  15. MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs 17 upvotes, #14 of 2025-06-02
  16. UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation 15 upvotes, #16 of 2025-06-02
  17. More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models 14 upvotes, #17 of 2025-06-02
  18. CLaSp: In-Context Layer Skip for Self-Speculative Decoding 13 upvotes, #18 of 2025-06-02
  19. Large Language Models are Locally Linear Mappings 13 upvotes, #18 of 2025-06-02
  20. EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering 12 upvotes, #20 of 2025-06-02
  21. Fork-Merge Decoding: Enhancing Multimodal Understanding in Audio-Visual Large Language Models 10 upvotes, #21 of 2025-06-02
  22. ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL 10 upvotes, #21 of 2025-06-02
  23. DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation 9 upvotes, #23 of 2025-06-02
  24. Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning 9 upvotes, #23 of 2025-06-02
  25. Evaluating and Steering Modality Preferences in Multimodal Large Language Model 8 upvotes, #25 of 2025-06-02
  26. Role-Playing Evaluation for Large Language Models 7 upvotes, #26 of 2025-06-02
  27. ChARM: Character-based Act-adaptive Reward Modeling for Advanced Role-Playing Language Agents 7 upvotes, #26 of 2025-06-02
  28. Enabling Flexible Multi-LLM Integration for Scalable Knowledge Aggregation 5 upvotes, #28 of 2025-06-02
  29. Point-MoE: Towards Cross-Domain Generalization in 3D Semantic Segmentation via Mixture-of-Experts 5 upvotes, #28 of 2025-06-02
  30. Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows 5 upvotes, #28 of 2025-06-02
  31. un^2CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP 5 upvotes, #28 of 2025-06-02
  32. Harnessing Large Language Models for Scientific Novelty Detection 5 upvotes, #28 of 2025-06-02
  33. SiLVR: A Simple Language-based Video Reasoning Framework 5 upvotes, #28 of 2025-06-02
  34. Revisiting Bi-Linear State Transitions in Recurrent Neural Networks 4 upvotes, #34 of 2025-06-02
  35. Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks 3 upvotes, #35 of 2025-06-02
  36. GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss Training 3 upvotes, #35 of 2025-06-02
  37. TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis 3 upvotes, #35 of 2025-06-02
  38. LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements Generation 2 upvotes, #38 of 2025-06-02
  39. OMNIGUARD: An Efficient Approach for AI Safety Moderation Across Modalities 2 upvotes, #38 of 2025-06-02
  40. The Automated but Risky Game: Modeling Agent-to-Agent Negotiations and Transactions in Consumer Markets 2 upvotes, #38 of 2025-06-02
  41. The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It 1 upvotes, #41 of 2025-06-02
  42. Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings 1 upvotes, #41 of 2025-06-02

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.