Daily Papers of 2025-06-02
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models 118 upvotes, #1 of 2025-06-02
- AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time 89 upvotes, #2 of 2025-06-02
- Time Blindness: Why Video-Language Models Can't See What Humans Can? 75 upvotes, #3 of 2025-06-02
- Large Language Models for Data Synthesis 48 upvotes, #4 of 2025-06-02
- HardTests: Synthesizing High-Quality Test Cases for LLM Coding 42 upvotes, #5 of 2025-06-02
- Don't Look Only Once: Towards Multimodal Interactive Reasoning with Selective Visual Revisitation 36 upvotes, #6 of 2025-06-02
- ViStoryBench: Comprehensive Benchmark Suite for Story Visualization 31 upvotes, #7 of 2025-06-02
- DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models 25 upvotes, #8 of 2025-06-02
- EXP-Bench: Can AI Conduct AI Research Experiments? 23 upvotes, #9 of 2025-06-02
- Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents 22 upvotes, #10 of 2025-06-02
- CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated Objects 21 upvotes, #11 of 2025-06-02
- Vision Language Models are Biased 20 upvotes, #12 of 2025-06-02
- MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning 20 upvotes, #12 of 2025-06-02
- EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-Judge 17 upvotes, #14 of 2025-06-02
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMs 17 upvotes, #14 of 2025-06-02
- UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation 15 upvotes, #16 of 2025-06-02
- More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models 14 upvotes, #17 of 2025-06-02
- CLaSp: In-Context Layer Skip for Self-Speculative Decoding 13 upvotes, #18 of 2025-06-02
- Large Language Models are Locally Linear Mappings 13 upvotes, #18 of 2025-06-02
- EasyText: Controllable Diffusion Transformer for Multilingual Text Rendering 12 upvotes, #20 of 2025-06-02
- Fork-Merge Decoding: Enhancing Multimodal Understanding in Audio-Visual Large Language Models 10 upvotes, #21 of 2025-06-02
- ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL 10 upvotes, #21 of 2025-06-02
- DexUMI: Using Human Hand as the Universal Manipulation Interface for Dexterous Manipulation 9 upvotes, #23 of 2025-06-02
- Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning 9 upvotes, #23 of 2025-06-02
- Evaluating and Steering Modality Preferences in Multimodal Large Language Model 8 upvotes, #25 of 2025-06-02
- Role-Playing Evaluation for Large Language Models 7 upvotes, #26 of 2025-06-02
- ChARM: Character-based Act-adaptive Reward Modeling for Advanced Role-Playing Language Agents 7 upvotes, #26 of 2025-06-02
- Enabling Flexible Multi-LLM Integration for Scalable Knowledge Aggregation 5 upvotes, #28 of 2025-06-02
- Point-MoE: Towards Cross-Domain Generalization in 3D Semantic Segmentation via Mixture-of-Experts 5 upvotes, #28 of 2025-06-02
- Fine-Tune an SLM or Prompt an LLM? The Case of Generating Low-Code Workflows 5 upvotes, #28 of 2025-06-02
- un^2CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP 5 upvotes, #28 of 2025-06-02
- Harnessing Large Language Models for Scientific Novelty Detection 5 upvotes, #28 of 2025-06-02
- SiLVR: A Simple Language-based Video Reasoning Framework 5 upvotes, #28 of 2025-06-02
- Revisiting Bi-Linear State Transitions in Recurrent Neural Networks 4 upvotes, #34 of 2025-06-02
- Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks 3 upvotes, #35 of 2025-06-02
- GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss Training 3 upvotes, #35 of 2025-06-02
- TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis 3 upvotes, #35 of 2025-06-02
- LegalSearchLM: Rethinking Legal Case Retrieval as Legal Elements Generation 2 upvotes, #38 of 2025-06-02
- OMNIGUARD: An Efficient Approach for AI Safety Moderation Across Modalities 2 upvotes, #38 of 2025-06-02
- The Automated but Risky Game: Modeling Agent-to-Agent Negotiations and Transactions in Consumer Markets 2 upvotes, #38 of 2025-06-02
- The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It 1 upvotes, #41 of 2025-06-02
- Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings 1 upvotes, #41 of 2025-06-02
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.