Daily Papers of 2025-05-22

  1. Web-Shepherd: Advancing PRMs for Reinforcing Web Agents 96 upvotes, #1 of 2025-05-22
  2. MMaDA: Multimodal Large Diffusion Language Models 83 upvotes, #2 of 2025-05-22
  3. Scaling Law for Quantization-Aware Training 67 upvotes, #3 of 2025-05-22
  4. Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective 52 upvotes, #4 of 2025-05-22
  5. UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning 50 upvotes, #5 of 2025-05-22
  6. Efficient Agent Training for Computer Use 41 upvotes, #6 of 2025-05-22
  7. This Time is Different: An Observability Perspective on Time Series Foundation Models 36 upvotes, #7 of 2025-05-22
  8. Learn to Reason Efficiently with Adaptive Length-based Reward Shaping 31 upvotes, #8 of 2025-05-22
  9. Vid2World: Crafting Video Diffusion Models to Interactive World Models 23 upvotes, #9 of 2025-05-22
  10. Constructing a 3D Town from a Single Image 23 upvotes, #9 of 2025-05-22
  11. When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning 22 upvotes, #11 of 2025-05-22
  12. lmgame-Bench: How Good are LLMs at Playing Games? 19 upvotes, #12 of 2025-05-22
  13. VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models 17 upvotes, #13 of 2025-05-22
  14. Learning to Reason via Mixture-of-Thought for Logical Reasoning 17 upvotes, #13 of 2025-05-22
  15. dKV-Cache: The Cache for Diffusion Language Models 16 upvotes, #15 of 2025-05-22
  16. Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs 15 upvotes, #16 of 2025-05-22
  17. IA-T2I: Internet-Augmented Text-to-Image Generation 15 upvotes, #16 of 2025-05-22
  18. RLVR-World: Training World Models with Reinforcement Learning 14 upvotes, #18 of 2025-05-22
  19. How Should We Enhance the Safety of Large Reasoning Models: An Empirical Study 13 upvotes, #19 of 2025-05-22
  20. Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen! 13 upvotes, #19 of 2025-05-22
  21. Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space 13 upvotes, #19 of 2025-05-22
  22. DiCo: Revitalizing ConvNets for Scalable and Efficient Diffusion Modeling 12 upvotes, #22 of 2025-05-22
  23. BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMs 11 upvotes, #23 of 2025-05-22
  24. ConvSearch-R1: Enhancing Query Reformulation for Conversational Search with Reasoning via Reinforcement Learning 10 upvotes, #24 of 2025-05-22
  25. Text Generation Beyond Discrete Token Sampling 9 upvotes, #25 of 2025-05-22
  26. AutoMat: Enabling Automated Crystal Structure Reconstruction from Microscopy via Agentic Tool Use 7 upvotes, #26 of 2025-05-22
  27. Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs 7 upvotes, #26 of 2025-05-22
  28. The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning 6 upvotes, #28 of 2025-05-22
  29. Prior Prompt Engineering for Reinforcement Fine-Tuning 5 upvotes, #29 of 2025-05-22
  30. Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models 5 upvotes, #29 of 2025-05-22
  31. VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL 5 upvotes, #29 of 2025-05-22
  32. BLEUBERI: BLEU is a surprisingly effective reward for instruction following 4 upvotes, #32 of 2025-05-22
  33. HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation 4 upvotes, #32 of 2025-05-22
  34. WebNovelBench: Placing LLM Novelists on the Web Novel Distribution 4 upvotes, #32 of 2025-05-22
  35. RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning 4 upvotes, #32 of 2025-05-22
  36. PiFlow: Principle-aware Scientific Discovery with Multi-Agent Collaboration 4 upvotes, #32 of 2025-05-22
  37. BanditSpec: Adaptive Speculative Decoding via Bandit Algorithms 4 upvotes, #32 of 2025-05-22
  38. Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach 3 upvotes, #38 of 2025-05-22
  39. Streamline Without Sacrifice - Squeeze out Computation Redundancy in LMM 3 upvotes, #38 of 2025-05-22
  40. MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations 2 upvotes, #40 of 2025-05-22
  41. In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties 2 upvotes, #40 of 2025-05-22
  42. Language Specific Knowledge: Do Models Know Better in X than in English? 1 upvotes, #42 of 2025-05-22

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.