Daily Papers of 2025-01-31

  1. GuardReasoner: Towards Reasoning-based LLM Safeguards 79 upvotes, #1 of 2025-01-31
  2. Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs 51 upvotes, #2 of 2025-01-31
  3. Streaming DiLoCo with overlapping communication: Towards a Distributed Free Lunch 25 upvotes, #3 of 2025-01-31
  4. Large Language Models Think Too Fast To Explore Effectively 22 upvotes, #4 of 2025-01-31
  5. o3-mini vs DeepSeek-R1: Which One is Safer? 21 upvotes, #5 of 2025-01-31
  6. MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding 19 upvotes, #6 of 2025-01-31
  7. PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding 17 upvotes, #7 of 2025-01-31
  8. WILDCHAT-50M: A Deep Dive Into the Role of Synthetic Data in Post-Training 17 upvotes, #7 of 2025-01-31
  9. SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer 16 upvotes, #9 of 2025-01-31
  10. CowPilot: A Framework for Autonomous and Human-Agent Collaborative Web Navigation 6 upvotes, #10 of 2025-01-31

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.