Daily Papers of 2025-05-19

  1. Qwen3 Technical Report 152 upvotes, #1 of 2025-05-19
  2. MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and Thoroughly 52 upvotes, #2 of 2025-05-19
  3. GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning 50 upvotes, #3 of 2025-05-19
  4. Visual Planning: Let's Think Only with Images 50 upvotes, #3 of 2025-05-19
  5. Group Think: Multiple Concurrent Reasoning Agents Collaborating at Token Level Granularity 28 upvotes, #5 of 2025-05-19
  6. Simple Semi-supervised Knowledge Distillation from Vision-Language Models via texttt{D}ual-texttt{H}ead texttt{O}ptimization 19 upvotes, #6 of 2025-05-19
  7. Multi-Token Prediction Needs Registers 12 upvotes, #7 of 2025-05-19
  8. Mergenetic: a Simple Evolutionary Model Merging Library 12 upvotes, #7 of 2025-05-19
  9. InstanceGen: Image Generation with Instance-level Instructions 7 upvotes, #9 of 2025-05-19
  10. MPS-Prover: Advancing Stepwise Theorem Proving by Multi-Perspective Search and Data Curation 7 upvotes, #9 of 2025-05-19
  11. Improving Assembly Code Performance with Large Language Models via Reinforcement Learning 7 upvotes, #9 of 2025-05-19
  12. MatTools: Benchmarking Large Language Models for Materials Science Tools 6 upvotes, #12 of 2025-05-19
  13. Scaling Reasoning can Improve Factuality in Large Language Models 6 upvotes, #12 of 2025-05-19
  14. Humans expect rationality and cooperation from LLM opponents in strategic games 4 upvotes, #14 of 2025-05-19
  15. Learning Dense Hand Contact Estimation from Imbalanced Data 3 upvotes, #15 of 2025-05-19
  16. GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing 3 upvotes, #15 of 2025-05-19
  17. From Trade-off to Synergy: A Versatile Symbiotic Watermarking Framework for Large Language Models 2 upvotes, #17 of 2025-05-19
  18. CheXGenBench: A Unified Benchmark For Fidelity, Privacy and Utility of Synthetic Chest Radiographs 2 upvotes, #17 of 2025-05-19
  19. Unifying Segment Anything in Microscopy with Multimodal Large Language Model 2 upvotes, #17 of 2025-05-19
  20. Improving Inference-Time Optimisation for Vocal Effects Style Transfer with a Gaussian Prior 0 upvotes, #20 of 2025-05-19

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.