Daily Papers of 2025-06-10

  1. Reinforcement Pre-Training 209 upvotes, #1 of 2025-06-10
  2. Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning 100 upvotes, #2 of 2025-06-10
  3. MiniCPM4: Ultra-Efficient LLMs on End Devices 78 upvotes, #3 of 2025-06-10
  4. Saffron-1: Towards an Inference Scaling Paradigm for LLM Safety Assurance 69 upvotes, #4 of 2025-06-10
  5. OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation 38 upvotes, #5 of 2025-06-10
  6. SpatialLM: Training Large Language Models for Structured Indoor Modeling 36 upvotes, #6 of 2025-06-10
  7. Astra: Toward General-Purpose Mobile Robots via Hierarchical Multimodal Learning 28 upvotes, #7 of 2025-06-10
  8. Image Reconstruction as a Tool for Feature Analysis 28 upvotes, #7 of 2025-06-10
  9. Pre-trained Large Language Models Learn Hidden Markov Models In-context 21 upvotes, #9 of 2025-06-10
  10. BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation 18 upvotes, #10 of 2025-06-10
  11. Through the Valley: Path to Effective Long CoT Training for Small Language Models 18 upvotes, #10 of 2025-06-10
  12. Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers 17 upvotes, #12 of 2025-06-10
  13. Vision Transformers Don't Need Trained Registers 16 upvotes, #13 of 2025-06-10
  14. Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation 14 upvotes, #14 of 2025-06-10
  15. Bootstrapping World Models from Dynamics Models in Multimodal Foundation Models 13 upvotes, #15 of 2025-06-10
  16. GTR-CoT: Graph Traversal as Visual Chain of Thought for Molecular Structure Recognition 13 upvotes, #15 of 2025-06-10
  17. Play to Generalize: Learning to Reason Through Game Play 13 upvotes, #15 of 2025-06-10
  18. The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity 12 upvotes, #18 of 2025-06-10
  19. ConfQA: Answer Only If You Are Confident 10 upvotes, #19 of 2025-06-10
  20. ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists 9 upvotes, #20 of 2025-06-10
  21. CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models 9 upvotes, #20 of 2025-06-10
  22. Model Immunization from a Condition Number Perspective 8 upvotes, #22 of 2025-06-10
  23. SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs 7 upvotes, #23 of 2025-06-10
  24. Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding 7 upvotes, #23 of 2025-06-10
  25. Dreamland: Controllable World Creation with Simulator and Generative Models 7 upvotes, #23 of 2025-06-10
  26. GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior 7 upvotes, #23 of 2025-06-10
  27. Agents of Change: Self-Evolving LLM Agents for Strategic Planning 6 upvotes, #27 of 2025-06-10
  28. SAFEFLOW: A Principled Protocol for Trustworthy and Transactional Autonomous Agent Systems 6 upvotes, #27 of 2025-06-10
  29. Cartridges: Lightweight and general-purpose long context representations via self-study 5 upvotes, #29 of 2025-06-10
  30. What Is Seen Cannot Be Unseen: The Disruptive Effect of Knowledge Conflict on Large Language Models 5 upvotes, #29 of 2025-06-10
  31. Dynamic View Synthesis as an Inverse Problem 5 upvotes, #29 of 2025-06-10
  32. Self-Adapting Improvement Loops for Robotic Learning 4 upvotes, #32 of 2025-06-10
  33. Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs 4 upvotes, #32 of 2025-06-10
  34. PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement 4 upvotes, #32 of 2025-06-10
  35. CyberV: Cybernetics for Test-time Scaling in Video Understanding 4 upvotes, #32 of 2025-06-10
  36. τ^2-Bench: Evaluating Conversational Agents in a Dual-Control Environment 4 upvotes, #32 of 2025-06-10
  37. Hidden in Plain Sight: Probing Implicit Reasoning in Multimodal Language Models 3 upvotes, #37 of 2025-06-10
  38. NetPress: Dynamically Generated LLM Benchmarks for Network Applications 3 upvotes, #37 of 2025-06-10
  39. GeometryZero: Improving Geometry Solving for LLM with Group Contrastive Policy Optimization 3 upvotes, #37 of 2025-06-10
  40. Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions 3 upvotes, #37 of 2025-06-10
  41. Improving large language models with concept-aware fine-tuning 3 upvotes, #37 of 2025-06-10
  42. EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions 2 upvotes, #42 of 2025-06-10
  43. Robust Preference Optimization via Dynamic Target Margins 2 upvotes, #42 of 2025-06-10
  44. MegaHan97K: A Large-Scale Dataset for Mega-Category Chinese Character Recognition with over 97K Categories 2 upvotes, #42 of 2025-06-10
  45. Proactive Assistant Dialogue Generation from Streaming Egocentric Videos 2 upvotes, #42 of 2025-06-10
  46. Training-Free Tokenizer Transplantation via Orthogonal Matching Pursuit 2 upvotes, #42 of 2025-06-10
  47. Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering 2 upvotes, #42 of 2025-06-10
  48. Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models 2 upvotes, #42 of 2025-06-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.