Daily Papers of 2025-06-10
- Reinforcement Pre-Training 209 upvotes, #1 of 2025-06-10
- Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning 100 upvotes, #2 of 2025-06-10
- MiniCPM4: Ultra-Efficient LLMs on End Devices 78 upvotes, #3 of 2025-06-10
- Saffron-1: Towards an Inference Scaling Paradigm for LLM Safety Assurance 69 upvotes, #4 of 2025-06-10
- OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation 38 upvotes, #5 of 2025-06-10
- SpatialLM: Training Large Language Models for Structured Indoor Modeling 36 upvotes, #6 of 2025-06-10
- Astra: Toward General-Purpose Mobile Robots via Hierarchical Multimodal Learning 28 upvotes, #7 of 2025-06-10
- Image Reconstruction as a Tool for Feature Analysis 28 upvotes, #7 of 2025-06-10
- Pre-trained Large Language Models Learn Hidden Markov Models In-context 21 upvotes, #9 of 2025-06-10
- BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation 18 upvotes, #10 of 2025-06-10
- Through the Valley: Path to Effective Long CoT Training for Small Language Models 18 upvotes, #10 of 2025-06-10
- Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers 17 upvotes, #12 of 2025-06-10
- Vision Transformers Don't Need Trained Registers 16 upvotes, #13 of 2025-06-10
- Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation 14 upvotes, #14 of 2025-06-10
- Bootstrapping World Models from Dynamics Models in Multimodal Foundation Models 13 upvotes, #15 of 2025-06-10
- GTR-CoT: Graph Traversal as Visual Chain of Thought for Molecular Structure Recognition 13 upvotes, #15 of 2025-06-10
- Play to Generalize: Learning to Reason Through Game Play 13 upvotes, #15 of 2025-06-10
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity 12 upvotes, #18 of 2025-06-10
- ConfQA: Answer Only If You Are Confident 10 upvotes, #19 of 2025-06-10
- ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists 9 upvotes, #20 of 2025-06-10
- CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models 9 upvotes, #20 of 2025-06-10
- Model Immunization from a Condition Number Perspective 8 upvotes, #22 of 2025-06-10
- SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs 7 upvotes, #23 of 2025-06-10
- Well Begun is Half Done: Low-resource Preference Alignment by Weak-to-Strong Decoding 7 upvotes, #23 of 2025-06-10
- Dreamland: Controllable World Creation with Simulator and Generative Models 7 upvotes, #23 of 2025-06-10
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior 7 upvotes, #23 of 2025-06-10
- Agents of Change: Self-Evolving LLM Agents for Strategic Planning 6 upvotes, #27 of 2025-06-10
- SAFEFLOW: A Principled Protocol for Trustworthy and Transactional Autonomous Agent Systems 6 upvotes, #27 of 2025-06-10
- Cartridges: Lightweight and general-purpose long context representations via self-study 5 upvotes, #29 of 2025-06-10
- What Is Seen Cannot Be Unseen: The Disruptive Effect of Knowledge Conflict on Large Language Models 5 upvotes, #29 of 2025-06-10
- Dynamic View Synthesis as an Inverse Problem 5 upvotes, #29 of 2025-06-10
- Self-Adapting Improvement Loops for Robotic Learning 4 upvotes, #32 of 2025-06-10
- Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs 4 upvotes, #32 of 2025-06-10
- PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement 4 upvotes, #32 of 2025-06-10
- CyberV: Cybernetics for Test-time Scaling in Video Understanding 4 upvotes, #32 of 2025-06-10
- τ^2-Bench: Evaluating Conversational Agents in a Dual-Control Environment 4 upvotes, #32 of 2025-06-10
- Hidden in Plain Sight: Probing Implicit Reasoning in Multimodal Language Models 3 upvotes, #37 of 2025-06-10
- NetPress: Dynamically Generated LLM Benchmarks for Network Applications 3 upvotes, #37 of 2025-06-10
- GeometryZero: Improving Geometry Solving for LLM with Group Contrastive Policy Optimization 3 upvotes, #37 of 2025-06-10
- Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questions 3 upvotes, #37 of 2025-06-10
- Improving large language models with concept-aware fine-tuning 3 upvotes, #37 of 2025-06-10
- EVOREFUSE: Evolutionary Prompt Optimization for Evaluation and Mitigation of LLM Over-Refusal to Pseudo-Malicious Instructions 2 upvotes, #42 of 2025-06-10
- Robust Preference Optimization via Dynamic Target Margins 2 upvotes, #42 of 2025-06-10
- MegaHan97K: A Large-Scale Dataset for Mega-Category Chinese Character Recognition with over 97K Categories 2 upvotes, #42 of 2025-06-10
- Proactive Assistant Dialogue Generation from Streaming Egocentric Videos 2 upvotes, #42 of 2025-06-10
- Training-Free Tokenizer Transplantation via Orthogonal Matching Pursuit 2 upvotes, #42 of 2025-06-10
- Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering 2 upvotes, #42 of 2025-06-10
- Evaluating LLMs Robustness in Less Resourced Languages with Proxy Models 2 upvotes, #42 of 2025-06-10
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.