Daily Papers of 2025-09-10
- Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing 559 upvotes, #1 of 2025-09-10
- Parallel-R1: Towards Parallel Thinking via Reinforcement Learning 95 upvotes, #2 of 2025-09-10
- Visual Representation Alignment for Multimodal Large Language Models 77 upvotes, #3 of 2025-09-10
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search 59 upvotes, #4 of 2025-09-10
- Reconstruction Alignment Improves Unified Multimodal Models 38 upvotes, #5 of 2025-09-10
- F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions 30 upvotes, #6 of 2025-09-10
- UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward 29 upvotes, #7 of 2025-09-10
- Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning 26 upvotes, #8 of 2025-09-10
- Language Self-Play For Data-Free Training 26 upvotes, #8 of 2025-09-10
- Curia: A Multi-Modal Foundation Model for Radiology 20 upvotes, #10 of 2025-09-10
- Staying in the Sweet Spot: Responsive Reasoning Evolution via Capability-Adaptive Hint Scaffolding 20 upvotes, #10 of 2025-09-10
- Causal Attention with Lookahead Keys 20 upvotes, #10 of 2025-09-10
- Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference 15 upvotes, #13 of 2025-09-10
- SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge 11 upvotes, #14 of 2025-09-10
- Q-Sched: Pushing the Boundaries of Few-Step Diffusion Models with Quantization-Aware Scheduling 7 upvotes, #15 of 2025-09-10
- ΔL Normalization: Rethink Loss Aggregation in RLVR 6 upvotes, #16 of 2025-09-10
- From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers 4 upvotes, #17 of 2025-09-10
- Benchmarking Information Retrieval Models on Complex Retrieval Tasks 3 upvotes, #18 of 2025-09-10
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.