Daily Papers of 2025-06-12
- Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models 120 upvotes, #1 of 2025-06-12
- Seedance 1.0: Exploring the Boundaries of Video Generation Models 87 upvotes, #2 of 2025-06-12
- Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge Generation 55 upvotes, #3 of 2025-06-12
- ComfyUI-R1: Exploring Reasoning Models for Workflow Generation 49 upvotes, #4 of 2025-06-12
- Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation 46 upvotes, #5 of 2025-06-12
- PlayerOne: Egocentric World Simulator 32 upvotes, #6 of 2025-06-12
- Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation 29 upvotes, #7 of 2025-06-12
- SeerAttention-R: Sparse Attention Adaptation for Long Reasoning 24 upvotes, #8 of 2025-06-12
- CoRT: Code-integrated Reasoning within Thinking 18 upvotes, #9 of 2025-06-12
- SWE-Flow: Synthesizing Software Engineering Data in a Test-Driven Manner 17 upvotes, #10 of 2025-06-12
- Give Me FP32 or Give Me Death? Challenges and Solutions for Reproducible Reasoning 16 upvotes, #11 of 2025-06-12
- Time to Talk: LLM Agents for Asynchronous Group Communication in Mafia Games 14 upvotes, #12 of 2025-06-12
- InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions 13 upvotes, #13 of 2025-06-12
- Vision Matters: Simple Visual Perturbations Can Boost Multimodal Math Reasoning 10 upvotes, #14 of 2025-06-12
- SAFE: Multitask Failure Detection for Vision-Language-Action Models 9 upvotes, #15 of 2025-06-12
- Hidden in plain sight: VLMs overlook their visual representations 8 upvotes, #16 of 2025-06-12
- Efficient Part-level 3D Object Generation via Dual Volume Packing 8 upvotes, #16 of 2025-06-12
- UFM: A Simple Path towards Unified Dense Correspondence with Flow 6 upvotes, #18 of 2025-06-12
- Can Vision Language Models Infer Human Gaze Direction? A Controlled Study 4 upvotes, #19 of 2025-06-12
- Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models 3 upvotes, #20 of 2025-06-12
- Reparameterized LLM Training via Orthogonal Equivalence Transformation 2 upvotes, #21 of 2025-06-12
- MIRAGE: Multimodal foundation model and benchmark for comprehensive retinal OCT image analysis 2 upvotes, #21 of 2025-06-12
- Query-Level Uncertainty in Large Language Models 2 upvotes, #21 of 2025-06-12
- When to Trust Context: Self-Reflective Debates for Context Reliability 1 upvotes, #24 of 2025-06-12
- Branched Schrödinger Bridge Matching 1 upvotes, #24 of 2025-06-12
- Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy 1 upvotes, #24 of 2025-06-12
- A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy 0 upvotes, #27 of 2025-06-12
- TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games 0 upvotes, #27 of 2025-06-12
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.