Daily Papers of 2025-06-12

  1. Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models 120 upvotes, #1 of 2025-06-12
  2. Seedance 1.0: Exploring the Boundaries of Video Generation Models 87 upvotes, #2 of 2025-06-12
  3. Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge Generation 55 upvotes, #3 of 2025-06-12
  4. ComfyUI-R1: Exploring Reasoning Models for Workflow Generation 49 upvotes, #4 of 2025-06-12
  5. Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation 46 upvotes, #5 of 2025-06-12
  6. PlayerOne: Egocentric World Simulator 32 upvotes, #6 of 2025-06-12
  7. Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation 29 upvotes, #7 of 2025-06-12
  8. SeerAttention-R: Sparse Attention Adaptation for Long Reasoning 24 upvotes, #8 of 2025-06-12
  9. CoRT: Code-integrated Reasoning within Thinking 18 upvotes, #9 of 2025-06-12
  10. SWE-Flow: Synthesizing Software Engineering Data in a Test-Driven Manner 17 upvotes, #10 of 2025-06-12
  11. Give Me FP32 or Give Me Death? Challenges and Solutions for Reproducible Reasoning 16 upvotes, #11 of 2025-06-12
  12. Time to Talk: LLM Agents for Asynchronous Group Communication in Mafia Games 14 upvotes, #12 of 2025-06-12
  13. InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions 13 upvotes, #13 of 2025-06-12
  14. Vision Matters: Simple Visual Perturbations Can Boost Multimodal Math Reasoning 10 upvotes, #14 of 2025-06-12
  15. SAFE: Multitask Failure Detection for Vision-Language-Action Models 9 upvotes, #15 of 2025-06-12
  16. Hidden in plain sight: VLMs overlook their visual representations 8 upvotes, #16 of 2025-06-12
  17. Efficient Part-level 3D Object Generation via Dual Volume Packing 8 upvotes, #16 of 2025-06-12
  18. UFM: A Simple Path towards Unified Dense Correspondence with Flow 6 upvotes, #18 of 2025-06-12
  19. Can Vision Language Models Infer Human Gaze Direction? A Controlled Study 4 upvotes, #19 of 2025-06-12
  20. Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models 3 upvotes, #20 of 2025-06-12
  21. Reparameterized LLM Training via Orthogonal Equivalence Transformation 2 upvotes, #21 of 2025-06-12
  22. MIRAGE: Multimodal foundation model and benchmark for comprehensive retinal OCT image analysis 2 upvotes, #21 of 2025-06-12
  23. Query-Level Uncertainty in Large Language Models 2 upvotes, #21 of 2025-06-12
  24. When to Trust Context: Self-Reflective Debates for Context Reliability 1 upvotes, #24 of 2025-06-12
  25. Branched Schrödinger Bridge Matching 1 upvotes, #24 of 2025-06-12
  26. Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy 1 upvotes, #24 of 2025-06-12
  27. A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy 0 upvotes, #27 of 2025-06-12
  28. TTT-Bench: A Benchmark for Evaluating Reasoning Ability with Simple and Novel Tic-Tac-Toe-style Games 0 upvotes, #27 of 2025-06-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.