Daily Papers of 2025-06-05

  1. MiMo-VL Technical Report 70 upvotes, #1 of 2025-06-05
  2. AmbiK: Dataset of Ambiguous Tasks in Kitchen Environment 45 upvotes, #2 of 2025-06-05
  3. Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning 45 upvotes, #2 of 2025-06-05
  4. OpenThoughts: Data Recipes for Reasoning Models 39 upvotes, #4 of 2025-06-05
  5. A Controllable Examination for Long-Context Language Models 32 upvotes, #5 of 2025-06-05
  6. SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models 32 upvotes, #5 of 2025-06-05
  7. MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos 29 upvotes, #7 of 2025-06-05
  8. Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis 26 upvotes, #8 of 2025-06-05
  9. Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation 24 upvotes, #9 of 2025-06-05
  10. VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation 22 upvotes, #10 of 2025-06-05
  11. Image Editing As Programs with Diffusion Models 22 upvotes, #10 of 2025-06-05
  12. IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation 21 upvotes, #12 of 2025-06-05
  13. Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem 17 upvotes, #13 of 2025-06-05
  14. Ψ-Sampler: Initial Particle Sampling for SMC-Based Inference-Time Reward Alignment in Score Models 16 upvotes, #14 of 2025-06-05
  15. SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation 14 upvotes, #15 of 2025-06-05
  16. DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models 13 upvotes, #16 of 2025-06-05
  17. LayerFlow: A Unified Model for Layer-aware Video Generation 13 upvotes, #16 of 2025-06-05
  18. TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence 11 upvotes, #18 of 2025-06-05
  19. Rectified Sparse Attention 10 upvotes, #19 of 2025-06-05
  20. TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models 9 upvotes, #20 of 2025-06-05
  21. Orak: A Foundational Benchmark for Training and Evaluating LLM Agents on Diverse Video Games 9 upvotes, #20 of 2025-06-05
  22. BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation 8 upvotes, #22 of 2025-06-05
  23. Beyond the Surface: Measuring Self-Preference in LLM Judgments 8 upvotes, #22 of 2025-06-05
  24. DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers 7 upvotes, #24 of 2025-06-05
  25. CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech 6 upvotes, #25 of 2025-06-05
  26. Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback 6 upvotes, #25 of 2025-06-05
  27. Robustness in Both Domains: CLIP Needs a Robust Text Encoder 6 upvotes, #25 of 2025-06-05
  28. Video-Skill-CoT: Skill-based Chain-of-Thoughts for Domain-Adaptive Video Reasoning 6 upvotes, #25 of 2025-06-05
  29. POSS: Position Specialist Generates Better Draft for Speculative Decoding 6 upvotes, #25 of 2025-06-05
  30. Improving Knowledge Distillation Under Unknown Covariate Shift Through Confidence-Guided Data Augmentation 5 upvotes, #30 of 2025-06-05
  31. Quantitative LLM Judges 5 upvotes, #30 of 2025-06-05
  32. Adapt before Continual Learning 5 upvotes, #30 of 2025-06-05
  33. DLP: Dynamic Layerwise Pruning in Large Language Models 4 upvotes, #33 of 2025-06-05
  34. Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents 4 upvotes, #33 of 2025-06-05
  35. RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions 4 upvotes, #33 of 2025-06-05
  36. Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models 3 upvotes, #36 of 2025-06-05
  37. Small Language Models are the Future of Agentic AI 3 upvotes, #36 of 2025-06-05
  38. HTSC-2025: A Benchmark Dataset of Ambient-Pressure High-Temperature Superconductors for AI-Driven Critical Temperature Prediction 3 upvotes, #36 of 2025-06-05
  39. TRiSM for Agentic AI: A Review of Trust, Risk, and Security Management in LLM-based Agentic Multi-Agent Systems 3 upvotes, #36 of 2025-06-05
  40. Unleashing Hour-Scale Video Training for Long Video-Language Understanding 3 upvotes, #36 of 2025-06-05
  41. FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning 2 upvotes, #41 of 2025-06-05
  42. Solving Inverse Problems with FLAIR 2 upvotes, #41 of 2025-06-05
  43. Robust Neural Rendering in the Wild with Asymmetric Dual 3D Gaussian Splatting 2 upvotes, #41 of 2025-06-05
  44. VLMs Can Aggregate Scattered Training Patches 2 upvotes, #41 of 2025-06-05
  45. CRAWLDoc: A Dataset for Robust Ranking of Bibliographic Documents 2 upvotes, #41 of 2025-06-05
  46. Rethinking the Stability-Plasticity Trade-off in Continual Learning from an Architectural Perspective 2 upvotes, #41 of 2025-06-05
  47. Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning 2 upvotes, #41 of 2025-06-05
  48. RiOSWorld: Benchmarking the Risk of Multimodal Compter-Use Agents 1 upvotes, #48 of 2025-06-05
  49. Survey of Active Learning Hyperparameters: Insights from a Large-Scale Experimental Grid 1 upvotes, #48 of 2025-06-05
  50. Sounding that Object: Interactive Object-Aware Image to Audio Generation 1 upvotes, #48 of 2025-06-05

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.