Daily Papers of 2025-06-04

  1. Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning 186 upvotes, #1 of 2025-06-04
  2. UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation 58 upvotes, #2 of 2025-06-04
  3. VS-Bench: Evaluating VLMs for Strategic Reasoning and Decision-Making in Multi-Agent Environments 56 upvotes, #3 of 2025-06-04
  4. SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis 51 upvotes, #4 of 2025-06-04
  5. CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs 48 upvotes, #5 of 2025-06-04
  6. GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents 44 upvotes, #6 of 2025-06-04
  7. OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models 37 upvotes, #7 of 2025-06-04
  8. FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation 35 upvotes, #8 of 2025-06-04
  9. OThink-R1: Intrinsic Fast/Slow Thinking Mode Switching for Over-Reasoning Mitigation 35 upvotes, #8 of 2025-06-04
  10. Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces 32 upvotes, #10 of 2025-06-04
  11. DINGO: Constrained Inference for Diffusion LLMs 29 upvotes, #11 of 2025-06-04
  12. Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics 28 upvotes, #12 of 2025-06-04
  13. Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers 27 upvotes, #13 of 2025-06-04
  14. MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs 26 upvotes, #14 of 2025-06-04
  15. AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation 22 upvotes, #15 of 2025-06-04
  16. Co-Evolving LLM Coder and Unit Tester via Reinforcement Learning 22 upvotes, #15 of 2025-06-04
  17. Negative-Guided Subject Fidelity Optimization for Zero-Shot Subject-Driven Generation 22 upvotes, #15 of 2025-06-04
  18. LumosFlow: Motion-Guided Long Video Generation 18 upvotes, #18 of 2025-06-04
  19. Native-Resolution Image Synthesis 18 upvotes, #18 of 2025-06-04
  20. RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers 15 upvotes, #20 of 2025-06-04
  21. FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation 14 upvotes, #21 of 2025-06-04
  22. DCM: Dual-Expert Consistency Model for Efficient and High-Quality Video Generation 14 upvotes, #21 of 2025-06-04
  23. Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability 13 upvotes, #23 of 2025-06-04
  24. Ctrl-Crash: Controllable Diffusion for Realistic Car Crashes 11 upvotes, #24 of 2025-06-04
  25. Training Language Models to Generate Quality Code with Program Analysis Feedback 10 upvotes, #25 of 2025-06-04
  26. PCoreSet: Effective Active Learning through Knowledge Distillation from Vision-Language Models 10 upvotes, #25 of 2025-06-04
  27. Self-Challenging Language Model Agents 9 upvotes, #27 of 2025-06-04
  28. Motion-Aware Concept Alignment for Consistent Video Editing 7 upvotes, #28 of 2025-06-04
  29. SHARE: An SLM-based Hierarchical Action CorREction Assistant for Text-to-SQL 6 upvotes, #29 of 2025-06-04
  30. Accelerating Diffusion LLMs via Adaptive Parallel Decoding 6 upvotes, #29 of 2025-06-04
  31. ORV: 4D Occupancy-centric Robot Video Generation 6 upvotes, #29 of 2025-06-04
  32. How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning 4 upvotes, #32 of 2025-06-04
  33. One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL 4 upvotes, #32 of 2025-06-04
  34. Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports From Scratch with Agentic Framework 4 upvotes, #32 of 2025-06-04
  35. Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding 3 upvotes, #35 of 2025-06-04
  36. Control-R: Towards controllable test-time scaling 3 upvotes, #35 of 2025-06-04
  37. ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding 3 upvotes, #35 of 2025-06-04
  38. Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation 3 upvotes, #35 of 2025-06-04
  39. Angles Don't Lie: Unlocking Training-Efficient RL Through the Model's Own Signals 3 upvotes, #35 of 2025-06-04
  40. TL;DR: Too Long, Do Re-weighting for Effcient LLM Reasoning Compression 3 upvotes, #35 of 2025-06-04
  41. FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens 3 upvotes, #35 of 2025-06-04
  42. MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query 3 upvotes, #35 of 2025-06-04
  43. R^2ec: Towards Large Recommender Models with Reasoning 2 upvotes, #43 of 2025-06-04
  44. Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion 2 upvotes, #43 of 2025-06-04
  45. QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation 2 upvotes, #43 of 2025-06-04
  46. M^3FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset 2 upvotes, #43 of 2025-06-04
  47. Controllable Human-centric Keyframe Interpolation with Generative Prior 2 upvotes, #43 of 2025-06-04
  48. Beyond In-Context Learning: Aligning Long-form Generation of Large Language Models via Task-Inherent Attribute Guidelines 1 upvotes, #48 of 2025-06-04
  49. Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability 1 upvotes, #48 of 2025-06-04
  50. ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions 1 upvotes, #48 of 2025-06-04

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.