Daily Papers of 2025-06-04
- Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning 186 upvotes, #1 of 2025-06-04
- UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation 58 upvotes, #2 of 2025-06-04
- VS-Bench: Evaluating VLMs for Strategic Reasoning and Decision-Making in Multi-Agent Environments 56 upvotes, #3 of 2025-06-04
- SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis 51 upvotes, #4 of 2025-06-04
- CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs 48 upvotes, #5 of 2025-06-04
- GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents 44 upvotes, #6 of 2025-06-04
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models 37 upvotes, #7 of 2025-06-04
- FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation 35 upvotes, #8 of 2025-06-04
- OThink-R1: Intrinsic Fast/Slow Thinking Mode Switching for Over-Reasoning Mitigation 35 upvotes, #8 of 2025-06-04
- Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces 32 upvotes, #10 of 2025-06-04
- DINGO: Constrained Inference for Diffusion LLMs 29 upvotes, #11 of 2025-06-04
- Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics 28 upvotes, #12 of 2025-06-04
- Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers 27 upvotes, #13 of 2025-06-04
- MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs 26 upvotes, #14 of 2025-06-04
- AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation 22 upvotes, #15 of 2025-06-04
- Co-Evolving LLM Coder and Unit Tester via Reinforcement Learning 22 upvotes, #15 of 2025-06-04
- Negative-Guided Subject Fidelity Optimization for Zero-Shot Subject-Driven Generation 22 upvotes, #15 of 2025-06-04
- LumosFlow: Motion-Guided Long Video Generation 18 upvotes, #18 of 2025-06-04
- Native-Resolution Image Synthesis 18 upvotes, #18 of 2025-06-04
- RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers 15 upvotes, #20 of 2025-06-04
- FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation 14 upvotes, #21 of 2025-06-04
- DCM: Dual-Expert Consistency Model for Efficient and High-Quality Video Generation 14 upvotes, #21 of 2025-06-04
- Datasheets Aren't Enough: DataRubrics for Automated Quality Metrics and Accountability 13 upvotes, #23 of 2025-06-04
- Ctrl-Crash: Controllable Diffusion for Realistic Car Crashes 11 upvotes, #24 of 2025-06-04
- Training Language Models to Generate Quality Code with Program Analysis Feedback 10 upvotes, #25 of 2025-06-04
- PCoreSet: Effective Active Learning through Knowledge Distillation from Vision-Language Models 10 upvotes, #25 of 2025-06-04
- Self-Challenging Language Model Agents 9 upvotes, #27 of 2025-06-04
- Motion-Aware Concept Alignment for Consistent Video Editing 7 upvotes, #28 of 2025-06-04
- SHARE: An SLM-based Hierarchical Action CorREction Assistant for Text-to-SQL 6 upvotes, #29 of 2025-06-04
- Accelerating Diffusion LLMs via Adaptive Parallel Decoding 6 upvotes, #29 of 2025-06-04
- ORV: 4D Occupancy-centric Robot Video Generation 6 upvotes, #29 of 2025-06-04
- How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning 4 upvotes, #32 of 2025-06-04
- One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL 4 upvotes, #32 of 2025-06-04
- Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports From Scratch with Agentic Framework 4 upvotes, #32 of 2025-06-04
- Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding 3 upvotes, #35 of 2025-06-04
- Control-R: Towards controllable test-time scaling 3 upvotes, #35 of 2025-06-04
- ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding 3 upvotes, #35 of 2025-06-04
- Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation 3 upvotes, #35 of 2025-06-04
- Angles Don't Lie: Unlocking Training-Efficient RL Through the Model's Own Signals 3 upvotes, #35 of 2025-06-04
- TL;DR: Too Long, Do Re-weighting for Effcient LLM Reasoning Compression 3 upvotes, #35 of 2025-06-04
- FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens 3 upvotes, #35 of 2025-06-04
- MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query 3 upvotes, #35 of 2025-06-04
- R^2ec: Towards Large Recommender Models with Reasoning 2 upvotes, #43 of 2025-06-04
- Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion 2 upvotes, #43 of 2025-06-04
- QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation 2 upvotes, #43 of 2025-06-04
- M^3FinMeeting: A Multilingual, Multi-Sector, and Multi-Task Financial Meeting Understanding Evaluation Dataset 2 upvotes, #43 of 2025-06-04
- Controllable Human-centric Keyframe Interpolation with Generative Prior 2 upvotes, #43 of 2025-06-04
- Beyond In-Context Learning: Aligning Long-form Generation of Large Language Models via Task-Inherent Attribute Guidelines 1 upvotes, #48 of 2025-06-04
- Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability 1 upvotes, #48 of 2025-06-04
- ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions 1 upvotes, #48 of 2025-06-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.