Daily Papers of 2025-03-28

  1. Video-R1: Reinforcing Video Reasoning in MLLMs 74 upvotes, #1 of 2025-03-28
  2. Large Language Model Agent: A Survey on Methodology, Applications and Challenges 67 upvotes, #2 of 2025-03-28
  3. UI-R1: Enhancing Action Prediction of GUI Agents by Reinforcement Learning 54 upvotes, #3 of 2025-03-28
  4. Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models 35 upvotes, #4 of 2025-03-28
  5. VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness 30 upvotes, #5 of 2025-03-28
  6. ReaRAG: Knowledge-guided Reasoning Enhances Factuality of Large Reasoning Models with Iterative Retrieval Augmented Generation 25 upvotes, #6 of 2025-03-28
  7. LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis 25 upvotes, #6 of 2025-03-28
  8. ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model 23 upvotes, #8 of 2025-03-28
  9. Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks 21 upvotes, #9 of 2025-03-28
  10. ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition 19 upvotes, #10 of 2025-03-28
  11. FinAudio: A Benchmark for Audio Large Language Models in Financial Applications 18 upvotes, #11 of 2025-03-28
  12. Lumina-Image 2.0: A Unified and Efficient Image Generative Framework 18 upvotes, #11 of 2025-03-28
  13. Synthetic Video Enhances Physical Fidelity in Video Synthesis 15 upvotes, #13 of 2025-03-28
  14. Optimal Stepsize for Diffusion Sampling 13 upvotes, #14 of 2025-03-28
  15. Exploring the Evolution of Physics Cognition in Video Generation: A Survey 11 upvotes, #15 of 2025-03-28
  16. Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields 8 upvotes, #16 of 2025-03-28
  17. Unified Multimodal Discrete Diffusion 8 upvotes, #16 of 2025-03-28
  18. ZJUKLAB at SemEval-2025 Task 4: Unlearning via Model Merging 6 upvotes, #18 of 2025-03-28
  19. Semantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic Segmentation 6 upvotes, #18 of 2025-03-28
  20. LLPut: Investigating Large Language Models for Bug Report-Based Input Generation 4 upvotes, #20 of 2025-03-28
  21. Tracktention: Leveraging Point Tracking to Attend Videos Faster and Better 2 upvotes, #21 of 2025-03-28
  22. LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Text-Guided Image Editing 1 upvotes, #22 of 2025-03-28

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.