Daily Papers of 2025-03-28
- Video-R1: Reinforcing Video Reasoning in MLLMs 74 upvotes, #1 of 2025-03-28
- Large Language Model Agent: A Survey on Methodology, Applications and Challenges 67 upvotes, #2 of 2025-03-28
- UI-R1: Enhancing Action Prediction of GUI Agents by Reinforcement Learning 54 upvotes, #3 of 2025-03-28
- Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models 35 upvotes, #4 of 2025-03-28
- VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness 30 upvotes, #5 of 2025-03-28
- ReaRAG: Knowledge-guided Reasoning Enhances Factuality of Large Reasoning Models with Iterative Retrieval Augmented Generation 25 upvotes, #6 of 2025-03-28
- LeX-Art: Rethinking Text Generation via Scalable High-Quality Data Synthesis 25 upvotes, #6 of 2025-03-28
- ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model 23 upvotes, #8 of 2025-03-28
- Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks 21 upvotes, #9 of 2025-03-28
- ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition 19 upvotes, #10 of 2025-03-28
- FinAudio: A Benchmark for Audio Large Language Models in Financial Applications 18 upvotes, #11 of 2025-03-28
- Lumina-Image 2.0: A Unified and Efficient Image Generative Framework 18 upvotes, #11 of 2025-03-28
- Synthetic Video Enhances Physical Fidelity in Video Synthesis 15 upvotes, #13 of 2025-03-28
- Optimal Stepsize for Diffusion Sampling 13 upvotes, #14 of 2025-03-28
- Exploring the Evolution of Physics Cognition in Video Generation: A Survey 11 upvotes, #15 of 2025-03-28
- Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields 8 upvotes, #16 of 2025-03-28
- Unified Multimodal Discrete Diffusion 8 upvotes, #16 of 2025-03-28
- ZJUKLAB at SemEval-2025 Task 4: Unlearning via Model Merging 6 upvotes, #18 of 2025-03-28
- Semantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic Segmentation 6 upvotes, #18 of 2025-03-28
- LLPut: Investigating Large Language Models for Bug Report-Based Input Generation 4 upvotes, #20 of 2025-03-28
- Tracktention: Leveraging Point Tracking to Attend Videos Faster and Better 2 upvotes, #21 of 2025-03-28
- LOCATEdit: Graph Laplacian Optimized Cross Attention for Localized Text-Guided Image Editing 1 upvotes, #22 of 2025-03-28
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.