Daily Papers of 2025-05-08
- On Path to Multimodal Generalist: General-Level and General-Bench 72 upvotes, #1 of 2025-05-08
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities 67 upvotes, #2 of 2025-05-08
- ZeroSearch: Incentivize the Search Capability of LLMs without Searching 58 upvotes, #3 of 2025-05-08
- HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation 33 upvotes, #4 of 2025-05-08
- PrimitiveAnything: Human-Crafted 3D Primitive Assembly Generation with Auto-Regressive Transformer 25 upvotes, #5 of 2025-05-08
- Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models 22 upvotes, #6 of 2025-05-08
- R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training 20 upvotes, #7 of 2025-05-08
- OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning 20 upvotes, #7 of 2025-05-08
- Benchmarking LLMs' Swarm intelligence 18 upvotes, #9 of 2025-05-08
- LLM-Independent Adaptive RAG: Let the Question Speak for Itself 11 upvotes, #10 of 2025-05-08
- Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving 10 upvotes, #11 of 2025-05-08
- Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey 8 upvotes, #12 of 2025-05-08
- OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation 8 upvotes, #12 of 2025-05-08
- OSUniverse: Benchmark for Multimodal GUI-navigation AI Agents 7 upvotes, #14 of 2025-05-08
- OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution 6 upvotes, #15 of 2025-05-08
- COSMOS: Predictable and Cost-Effective Adaptation of LLMs 3 upvotes, #16 of 2025-05-08
- AutoLibra: Agent Metric Induction from Open-Ended Feedback 3 upvotes, #16 of 2025-05-08
- Uncertainty-Weighted Image-Event Multimodal Fusion for Video Anomaly Detection 2 upvotes, #18 of 2025-05-08
- RAIL: Region-Aware Instructive Learning for Semi-Supervised Tooth Segmentation in CBCT 2 upvotes, #18 of 2025-05-08
- Cognitio Emergens: Agency, Dimensions, and Dynamics in Human-AI Knowledge Co-Creation 1 upvotes, #20 of 2025-05-08
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.