Daily Papers of 2025-05-08

  1. On Path to Multimodal Generalist: General-Level and General-Bench 72 upvotes, #1 of 2025-05-08
  2. Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities 67 upvotes, #2 of 2025-05-08
  3. ZeroSearch: Incentivize the Search Capability of LLMs without Searching 58 upvotes, #3 of 2025-05-08
  4. HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation 33 upvotes, #4 of 2025-05-08
  5. PrimitiveAnything: Human-Crafted 3D Primitive Assembly Generation with Auto-Regressive Transformer 25 upvotes, #5 of 2025-05-08
  6. Beyond Recognition: Evaluating Visual Perspective Taking in Vision Language Models 22 upvotes, #6 of 2025-05-08
  7. R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training 20 upvotes, #7 of 2025-05-08
  8. OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning 20 upvotes, #7 of 2025-05-08
  9. Benchmarking LLMs' Swarm intelligence 18 upvotes, #9 of 2025-05-08
  10. LLM-Independent Adaptive RAG: Let the Question Speak for Itself 11 upvotes, #10 of 2025-05-08
  11. Beyond Theorem Proving: Formulation, Framework and Benchmark for Formal Problem-Solving 10 upvotes, #11 of 2025-05-08
  12. Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey 8 upvotes, #12 of 2025-05-08
  13. OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation 8 upvotes, #12 of 2025-05-08
  14. OSUniverse: Benchmark for Multimodal GUI-navigation AI Agents 7 upvotes, #14 of 2025-05-08
  15. OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution 6 upvotes, #15 of 2025-05-08
  16. COSMOS: Predictable and Cost-Effective Adaptation of LLMs 3 upvotes, #16 of 2025-05-08
  17. AutoLibra: Agent Metric Induction from Open-Ended Feedback 3 upvotes, #16 of 2025-05-08
  18. Uncertainty-Weighted Image-Event Multimodal Fusion for Video Anomaly Detection 2 upvotes, #18 of 2025-05-08
  19. RAIL: Region-Aware Instructive Learning for Semi-Supervised Tooth Segmentation in CBCT 2 upvotes, #18 of 2025-05-08
  20. Cognitio Emergens: Agency, Dimensions, and Dynamics in Human-AI Knowledge Co-Creation 1 upvotes, #20 of 2025-05-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.