Chunyuan Li

Chunyuan Li on Hugging Face Daily Papers: 23 papers, 11 in the top 3 of their day, 726 upvotes.

  1. UniT: Unified Multimodal Chain-of-Thought Test-time Scaling 19 upvotes, #8 of 2026-02-18
  2. Video Instruction Tuning With Synthetic Data 33 upvotes, #5 of 2024-10-04
  3. LLaVA-Critic: Learning to Evaluate Multimodal Models 31 upvotes, #6 of 2024-10-04
  4. MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines 33 upvotes, #3 of 2024-09-20
  5. SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners 25 upvotes, #5 of 2024-08-30
  6. LLaVA-OneVision: Easy Visual Task Transfer 50 upvotes, #2 of 2024-08-07
  7. LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models 34 upvotes, #3 of 2024-07-11
  8. Long Context Transfer from Language to Vision 30 upvotes, #4 of 2024-06-25
  9. MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding 17 upvotes, #9 of 2024-06-14
  10. TrustLLM: Trustworthiness in Large Language Models 69 upvotes, #1 of 2024-01-12
  11. LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models 14 upvotes, #10 of 2023-12-06
  12. Visual In-Context Prompting 18 upvotes, #7 of 2023-11-23
  13. LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents 50 upvotes, #2 of 2023-11-10
  14. LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing 42 upvotes, #2 of 2023-11-02
  15. Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V 29 upvotes, #4 of 2023-10-18
  16. Improved Baselines with Visual Instruction Tuning 39 upvotes, #2 of 2023-10-06
  17. Aligning Large Multimodal Models with Factually Augmented RLHF 32 upvotes, #4 of 2023-09-27
  18. Multimodal Foundation Models: From Specialists to General-Purpose Assistants 41 upvotes, #2 of 2023-09-20
  19. An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models 19 upvotes, #6 of 2023-09-19
  20. Semantic-SAM: Segment and Recognize Anything at Any Granularity 23 upvotes, #3 of 2023-07-11
  21. MIMIC-IT: Multi-Modal In-Context Instruction Tuning 12 upvotes, #2 of 2023-06-09
  22. LLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One Day 14 upvotes, #3 of 2023-06-02
  23. Towards Building the Federated GPT: Federated Instruction Tuning 6 upvotes, #4 of 2023-05-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.