Linjie Li

Linjie Li on Hugging Face Daily Papers: 19 papers, 6 in the top 3 of their day, 437 upvotes.

  1. OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning 39 upvotes, #4 of 2025-05-16
  2. SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement 14 upvotes, #10 of 2025-04-11
  3. V-MAGE: A Game Evaluation Framework for Assessing Visual-Centric Capabilities in Multimodal Large Language Models 12 upvotes, #7 of 2025-04-09
  4. Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models 4 upvotes, #20 of 2025-03-27
  5. Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback 50 upvotes, #5 of 2025-01-23
  6. GenXD: Generating Any 3D and 4D Scenes 18 upvotes, #9 of 2024-11-05
  7. MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models 48 upvotes, #2 of 2024-10-15
  8. MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities 10 upvotes, #7 of 2024-08-02
  9. VideoGUI: A Benchmark for GUI Automation from Instructional Videos 8 upvotes, #13 of 2024-06-17
  10. MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos 23 upvotes, #10 of 2024-06-13
  11. Interfacing Foundation Models' Embeddings 10 upvotes, #11 of 2023-12-13
  12. GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation 14 upvotes, #7 of 2023-11-14
  13. The Generative AI Paradox: "What It Can Create, It May Not Understand" 18 upvotes, #5 of 2023-11-02
  14. MM-VID: Advancing Video Understanding with GPT-4V(ision) 20 upvotes, #2 of 2023-10-31
  15. DEsignBench: Exploring and Benchmarking DALL-E 3 for Imagining Visual Design 14 upvotes, #3 of 2023-10-24
  16. Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation 17 upvotes, #4 of 2023-10-13
  17. Multimodal Foundation Models: From Specialists to General-Purpose Assistants 41 upvotes, #2 of 2023-09-20
  18. MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities 19 upvotes, #2 of 2023-08-07
  19. DisCo: Disentangled Control for Referring Human Dance Generation in Real World 27 upvotes, #3 of 2023-07-04

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.