Linjie Li
Linjie Li on Hugging Face Daily Papers: 19 papers, 6 in the top 3 of their day, 437 upvotes.
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning 39 upvotes, #4 of 2025-05-16
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement 14 upvotes, #10 of 2025-04-11
- V-MAGE: A Game Evaluation Framework for Assessing Visual-Centric Capabilities in Multimodal Large Language Models 12 upvotes, #7 of 2025-04-09
- Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models 4 upvotes, #20 of 2025-03-27
- Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback 50 upvotes, #5 of 2025-01-23
- GenXD: Generating Any 3D and 4D Scenes 18 upvotes, #9 of 2024-11-05
- MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models 48 upvotes, #2 of 2024-10-15
- MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities 10 upvotes, #7 of 2024-08-02
- VideoGUI: A Benchmark for GUI Automation from Instructional Videos 8 upvotes, #13 of 2024-06-17
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos 23 upvotes, #10 of 2024-06-13
- Interfacing Foundation Models' Embeddings 10 upvotes, #11 of 2023-12-13
- GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation 14 upvotes, #7 of 2023-11-14
- The Generative AI Paradox: "What It Can Create, It May Not Understand" 18 upvotes, #5 of 2023-11-02
- MM-VID: Advancing Video Understanding with GPT-4V(ision) 20 upvotes, #2 of 2023-10-31
- DEsignBench: Exploring and Benchmarking DALL-E 3 for Imagining Visual Design 14 upvotes, #3 of 2023-10-24
- Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation 17 upvotes, #4 of 2023-10-13
- Multimodal Foundation Models: From Specialists to General-Purpose Assistants 41 upvotes, #2 of 2023-09-20
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities 19 upvotes, #2 of 2023-08-07
- DisCo: Disentangled Control for Referring Human Dance Generation in Real World 27 upvotes, #3 of 2023-07-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.