Lijuan Wang
Lijuan Wang on Hugging Face Daily Papers: 17 papers, 5 in the top 3 of their day, 295 upvotes.
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs 20 upvotes, #5 of 2025-06-16
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement 14 upvotes, #10 of 2025-04-11
- Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models 4 upvotes, #20 of 2025-03-27
- MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities 10 upvotes, #7 of 2024-08-02
- VideoGUI: A Benchmark for GUI Automation from Instructional Videos 8 upvotes, #13 of 2024-06-17
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos 23 upvotes, #10 of 2024-06-13
- List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs 14 upvotes, #8 of 2024-04-26
- StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis 20 upvotes, #4 of 2024-01-31
- Interfacing Foundation Models' Embeddings 10 upvotes, #11 of 2023-12-13
- Segment and Caption Anything 21 upvotes, #6 of 2023-12-05
- GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation 14 upvotes, #7 of 2023-11-14
- MM-VID: Advancing Video Understanding with GPT-4V(ision) 20 upvotes, #2 of 2023-10-31
- DEsignBench: Exploring and Benchmarking DALL-E 3 for Imagining Visual Design 14 upvotes, #3 of 2023-10-24
- Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation 17 upvotes, #4 of 2023-10-13
- Multimodal Foundation Models: From Specialists to General-Purpose Assistants 41 upvotes, #2 of 2023-09-20
- ORES: Open-vocabulary Responsible Visual Synthesis 8 upvotes, #2 of 2023-08-29
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities 19 upvotes, #2 of 2023-08-07
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.