Zhengyuan Yang
Zhengyuan Yang on Hugging Face Daily Papers: 34 papers, 9 in the top 3 of their day, 869 upvotes.
- Computer-Use Agents as Judges for Generative User Interface 50 upvotes, #5 of 2025-11-25
- InfoAgent: Advancing Autonomous Information-Seeking Agents 10 upvotes, #27 of 2025-10-01
- EdiVal-Agent: An Object-Centric Framework for Automated, Scalable, Fine-Grained Evaluation of Multi-Turn Editing 3 upvotes, #18 of 2025-09-19
- STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models 25 upvotes, #10 of 2025-07-22
- Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers 76 upvotes, #2 of 2025-07-04
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs 20 upvotes, #5 of 2025-06-16
- Audio-Aware Large Language Models as Judges for Speaking Styles 14 upvotes, #11 of 2025-06-09
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning 39 upvotes, #4 of 2025-05-16
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning 8 upvotes, #11 of 2025-04-30
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement 14 upvotes, #10 of 2025-04-11
- V-MAGE: A Game Evaluation Framework for Assessing Visual-Centric Capabilities in Multimodal Large Language Models 12 upvotes, #7 of 2025-04-09
- Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models 4 upvotes, #20 of 2025-03-27
- TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation 42 upvotes, #3 of 2025-02-13
- ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding 14 upvotes, #9 of 2025-01-13
- OLA-VLM: Elevating Visual Perception in Multimodal LLMs with Auxiliary Embedding Distillation 10 upvotes, #14 of 2024-12-13
- ShowUI: One Vision-Language-Action Model for GUI Visual Agent 68 upvotes, #1 of 2024-11-27
- GenXD: Generating Any 3D and 4D Scenes 18 upvotes, #9 of 2024-11-05
- SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation 7 upvotes, #9 of 2024-10-31
- MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models 7 upvotes, #15 of 2024-10-15
- MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities 10 upvotes, #7 of 2024-08-02
- VideoGUI: A Benchmark for GUI Automation from Instructional Videos 8 upvotes, #13 of 2024-06-17
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos 23 upvotes, #10 of 2024-06-13
- List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs 14 upvotes, #8 of 2024-04-26
- Design2Code: How Far Are We From Automating Front-End Engineering? 85 upvotes, #1 of 2024-03-06
- StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis 20 upvotes, #4 of 2024-01-31
- COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training 17 upvotes, #4 of 2024-01-02
- Interfacing Foundation Models' Embeddings 10 upvotes, #11 of 2023-12-13
- GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation 14 upvotes, #7 of 2023-11-14
- MM-VID: Advancing Video Understanding with GPT-4V(ision) 20 upvotes, #2 of 2023-10-31
- DEsignBench: Exploring and Benchmarking DALL-E 3 for Imagining Visual Design 14 upvotes, #3 of 2023-10-24
- Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation 17 upvotes, #4 of 2023-10-13
- Multimodal Foundation Models: From Specialists to General-Purpose Assistants 41 upvotes, #2 of 2023-09-20
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities 19 upvotes, #2 of 2023-08-07
- DisCo: Disentangled Control for Referring Human Dance Generation in Real World 27 upvotes, #3 of 2023-07-04
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.