Zhengyuan Yang

Zhengyuan Yang on Hugging Face Daily Papers: 34 papers, 9 in the top 3 of their day, 869 upvotes.

  1. Computer-Use Agents as Judges for Generative User Interface 50 upvotes, #5 of 2025-11-25
  2. InfoAgent: Advancing Autonomous Information-Seeking Agents 10 upvotes, #27 of 2025-10-01
  3. EdiVal-Agent: An Object-Centric Framework for Automated, Scalable, Fine-Grained Evaluation of Multi-Turn Editing 3 upvotes, #18 of 2025-09-19
  4. STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models 25 upvotes, #10 of 2025-07-22
  5. Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers 76 upvotes, #2 of 2025-07-04
  6. ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs 20 upvotes, #5 of 2025-06-16
  7. Audio-Aware Large Language Models as Judges for Speaking Styles 14 upvotes, #11 of 2025-06-09
  8. OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning 39 upvotes, #4 of 2025-05-16
  9. RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning 8 upvotes, #11 of 2025-04-30
  10. SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement 14 upvotes, #10 of 2025-04-11
  11. V-MAGE: A Game Evaluation Framework for Assessing Visual-Centric Capabilities in Multimodal Large Language Models 12 upvotes, #7 of 2025-04-09
  12. Beyond Words: Advancing Long-Text Image Generation via Multimodal Autoregressive Models 4 upvotes, #20 of 2025-03-27
  13. TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation 42 upvotes, #3 of 2025-02-13
  14. ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding 14 upvotes, #9 of 2025-01-13
  15. OLA-VLM: Elevating Visual Perception in Multimodal LLMs with Auxiliary Embedding Distillation 10 upvotes, #14 of 2024-12-13
  16. ShowUI: One Vision-Language-Action Model for GUI Visual Agent 68 upvotes, #1 of 2024-11-27
  17. GenXD: Generating Any 3D and 4D Scenes 18 upvotes, #9 of 2024-11-05
  18. SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation 7 upvotes, #9 of 2024-10-31
  19. MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models 7 upvotes, #15 of 2024-10-15
  20. MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities 10 upvotes, #7 of 2024-08-02
  21. VideoGUI: A Benchmark for GUI Automation from Instructional Videos 8 upvotes, #13 of 2024-06-17
  22. MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos 23 upvotes, #10 of 2024-06-13
  23. List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs 14 upvotes, #8 of 2024-04-26
  24. Design2Code: How Far Are We From Automating Front-End Engineering? 85 upvotes, #1 of 2024-03-06
  25. StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis 20 upvotes, #4 of 2024-01-31
  26. COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training 17 upvotes, #4 of 2024-01-02
  27. Interfacing Foundation Models' Embeddings 10 upvotes, #11 of 2023-12-13
  28. GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation 14 upvotes, #7 of 2023-11-14
  29. MM-VID: Advancing Video Understanding with GPT-4V(ision) 20 upvotes, #2 of 2023-10-31
  30. DEsignBench: Exploring and Benchmarking DALL-E 3 for Imagining Visual Design 14 upvotes, #3 of 2023-10-24
  31. Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation 17 upvotes, #4 of 2023-10-13
  32. Multimodal Foundation Models: From Specialists to General-Purpose Assistants 41 upvotes, #2 of 2023-09-20
  33. MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities 19 upvotes, #2 of 2023-08-07
  34. DisCo: Disentangled Control for Referring Human Dance Generation in Real World 27 upvotes, #3 of 2023-07-04

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.