Jianwei Yang

Jianwei Yang on Hugging Face Daily Papers: 18 papers, 6 in the top 3 of their day, 437 upvotes.

  1. ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding 14 upvotes, #9 of 2025-01-13
  2. TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies 2 upvotes, #17 of 2024-12-16
  3. OLA-VLM: Elevating Visual Perception in Multimodal LLMs with Auxiliary Embedding Distillation 10 upvotes, #14 of 2024-12-13
  4. Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion 50 upvotes, #2 of 2024-12-06
  5. TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models 14 upvotes, #11 of 2024-10-15
  6. OmniParser for Pure Vision Based GUI Agent 16 upvotes, #6 of 2024-08-02
  7. Matryoshka Multimodal Models 29 upvotes, #3 of 2024-05-28
  8. List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs 14 upvotes, #8 of 2024-04-26
  9. Pix2Gif: Motion-Guided Diffusion for GIF Generation 12 upvotes, #9 of 2024-03-08
  10. Interfacing Foundation Models' Embeddings 10 upvotes, #11 of 2023-12-13
  11. LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models 14 upvotes, #10 of 2023-12-06
  12. GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation 14 upvotes, #7 of 2023-11-14
  13. LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents 50 upvotes, #2 of 2023-11-10
  14. LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing 42 upvotes, #2 of 2023-11-02
  15. Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V 29 upvotes, #4 of 2023-10-18
  16. Multimodal Foundation Models: From Specialists to General-Purpose Assistants 41 upvotes, #2 of 2023-09-20
  17. An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models 19 upvotes, #6 of 2023-09-19
  18. Semantic-SAM: Segment and Recognize Anything at Any Granularity 23 upvotes, #3 of 2023-07-11

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.