Jianwei Yang
Jianwei Yang on Hugging Face Daily Papers: 18 papers, 6 in the top 3 of their day, 437 upvotes.
- ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding 14 upvotes, #9 of 2025-01-13
- TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies 2 upvotes, #17 of 2024-12-16
- OLA-VLM: Elevating Visual Perception in Multimodal LLMs with Auxiliary Embedding Distillation 10 upvotes, #14 of 2024-12-13
- Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion 50 upvotes, #2 of 2024-12-06
- TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models 14 upvotes, #11 of 2024-10-15
- OmniParser for Pure Vision Based GUI Agent 16 upvotes, #6 of 2024-08-02
- Matryoshka Multimodal Models 29 upvotes, #3 of 2024-05-28
- List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs 14 upvotes, #8 of 2024-04-26
- Pix2Gif: Motion-Guided Diffusion for GIF Generation 12 upvotes, #9 of 2024-03-08
- Interfacing Foundation Models' Embeddings 10 upvotes, #11 of 2023-12-13
- LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models 14 upvotes, #10 of 2023-12-06
- GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation 14 upvotes, #7 of 2023-11-14
- LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents 50 upvotes, #2 of 2023-11-10
- LLaVA-Interactive: An All-in-One Demo for Image Chat, Segmentation, Generation and Editing 42 upvotes, #2 of 2023-11-02
- Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V 29 upvotes, #4 of 2023-10-18
- Multimodal Foundation Models: From Specialists to General-Purpose Assistants 41 upvotes, #2 of 2023-09-20
- An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models 19 upvotes, #6 of 2023-09-19
- Semantic-SAM: Segment and Recognize Anything at Any Granularity 23 upvotes, #3 of 2023-07-11
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.