Yufan Zhou
Yufan Zhou on Hugging Face Daily Papers: 7 papers, 0 in the top 3 of their day, 55 upvotes.
- Towards Visual Text Grounding of Multimodal Large Language Model 13 upvotes, #11 of 2025-04-11
- SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner 5 upvotes, #10 of 2024-12-18
- LoRA-Contextualizing Adaptation of Large Multimodal Models for Long Document Understanding 4 upvotes, #22 of 2024-11-05
- Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models 5 upvotes, #11 of 2024-10-09
- Toffee: Efficient Million-Scale Dataset Construction for Subject-Driven Text-to-Image Generation 4 upvotes, #25 of 2024-06-14
- LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding 12 upvotes, #6 of 2023-06-30
- Enhancing Detail Preservation for Customized Text-to-Image Generation: A Regularization-Free Approach 4 upvotes, #5 of 2023-05-24
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.