Zhengfeng Lai
Zhengfeng Lai on Hugging Face Daily Papers: 8 papers, 4 in the top 3 of their day, 320 upvotes.
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant 13 upvotes, #9 of 2025-05-09
- ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering 11 upvotes, #13 of 2025-03-24
- STIV: Scalable Text and Image Conditioned Video Generation 67 upvotes, #1 of 2024-12-11
- Contrastive Localized Language-Image Pre-Training 28 upvotes, #7 of 2024-10-04
- Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models 51 upvotes, #1 of 2024-10-04
- MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning 46 upvotes, #2 of 2024-10-01
- MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains 37 upvotes, #6 of 2024-07-30
- SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models 32 upvotes, #1 of 2024-07-23
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.