Xiaohan Wang

Xiaohan Wang on Hugging Face Daily Papers: 9 papers, 3 in the top 3 of their day, 427 upvotes.

  1. FineVision: Open Data Is All You Need 59 upvotes, #4 of 2025-10-21
  2. SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models 8 upvotes, #30 of 2025-10-10
  3. Video Action Differencing 30 upvotes, #8 of 2025-03-12
  4. Temporal Preference Optimization for Long-Form Video Understanding 21 upvotes, #6 of 2025-01-24
  5. BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature 46 upvotes, #4 of 2025-01-14
  6. Apollo: An Exploration of Video Understanding in Large Multimodal Models 131 upvotes, #1 of 2024-12-16
  7. Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision 24 upvotes, #3 of 2024-07-10
  8. VideoAgent: Long-form Video Understanding with Large Language Model as Agent 27 upvotes, #3 of 2024-03-18
  9. Describing Differences in Image Sets with Natural Language 15 upvotes, #9 of 2023-12-06

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.