Can Qin

Can Qin on Hugging Face Daily Papers: 10 papers, 2 in the top 3 of their day, 405 upvotes.

  1. UNIDOC-BENCH: A Unified Benchmark for Document-Centric Multimodal RAG 15 upvotes, #25 of 2025-10-10
  2. CoDA: Coding LM via Diffusion Adaptation 39 upvotes, #6 of 2025-10-08
  3. When Tokens Talk Too Much: A Survey of Multimodal Long-Context Token Compression across Images, Videos, and Audios 24 upvotes, #4 of 2025-07-28
  4. VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents 15 upvotes, #12 of 2025-07-08
  5. HoliTom: Holistic Token Merging for Fast Video Large Language Models 18 upvotes, #24 of 2025-05-28
  6. BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset 80 upvotes, #1 of 2025-05-15
  7. Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models 23 upvotes, #15 of 2025-03-21
  8. xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs 14 upvotes, #4 of 2024-10-23
  9. xGen-VideoSyn-1: High-fidelity Text-to-Video Synthesis with Compressed Representations 32 upvotes, #5 of 2024-08-23
  10. xGen-MM (BLIP-3): A Family of Open Large Multimodal Models 91 upvotes, #1 of 2024-08-19

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.