xiangan

xiangan on Hugging Face Daily Papers: 15 papers, 0 in the top 3 of their day, 329 upvotes.

  1. StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding 9 upvotes, #24 of 2026-08-18
  2. Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model 35 upvotes, #5 of 2026-07-29
  3. LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence 27 upvotes, #11 of 2026-05-27
  4. 4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding 17 upvotes, #17 of 2026-05-11
  5. OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence 46 upvotes, #5 of 2026-02-16
  6. DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset 36 upvotes, #7 of 2026-01-16
  7. ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder 9 upvotes, #18 of 2025-10-22
  8. UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning 11 upvotes, #18 of 2025-10-16
  9. LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training 38 upvotes, #8 of 2025-09-29
  10. Gradient-Attention Guided Dual-Masking Synergetic Framework for Robust Text-based Person Retrieval 7 upvotes, #15 of 2025-09-12
  11. ForCenNet: Foreground-Centric Network for Document Image Rectification 11 upvotes, #14 of 2025-07-29
  12. Region-based Cluster Discrimination for Visual Representation Learning 17 upvotes, #10 of 2025-07-29
  13. RealSyn: An Effective and Scalable Multimodal Interleaved Document Transformation Paradigm 15 upvotes, #14 of 2025-02-19
  14. ORID: Organ-Regional Information Driven Framework for Radiology Report Generation 2 upvotes, #11 of 2024-11-21
  15. IDAdapter: Learning Mixed Features for Tuning-Free Personalization of Text-to-Image Models 20 upvotes, #6 of 2024-03-21

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.