kcz

kcz on Hugging Face Daily Papers: 13 papers, 5 in the top 3 of their day, 761 upvotes.

  1. StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding 9 upvotes, #24 of 2026-08-18
  2. Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model 35 upvotes, #5 of 2026-07-29
  3. Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing 72 upvotes, #5 of 2026-07-22
  4. ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning 34 upvotes, #10 of 2026-05-26
  5. Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling 89 upvotes, #2 of 2026-05-01
  6. LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling 150 upvotes, #2 of 2025-12-02
  7. OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe 88 upvotes, #3 of 2025-11-24
  8. UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning 11 upvotes, #18 of 2025-10-16
  9. Large Multi-modal Models Can Interpret Features in Large Multi-modal Models 14 upvotes, #7 of 2024-11-25
  10. MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures 72 upvotes, #2 of 2024-10-18
  11. LLaVA-OneVision: Easy Visual Task Transfer 50 upvotes, #2 of 2024-08-07
  12. LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models 28 upvotes, #5 of 2024-07-18
  13. Long Context Transfer from Language to Vision 30 upvotes, #4 of 2024-06-25

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.