Cihang Xie

Cihang Xie on Hugging Face Daily Papers: 20 papers, 3 in the top 3 of their day, 424 upvotes.

  1. In-Context Reinforcement Learning for Tool Use in Large Language Models 39 upvotes, #5 of 2026-03-12
  2. OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning 28 upvotes, #12 of 2025-09-03
  3. AHELM: A Holistic Evaluation of Audio-Language Models 9 upvotes, #11 of 2025-09-01
  4. GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset 20 upvotes, #9 of 2025-07-29
  5. OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning 20 upvotes, #7 of 2025-05-08
  6. Complex-Edit: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark 8 upvotes, #19 of 2025-04-18
  7. SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 26 upvotes, #5 of 2025-04-17
  8. ViLBench: A Suite for Vision-Language Process Reward Modeling 7 upvotes, #14 of 2025-03-27
  9. Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More 9 upvotes, #18 of 2025-02-10
  10. VHELM: A Holistic Evaluation of Vision Language Models 2 upvotes, #47 of 2024-10-10
  11. A Preliminary Study of o1 in Medicine: Are We Closer to an AI Doctor? 33 upvotes, #2 of 2024-09-24
  12. VideoLLaMB: Long-context Video Understanding with Recurrent Memory Bridges 26 upvotes, #8 of 2024-09-04
  13. MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for Medicine 23 upvotes, #4 of 2024-08-07
  14. VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models 22 upvotes, #6 of 2024-06-25
  15. What If We Recaption Billions of Web Images with LLaMA-3? 35 upvotes, #4 of 2024-06-13
  16. HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing 11 upvotes, #7 of 2024-04-16
  17. Scaling (Down) CLIP: A Comprehensive Analysis of Data, Architecture, and Training Strategies 24 upvotes, #3 of 2024-04-15
  18. Rejuvenating image-GPT as Strong Visual Representation Learners 6 upvotes, #20 of 2023-12-05
  19. CLIPA-v2: Scaling CLIP Training with 81.1% Zero-shot ImageNet Accuracy within a \10,000 Budget; An Extra 4,000 Unlocks 81.8% Accuracy 13 upvotes, #3 of 2023-06-28
  20. An Inverse Scaling Law for CLIP Training 3 upvotes, #6 of 2023-05-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.