Cao Yuhang

Cao Yuhang on Hugging Face Daily Papers: 17 papers, 10 in the top 3 of their day, 861 upvotes.

  1. Visual Agentic Reinforcement Fine-Tuning 31 upvotes, #6 of 2025-05-21
  2. MM-IFEngine: Towards Multimodal Instruction Following 31 upvotes, #6 of 2025-04-11
  3. HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance 11 upvotes, #10 of 2025-04-09
  4. Visual-RFT: Visual Reinforcement Fine-Tuning 62 upvotes, #2 of 2025-03-04
  5. SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation 35 upvotes, #4 of 2025-02-20
  6. VideoRoPE: What Makes for Good Video Rotary Position Embedding? 60 upvotes, #3 of 2025-02-10
  7. InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model 39 upvotes, #6 of 2025-01-22
  8. OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? 36 upvotes, #5 of 2025-01-13
  9. Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction 32 upvotes, #4 of 2025-01-07
  10. BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning 34 upvotes, #3 of 2025-01-07
  11. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions 89 upvotes, #1 of 2024-12-13
  12. MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models 34 upvotes, #1 of 2024-10-24
  13. PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction 42 upvotes, #1 of 2024-10-23
  14. SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree 61 upvotes, #2 of 2024-10-22
  15. InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 84 upvotes, #1 of 2024-07-04
  16. InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD 24 upvotes, #3 of 2024-04-10
  17. InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 55 upvotes, #1 of 2024-01-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.