Shiyu Huang
Shiyu Huang on Hugging Face Daily Papers: 8 papers, 2 in the top 3 of their day, 477 upvotes.
- X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation 42 upvotes, #5 of 2026-09-11
- GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning 176 upvotes, #1 of 2025-07-02
- MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models 40 upvotes, #4 of 2025-01-08
- VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation 18 upvotes, #4 of 2025-01-06
- DreamPolish: Domain Score Distillation With Progressive Geometry Generation 9 upvotes, #7 of 2024-11-06
- CogVLM2: Visual Language Models for Image and Video Understanding 55 upvotes, #2 of 2024-08-30
- CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer 31 upvotes, #5 of 2024-08-13
- SwiftSage: A Generative Agent with Fast and Slow Thinking for Complex Interactive Tasks 4 upvotes, #8 of 2023-05-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.