Xiaoyi Dong
Xiaoyi Dong on Hugging Face Daily Papers: 28 papers, 17 in the top 3 of their day, 1,352 upvotes.
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience 46 upvotes, #6 of 2025-08-07
- Beyond Fixed: Variable-Length Denoising for Diffusion Large Language Models 61 upvotes, #2 of 2025-08-04
- SeC: Advancing Complex Video Object Segmentation via Progressive Concept Construction 37 upvotes, #7 of 2025-07-22
- ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing 25 upvotes, #8 of 2025-06-25
- SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation 35 upvotes, #4 of 2025-02-20
- Light-A-Video: Training-free Video Relighting via Progressive Light Fusion 37 upvotes, #6 of 2025-02-13
- VideoRoPE: What Makes for Good Video Rotary Position Embedding? 60 upvotes, #3 of 2025-02-10
- InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model 39 upvotes, #6 of 2025-01-22
- OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding? 36 upvotes, #5 of 2025-01-13
- BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning 34 upvotes, #3 of 2025-01-07
- Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction 32 upvotes, #4 of 2025-01-07
- InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions 89 upvotes, #1 of 2024-12-13
- X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models 61 upvotes, #1 of 2024-12-03
- Open-Sora Plan: Open-Source Large Video Generation Model 30 upvotes, #3 of 2024-12-03
- MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models 34 upvotes, #1 of 2024-10-24
- PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction 42 upvotes, #1 of 2024-10-23
- SAM2Long: Enhancing SAM 2 for Long Video Segmentation with a Training-Free Memory Tree 61 upvotes, #2 of 2024-10-22
- Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate 36 upvotes, #7 of 2024-10-10
- BroadWay: Boost Your Text-to-Video Generation Model in a Training-free Way 10 upvotes, #24 of 2024-10-10
- VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models 11 upvotes, #6 of 2024-07-17
- InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 84 upvotes, #1 of 2024-07-04
- MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs 52 upvotes, #1 of 2024-06-18
- MotionClone: Training-Free Motion Cloning for Controllable Video Generation 38 upvotes, #3 of 2024-06-13
- ShareGPT4Video: Improving Video Understanding and Generation with Better Captions 61 upvotes, #1 of 2024-06-07
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites 47 upvotes, #1 of 2024-04-26
- InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD 24 upvotes, #3 of 2024-04-10
- InternLM2 Technical Report 22 upvotes, #3 of 2024-03-27
- InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 55 upvotes, #1 of 2024-01-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.