Yixiao Ge
Yixiao Ge on Hugging Face Daily Papers: 22 papers, 3 in the top 3 of their day, 482 upvotes.
- GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning 26 upvotes, #7 of 2025-06-25
- Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? 26 upvotes, #17 of 2025-05-28
- AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction 57 upvotes, #4 of 2025-04-03
- Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1 36 upvotes, #4 of 2025-04-02
- GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers 15 upvotes, #9 of 2025-03-27
- Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation 13 upvotes, #7 of 2024-12-10
- Moto: Latent Motion Token as the Bridging Language for Robot Manipulation 20 upvotes, #7 of 2024-12-09
- Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation 23 upvotes, #3 of 2024-09-09
- SEED-Story: Multimodal Long Story Generation with Large Language Model 18 upvotes, #7 of 2024-07-12
- VoCo-LLaMA: Towards Vision Compression with Large Language Models 28 upvotes, #5 of 2024-06-19
- Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots 15 upvotes, #5 of 2024-05-14
- SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension 7 upvotes, #10 of 2024-04-26
- SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation 16 upvotes, #7 of 2024-04-23
- YOLO-World: Real-Time Open-Vocabulary Object Detection 44 upvotes, #2 of 2024-01-31
- Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities 13 upvotes, #8 of 2024-01-26
- Towards A Better Metric for Text-to-Video Generation 15 upvotes, #7 of 2024-01-17
- Cached Transformers: Improving Transformers with Differentiable Memory Cache 13 upvotes, #7 of 2023-12-21
- VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation 7 upvotes, #19 of 2023-12-15
- Planting a SEED of Vision in Large Language Model 12 upvotes, #7 of 2023-07-18
- DreamDiffusion: Generating High-Quality Images from Brain EEG Signals 32 upvotes, #3 of 2023-06-30
- GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction 5 upvotes, #4 of 2023-05-31
- Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models 6 upvotes, #4 of 2023-05-30
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.