Yixiao Ge

Yixiao Ge on Hugging Face Daily Papers: 22 papers, 3 in the top 3 of their day, 482 upvotes.

  1. GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning 26 upvotes, #7 of 2025-06-25
  2. Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning? 26 upvotes, #17 of 2025-05-28
  3. AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction 57 upvotes, #4 of 2025-04-03
  4. Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1 36 upvotes, #4 of 2025-04-02
  5. GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers 15 upvotes, #9 of 2025-03-27
  6. Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation 13 upvotes, #7 of 2024-12-10
  7. Moto: Latent Motion Token as the Bridging Language for Robot Manipulation 20 upvotes, #7 of 2024-12-09
  8. Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation 23 upvotes, #3 of 2024-09-09
  9. SEED-Story: Multimodal Long Story Generation with Large Language Model 18 upvotes, #7 of 2024-07-12
  10. VoCo-LLaMA: Towards Vision Compression with Large Language Models 28 upvotes, #5 of 2024-06-19
  11. Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots 15 upvotes, #5 of 2024-05-14
  12. SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension 7 upvotes, #10 of 2024-04-26
  13. SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation 16 upvotes, #7 of 2024-04-23
  14. YOLO-World: Real-Time Open-Vocabulary Object Detection 44 upvotes, #2 of 2024-01-31
  15. Multimodal Pathway: Improve Transformers with Irrelevant Data from Other Modalities 13 upvotes, #8 of 2024-01-26
  16. Towards A Better Metric for Text-to-Video Generation 15 upvotes, #7 of 2024-01-17
  17. Cached Transformers: Improving Transformers with Differentiable Memory Cache 13 upvotes, #7 of 2023-12-21
  18. VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation 7 upvotes, #19 of 2023-12-15
  19. Planting a SEED of Vision in Large Language Model 12 upvotes, #7 of 2023-07-18
  20. DreamDiffusion: Generating High-Quality Images from Brain EEG Signals 32 upvotes, #3 of 2023-06-30
  21. GPT4Tools: Teaching Large Language Model to Use Tools via Self-instruction 5 upvotes, #4 of 2023-05-31
  22. Mix-of-Show: Decentralized Low-Rank Adaptation for Multi-Concept Customization of Diffusion Models 6 upvotes, #4 of 2023-05-30

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.