Xiangtai Li

Xiangtai Li on Hugging Face Daily Papers: 37 papers, 8 in the top 3 of their day, 1,027 upvotes.

  1. UniVR: Thinking in Visual Space for Unified Visual Reasoning 32 upvotes, #10 of 2026-07-17
  2. Prism: Efficient Test-Time Scaling via Hierarchical Search and Self-Verification for Discrete Diffusion Language Models 3 upvotes, #53 of 2026-02-03
  3. SAMTok: Representing Any Mask with Two Words 41 upvotes, #9 of 2026-01-23
  4. Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future 12 upvotes, #18 of 2025-12-18
  5. RecTok: Reconstruction Distillation along Rectified Flow 4 upvotes, #28 of 2025-12-16
  6. MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation 64 upvotes, #6 of 2025-11-18
  7. Visual Spatial Tuning 46 upvotes, #2 of 2025-11-10
  8. Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark 32 upvotes, #9 of 2025-10-31
  9. PairUni: Pairwise Training for Unified Multimodal Language Models 13 upvotes, #16 of 2025-10-30
  10. Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence 52 upvotes, #3 of 2025-10-24
  11. From Masks to Worlds: A Hitchhiker's Guide to World Models 6 upvotes, #20 of 2025-10-24
  12. Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs 35 upvotes, #9 of 2025-10-22
  13. DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training 29 upvotes, #9 of 2025-10-14
  14. Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology 44 upvotes, #3 of 2025-07-11
  15. CyberV: Cybernetics for Test-time Scaling in Video Understanding 4 upvotes, #32 of 2025-06-10
  16. DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers 7 upvotes, #24 of 2025-06-05
  17. MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query 3 upvotes, #35 of 2025-06-04
  18. Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model 14 upvotes, #25 of 2025-05-30
  19. PixelThink: Towards Efficient Chain-of-Pixel Reasoning 3 upvotes, #44 of 2025-05-29
  20. On Path to Multimodal Generalist: General-Level and General-Bench 72 upvotes, #1 of 2025-05-08
  21. DC-SAM: In-Context Segment Anything in Images and Videos via Dual Consistency 7 upvotes, #12 of 2025-04-28
  22. The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer 15 upvotes, #10 of 2025-04-16
  23. PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild 6 upvotes, #23 of 2025-04-16
  24. Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding 28 upvotes, #6 of 2025-04-16
  25. An Empirical Study of GPT-4o Image Generation Capabilities 59 upvotes, #4 of 2025-04-09
  26. Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos 40 upvotes, #4 of 2025-01-08
  27. EMOv2: Pushing 5M Vision Model Frontier 13 upvotes, #15 of 2024-12-11
  28. DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation 43 upvotes, #3 of 2024-12-11
  29. HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing 13 upvotes, #15 of 2024-12-06
  30. Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis 44 upvotes, #2 of 2024-10-14
  31. Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language 9 upvotes, #16 of 2024-07-02
  32. OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding 50 upvotes, #1 of 2024-06-28
  33. MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning 18 upvotes, #6 of 2024-06-26
  34. MotionBooth: Motion-Aware Customized Text-to-Video Generation 17 upvotes, #8 of 2024-06-26
  35. Face Adapter for Pre-Trained Diffusion Models with Fine-Grained ID and Attribute Control 20 upvotes, #4 of 2024-05-22
  36. Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively 23 upvotes, #5 of 2024-01-08
  37. MosaicFusion: Diffusion Models as Data Augmenters for Large Vocabulary Instance Segmentation 9 upvotes, #3 of 2023-09-25

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.