Ziyu Guo

Ziyu Guo on Hugging Face Daily Papers: 11 papers, 3 in the top 3 of their day, 325 upvotes.

  1. ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both 19 upvotes, #20 of 2026-05-15
  2. Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMs 19 upvotes, #9 of 2025-08-21
  3. T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT 39 upvotes, #3 of 2025-05-02
  4. Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step 31 upvotes, #4 of 2025-01-24
  5. IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models 13 upvotes, #8 of 2025-01-24
  6. MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines 33 upvotes, #3 of 2024-09-20
  7. SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners 25 upvotes, #5 of 2024-08-30
  8. MAVIS: Mathematical Visual Instruction Tuning 26 upvotes, #5 of 2024-07-12
  9. MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? 45 upvotes, #1 of 2024-03-22
  10. ImageBind-LLM: Multi-modality Instruction Tuning 17 upvotes, #7 of 2023-09-08
  11. Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following 13 upvotes, #7 of 2023-09-04

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.