Shoufa Chen

Shoufa Chen on Hugging Face Daily Papers: 14 papers, 3 in the top 3 of their day, 483 upvotes.

  1. TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents 47 upvotes, #6 of 2026-06-30
  2. WavFlow: Audio Generation in Waveform Space 10 upvotes, #25 of 2026-05-19
  3. Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation 8 upvotes, #15 of 2026-04-28
  4. HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming 20 upvotes, #8 of 2025-12-25
  5. TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models 60 upvotes, #5 of 2025-12-02
  6. PixelFlow: Pixel-Space Generative Models with Flow 17 upvotes, #6 of 2025-04-14
  7. FlashVideo:Flowing Fidelity to Detail for Efficient High-Resolution Video Generation 22 upvotes, #7 of 2025-02-10
  8. Goku: Flow Based Video Generative Foundation Models 82 upvotes, #2 of 2025-02-10
  9. ControlAR: Controllable Image Generation with Autoregressive Models 7 upvotes, #7 of 2024-10-09
  10. Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation 57 upvotes, #1 of 2024-06-11
  11. GenTron: Delving Deep into Diffusion Transformers for Image and Video Generation 13 upvotes, #7 of 2023-12-08
  12. GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest 13 upvotes, #3 of 2023-07-10
  13. Going Denser with Open-Vocabulary Part Segmentation 2 upvotes, #12 of 2023-05-19
  14. InternChat: Solving Vision-Centric Tasks by Interacting with Chatbots Beyond Language 5 upvotes, #5 of 2023-05-10

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.