Haiwen Diao

Haiwen Diao on Hugging Face Daily Papers: 22 papers, 5 in the top 3 of their day, 1,406 upvotes.

  1. Think Before You Score: Thinking Reward Model for Visual Generation 101 upvotes, #11 of 2026-09-30
  2. Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning 49 upvotes, #13 of 2026-09-29
  3. VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control 41 upvotes, #17 of 2026-09-18
  4. SenseNova-U1.5: Towards Native Unified Visual Intelligence 262 upvotes, #2 of 2026-09-11
  5. Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System 23 upvotes, #10 of 2026-09-02
  6. V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning 15 upvotes, #15 of 2026-08-27
  7. VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 266 upvotes, #1 of 2026-08-27
  8. Vision as Unified Multimodal Generation 46 upvotes, #6 of 2026-07-08
  9. Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion 21 upvotes, #10 of 2026-06-17
  10. From Pixels to Words -- Towards Native One-Vision Models at Scale 72 upvotes, #4 of 2026-05-28
  11. SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture 184 upvotes, #1 of 2026-05-13
  12. VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? 1 upvotes, #49 of 2026-02-11
  13. DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation 68 upvotes, #4 of 2026-01-30
  14. The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding 61 upvotes, #2 of 2025-12-23
  15. From Pixels to Words -- Towards Native Vision-Language Primitives at Scale 64 upvotes, #5 of 2025-10-17
  16. Visual Jigsaw Post-Training Improves MLLMs 34 upvotes, #11 of 2025-09-30
  17. End-to-End Vision Tokenizer Tuning 20 upvotes, #8 of 2025-05-16
  18. EVEv2: Improved Baselines for Encoder-Free Vision-Language Models 11 upvotes, #14 of 2025-02-11
  19. Autoregressive Video Generation without Vector Quantization 13 upvotes, #10 of 2024-12-19
  20. SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning 4 upvotes, #12 of 2024-07-16
  21. DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception 17 upvotes, #10 of 2024-07-12
  22. Unveiling Encoder-Free Vision-Language Models 45 upvotes, #1 of 2024-07-08

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.