Haiwen Diao
Haiwen Diao on Hugging Face Daily Papers: 22 papers, 5 in the top 3 of their day, 1,406 upvotes.
- Think Before You Score: Thinking Reward Model for Visual Generation 101 upvotes, #11 of 2026-09-30
- Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning 49 upvotes, #13 of 2026-09-29
- VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control 41 upvotes, #17 of 2026-09-18
- SenseNova-U1.5: Towards Native Unified Visual Intelligence 262 upvotes, #2 of 2026-09-11
- Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System 23 upvotes, #10 of 2026-09-02
- V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning 15 upvotes, #15 of 2026-08-27
- VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 266 upvotes, #1 of 2026-08-27
- Vision as Unified Multimodal Generation 46 upvotes, #6 of 2026-07-08
- Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion 21 upvotes, #10 of 2026-06-17
- From Pixels to Words -- Towards Native One-Vision Models at Scale 72 upvotes, #4 of 2026-05-28
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture 184 upvotes, #1 of 2026-05-13
- VISTA-Bench: Do Vision-Language Models Really Understand Visualized Text as Well as Pure Text? 1 upvotes, #49 of 2026-02-11
- DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation 68 upvotes, #4 of 2026-01-30
- The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding 61 upvotes, #2 of 2025-12-23
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale 64 upvotes, #5 of 2025-10-17
- Visual Jigsaw Post-Training Improves MLLMs 34 upvotes, #11 of 2025-09-30
- End-to-End Vision Tokenizer Tuning 20 upvotes, #8 of 2025-05-16
- EVEv2: Improved Baselines for Encoder-Free Vision-Language Models 11 upvotes, #14 of 2025-02-11
- Autoregressive Video Generation without Vector Quantization 13 upvotes, #10 of 2024-12-19
- SHERL: Synthesizing High Accuracy and Efficient Memory for Resource-Limited Transfer Learning 4 upvotes, #12 of 2024-07-16
- DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception 17 upvotes, #10 of 2024-07-12
- Unveiling Encoder-Free Vision-Language Models 45 upvotes, #1 of 2024-07-08
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.