Yang Shi

Yang Shi on Hugging Face Daily Papers: 38 papers, 2 in the top 3 of their day, 1,468 upvotes.

  1. Think Before You Score: Thinking Reward Model for Visual Generation 101 upvotes, #11 of 2026-09-30
  2. HiRAE: Hierarchical Representation Autoencoding with Residual Budgets 25 upvotes, #38 of 2026-09-30
  3. FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory 14 upvotes, #18 of 2026-08-06
  4. OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models 26 upvotes, #11 of 2026-08-05
  5. RefCaptioner: Multi-Reference Image-Grounded Video Captioning 30 upvotes, #14 of 2026-07-31
  6. Beacon: Knowing When and How to Perform Agentic Visual Reasoning 54 upvotes, #8 of 2026-07-31
  7. Flux-OPD: On-Policy Distillation with Evolving Contexts 44 upvotes, #10 of 2026-07-31
  8. MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation 34 upvotes, #8 of 2026-07-17
  9. KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation 42 upvotes, #6 of 2026-07-17
  10. DOPD: Dual On-policy Distillation 103 upvotes, #2 of 2026-07-01
  11. LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV 38 upvotes, #8 of 2026-05-27
  12. LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning 46 upvotes, #8 of 2026-05-22
  13. MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation 14 upvotes, #15 of 2026-05-20
  14. Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos 22 upvotes, #12 of 2026-05-20
  15. Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling 32 upvotes, #9 of 2026-05-14
  16. Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization 33 upvotes, #10 of 2026-05-13
  17. OpenWorldLib: A Unified Codebase and Definition of Advanced World Models 200 upvotes, #2 of 2026-04-07
  18. Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence? 36 upvotes, #5 of 2026-04-06
  19. VTC-Bench: Evaluating Agentic Multimodal Models via Compositional Visual Tool Chaining 21 upvotes, #14 of 2026-03-20
  20. BrowseComp-V^3: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agents 8 upvotes, #15 of 2026-02-17
  21. OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models 46 upvotes, #5 of 2026-02-05
  22. Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks 46 upvotes, #7 of 2026-02-04
  23. CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation 30 upvotes, #12 of 2026-01-16
  24. GRAN-TED: Generating Robust, Aligned, and Nuanced Text Embedding for Diffusion Models 24 upvotes, #9 of 2025-12-30
  25. Hybrid Attribution Priors for Explainable and Robust Model Training 2 upvotes, #28 of 2025-12-18
  26. Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling 40 upvotes, #4 of 2025-12-17
  27. Monet: Reasoning in Latent Visual Space Beyond Images and Language 15 upvotes, #7 of 2025-11-27
  28. When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs 24 upvotes, #5 of 2025-11-05
  29. MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning 19 upvotes, #13 of 2025-10-20
  30. AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration 28 upvotes, #10 of 2025-10-14
  31. OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing 49 upvotes, #4 of 2025-09-30
  32. RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark 44 upvotes, #6 of 2025-09-30
  33. BaseReward: A Strong Baseline for Multimodal Reward Model 21 upvotes, #5 of 2025-09-22
  34. MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios 38 upvotes, #12 of 2025-05-28
  35. Mavors: Multi-granularity Video Representation for Multimodal Large Language Model 30 upvotes, #7 of 2025-04-15
  36. MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models 13 upvotes, #9 of 2025-04-07
  37. MM-RLHF: The Next Step Forward in Multimodal LLM Alignment 30 upvotes, #6 of 2025-02-17
  38. EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents 5 upvotes, #16 of 2025-01-24

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.