Hongcheng Gao

Hongcheng Gao on Hugging Face Daily Papers: 20 papers, 11 in the top 3 of their day, 1,925 upvotes.

  1. Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence 107 upvotes, #6 of 2026-09-29
  2. StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding 9 upvotes, #24 of 2026-08-18
  3. PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models 18 upvotes, #11 of 2026-07-29
  4. Kimi K3: Open Frontier Intelligence 458 upvotes, #1 of 2026-07-28
  5. SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks 42 upvotes, #8 of 2026-06-09
  6. OpenWorldLib: A Unified Codebase and Definition of Advanced World Models 200 upvotes, #2 of 2026-04-07
  7. Research on World Models Is Not Merely Injecting World Knowledge into Specific Tasks 46 upvotes, #7 of 2026-02-04
  8. Kimi K2.5: Visual Agentic Intelligence 219 upvotes, #2 of 2026-02-03
  9. Pixels, Patterns, but No Poetry: To See The World like Humans 62 upvotes, #2 of 2025-07-24
  10. G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning 12 upvotes, #31 of 2025-05-27
  11. GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning 50 upvotes, #3 of 2025-05-19
  12. FlowReasoner: Reinforcing Query-Level Meta-Agents 46 upvotes, #3 of 2025-04-22
  13. Kimi-VL Technical Report 113 upvotes, #1 of 2025-04-11
  14. Efficient Inference for Large Reasoning Models: A Survey 45 upvotes, #5 of 2025-04-01
  15. Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation 29 upvotes, #4 of 2025-03-26
  16. GuardReasoner: Towards Reasoning-based LLM Safeguards 79 upvotes, #1 of 2025-01-31
  17. Kimi k1.5: Scaling Reinforcement Learning with LLMs 76 upvotes, #2 of 2025-01-23
  18. Spider2-V: How Far Are Multimodal Agents From Automating Data Science and Engineering Workflows? 5 upvotes, #10 of 2024-07-16
  19. McEval: Massively Multilingual Code Evaluation 38 upvotes, #3 of 2024-06-12
  20. Generative Pretraining in Multimodality 23 upvotes, #3 of 2023-07-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.