Jieyu Zhang

Jieyu Zhang on Hugging Face Daily Papers: 21 papers, 6 in the top 3 of their day, 621 upvotes.

  1. Vision-Language Grounding as Bidirectional Concept Correspondence 7 upvotes, #26 of 2026-08-11
  2. MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction 50 upvotes, #1 of 2026-06-18
  3. Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models 119 upvotes, #1 of 2026-06-08
  4. You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass 11 upvotes, #20 of 2026-04-15
  5. WildDet3D: Scaling Promptable 3D Detection in the Wild 239 upvotes, #1 of 2026-04-13
  6. MolmoPoint: Better Pointing for VLMs with Grounding Tokens 8 upvotes, #26 of 2026-03-31
  7. Video-Based Reward Modeling for Computer-Use Agents 41 upvotes, #4 of 2026-03-13
  8. Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos 4 upvotes, #29 of 2026-03-03
  9. Theory of Space: Can Foundation Models Construct Spatial Beliefs through Active Exploration? 21 upvotes, #19 of 2026-02-10
  10. Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding 26 upvotes, #15 of 2026-01-16
  11. SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning 16 upvotes, #12 of 2025-12-18
  12. MolmoAct: Action Reasoning Models that can Reason in Space 37 upvotes, #7 of 2025-08-12
  13. CoAct-1: Computer-using Agents with Coding as Actions 13 upvotes, #10 of 2025-08-08
  14. Spatial Mental Modeling from Limited Views 12 upvotes, #10 of 2025-06-30
  15. Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems 6 upvotes, #14 of 2025-05-07
  16. Discovering Knowledge Deficiencies of Language Models on Massive Knowledge Base 6 upvotes, #26 of 2025-04-02
  17. On the Trustworthiness of Generative Foundation Models: Guideline, Assessment, and Perspective 44 upvotes, #2 of 2025-02-20
  18. xGen-MM (BLIP-3): A Family of Open Large Multimodal Models 91 upvotes, #1 of 2024-08-19
  19. Task Me Anything 7 upvotes, #20 of 2024-06-18
  20. DataComp-LM: In search of the next generation of training sets for language models 33 upvotes, #2 of 2024-06-18
  21. SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models 9 upvotes, #7 of 2023-07-21

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.