Wenqi Zhang

Wenqi Zhang on Hugging Face Daily Papers: 21 papers, 5 in the top 3 of their day, 756 upvotes.

  1. Spatial-Interactor: Learning Spatial Reasoning through Interaction with the Observable Physical World 60 upvotes, #3 of 2026-09-24
  2. VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon 30 upvotes, #5 of 2026-07-06
  3. GFT: From Imitation to Reward Fine-Tuning with Unbiased Group Advantages and Dynamic Coefficient Rectification 23 upvotes, #9 of 2026-04-21
  4. UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization 6 upvotes, #23 of 2026-04-16
  5. KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation 47 upvotes, #12 of 2026-04-10
  6. GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts 27 upvotes, #13 of 2025-09-30
  7. EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering 26 upvotes, #16 of 2025-09-30
  8. Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models 16 upvotes, #9 of 2025-08-14
  9. OmniEAR: Benchmarking Agent Reasoning in Embodied Tasks 18 upvotes, #14 of 2025-08-12
  10. GUI-G^2: Gaussian Reward Modeling for GUI Grounding 122 upvotes, #1 of 2025-07-22
  11. TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence 11 upvotes, #18 of 2025-06-05
  12. SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation 14 upvotes, #15 of 2025-06-05
  13. ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models 11 upvotes, #33 of 2025-05-28
  14. Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning 23 upvotes, #12 of 2025-05-23
  15. Let LLMs Break Free from Overthinking via Self-Braking Tuning 23 upvotes, #12 of 2025-05-23
  16. Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency 10 upvotes, #7 of 2025-04-29
  17. Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks 21 upvotes, #9 of 2025-03-28
  18. VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding 75 upvotes, #3 of 2025-01-23
  19. 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining 91 upvotes, #1 of 2025-01-03
  20. Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model 37 upvotes, #3 of 2024-07-12
  21. VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs 28 upvotes, #8 of 2024-06-13

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.