Yongliang Shen

Yongliang Shen on Hugging Face Daily Papers: 27 papers, 6 in the top 3 of their day, 981 upvotes.

  1. ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents 141 upvotes, #1 of 2026-04-14
  2. KnowU-Bench: Towards Interactive, Proactive, and Personalized Mobile Agent Evaluation 47 upvotes, #12 of 2026-04-10
  3. SKILL0: In-Context Agentic Reinforcement Learning for Skill Internalization 92 upvotes, #4 of 2026-04-03
  4. Code-A1: Adversarial Evolving of Code LLM and Test LLM via Reinforcement Learning 10 upvotes, #20 of 2026-03-17
  5. How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities 22 upvotes, #8 of 2026-03-04
  6. InftyThink+: Effective and Efficient Infinite-Horizon Reasoning via Reinforcement Learning 12 upvotes, #16 of 2026-02-09
  7. IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video? 3 upvotes, #64 of 2025-09-30
  8. EasySteer: A Unified Framework for High-Performance and Extensible LLM Steering 26 upvotes, #16 of 2025-09-30
  9. GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts 27 upvotes, #13 of 2025-09-30
  10. UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning 45 upvotes, #2 of 2025-09-16
  11. Cooper: Co-Optimizing Policy and Reward Models in Reinforcement Learning for Large Language Models 16 upvotes, #9 of 2025-08-14
  12. Test-Time Reinforcement Learning for GUI Grounding via Region Consistency 20 upvotes, #12 of 2025-08-13
  13. Time Is a Feature: Exploiting Temporal Dynamics in Diffusion Language Models 34 upvotes, #6 of 2025-08-13
  14. OmniEAR: Benchmarking Agent Reasoning in Embodied Tasks 18 upvotes, #14 of 2025-08-12
  15. Hierarchical Budget Policy Optimization for Adaptive Reasoning 16 upvotes, #8 of 2025-07-25
  16. LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization 34 upvotes, #5 of 2025-07-25
  17. GUI-G^2: Gaussian Reward Modeling for GUI Grounding 122 upvotes, #1 of 2025-07-22
  18. TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence 11 upvotes, #18 of 2025-06-05
  19. SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation 14 upvotes, #15 of 2025-06-05
  20. ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models 11 upvotes, #33 of 2025-05-28
  21. Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning 23 upvotes, #12 of 2025-05-23
  22. Let LLMs Break Free from Overthinking via Self-Braking Tuning 23 upvotes, #12 of 2025-05-23
  23. VerifyBench: Benchmarking Reference-based Reward Systems for Large Language Models 17 upvotes, #13 of 2025-05-22
  24. Chain-of-Model Learning for Language Model 107 upvotes, #1 of 2025-05-20
  25. Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks 21 upvotes, #9 of 2025-03-28
  26. 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining 91 upvotes, #1 of 2025-01-03
  27. Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model 37 upvotes, #3 of 2024-07-12

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.