Cheng Qian

Cheng Qian on Hugging Face Daily Papers: 14 papers, 5 in the top 3 of their day, 665 upvotes.

  1. AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses 111 upvotes, #3 of 2026-08-13
  2. Trimming the Long-Tail of Visual World Modeling Evaluation 42 upvotes, #7 of 2026-06-30
  3. PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems 95 upvotes, #1 of 2026-06-23
  4. AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints 40 upvotes, #4 of 2026-06-05
  5. Advancing Creative Physical Intelligence in Large Multimodal Models 19 upvotes, #21 of 2026-05-28
  6. UserRL: Training Interactive User-Centric Agent via Reinforcement Learning 10 upvotes, #16 of 2025-09-26
  7. LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering 6 upvotes, #16 of 2025-09-12
  8. UserBench: An Interactive Gym Environment for User-Centric Agents 29 upvotes, #10 of 2025-08-12
  9. A Survey of Self-Evolving Agents: On Path to Artificial Super Intelligence 72 upvotes, #2 of 2025-07-29
  10. RM-R1: Reward Modeling as Reasoning 66 upvotes, #3 of 2025-05-06
  11. OTC: Optimal Tool Calls via Reinforcement Learning 33 upvotes, #5 of 2025-04-22
  12. ToolRL: Reward is All Tool Learning Needs 41 upvotes, #4 of 2025-04-22
  13. MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents 23 upvotes, #2 of 2025-03-05
  14. EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents 32 upvotes, #4 of 2025-02-14

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.