Xuandong Zhao

Xuandong Zhao on Hugging Face Daily Papers: 13 papers, 1 in the top 3 of their day, 208 upvotes.

  1. SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks 51 upvotes, #3 of 2026-02-18
  2. Clipping-Free Policy Optimization for Large Language Models 3 upvotes, #53 of 2026-02-03
  3. Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces 32 upvotes, #11 of 2026-01-23
  4. InfoSynth: Information-Guided Benchmark Synthesis for LLMs 2 upvotes, #14 of 2026-01-05
  5. Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models 16 upvotes, #13 of 2025-07-11
  6. The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation 5 upvotes, #23 of 2025-07-09
  7. AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents 6 upvotes, #23 of 2025-06-18
  8. Learning to Reason without External Rewards 25 upvotes, #13 of 2025-05-27
  9. SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning 7 upvotes, #33 of 2025-05-23
  10. Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs 12 upvotes, #10 of 2025-04-08
  11. The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1 5 upvotes, #26 of 2025-02-19
  12. Multimodal Situational Safety 8 upvotes, #26 of 2024-10-10
  13. Weak-to-Strong Jailbreaking on Large Language Models 16 upvotes, #8 of 2024-01-31

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.