Xuandong Zhao
Xuandong Zhao on Hugging Face Daily Papers: 13 papers, 1 in the top 3 of their day, 208 upvotes.
- SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks 51 upvotes, #3 of 2026-02-18
- Clipping-Free Policy Optimization for Large Language Models 3 upvotes, #53 of 2026-02-03
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces 32 upvotes, #11 of 2026-01-23
- InfoSynth: Information-Guided Benchmark Synthesis for LLMs 2 upvotes, #14 of 2026-01-05
- Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models 16 upvotes, #13 of 2025-07-11
- The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation 5 upvotes, #23 of 2025-07-09
- AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents 6 upvotes, #23 of 2025-06-18
- Learning to Reason without External Rewards 25 upvotes, #13 of 2025-05-27
- SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning 7 upvotes, #33 of 2025-05-23
- Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs 12 upvotes, #10 of 2025-04-08
- The Hidden Risks of Large Reasoning Models: A Safety Assessment of R1 5 upvotes, #26 of 2025-02-19
- Multimodal Situational Safety 8 upvotes, #26 of 2024-10-10
- Weak-to-Strong Jailbreaking on Large Language Models 16 upvotes, #8 of 2024-01-31
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.