Yuxuan Zhang
Yuxuan Zhang on Hugging Face Daily Papers: 15 papers, 6 in the top 3 of their day, 2,450 upvotes.
- S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? 39 upvotes, #8 of 2026-09-03
- Aspire: Can Models Self-Evolve from Vague Goals? 228 upvotes, #3 of 2026-09-03
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? 264 upvotes, #2 of 2026-09-03
- VGI-BENCH: Probing Visual Intelligence in Video Generation Models 177 upvotes, #2 of 2026-08-27
- Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models 107 upvotes, #1 of 2026-07-15
- Learning from the Self-future: On-policy Self-distillation for dLLMs 74 upvotes, #3 of 2026-06-17
- Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback 15 upvotes, #17 of 2026-06-12
- OpenSkill: Open-World Self-Evolution for LLM Agents 27 upvotes, #9 of 2026-06-08
- MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection 22 upvotes, #12 of 2026-06-03
- ClawBench: Can AI Agents Complete Everyday Online Tasks? 255 upvotes, #3 of 2026-04-10
- Watch Before You Answer: Learning from Visually Grounded Post-Training 35 upvotes, #11 of 2026-04-08
- A Rigorous Benchmark with Multidimensional Evaluation for Deep Research Agents: From Answers to Reports 18 upvotes, #17 of 2025-10-03
- VideoScore2: Think before You Score in Generative Video Evaluation 22 upvotes, #18 of 2025-09-30
- StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs 18 upvotes, #19 of 2025-05-27
- ScholarCopilot: Training Large Language Models for Academic Writing with Accurate Citations 38 upvotes, #5 of 2025-04-03
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.