Zhengyang Tang
Zhengyang Tang on Hugging Face Daily Papers: 18 papers, 5 in the top 3 of their day, 844 upvotes.
- JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents 124 upvotes, #3 of 2026-07-28
- Training Open Models for Agentic Phone Use 16 upvotes, #17 of 2026-06-23
- GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine? 55 upvotes, #5 of 2026-06-17
- PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions 12 upvotes, #23 of 2026-06-16
- PhoneWorld: Scaling Phone-Use Agent Environments 8 upvotes, #44 of 2026-05-29
- Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents 2 upvotes, #63 of 2026-05-12
- Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows 40 upvotes, #6 of 2026-05-01
- Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoning 23 upvotes, #7 of 2026-04-20
- Do Phone-Use Agents Respect Your Privacy? 9 upvotes, #17 of 2026-04-02
- Kimi K2.5: Visual Agentic Intelligence 219 upvotes, #2 of 2026-02-03
- CALM Before the STORM: Unlocking Native Reasoning for Optimization Modeling 19 upvotes, #16 of 2025-10-09
- CoRT: Code-integrated Reasoning within Thinking 18 upvotes, #9 of 2025-06-12
- Learning from Peers in Reasoning Models 41 upvotes, #4 of 2025-05-13
- RealCritic: Towards Effectiveness-Driven Evaluation of Language Model Critiques 29 upvotes, #3 of 2025-01-27
- Enabling Scalable Oversight via Self-Evolving Critic 66 upvotes, #1 of 2025-01-13
- Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models 26 upvotes, #6 of 2024-10-15
- MathScale: Scaling Instruction Tuning for Mathematical Reasoning 13 upvotes, #6 of 2024-03-06
- Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models 52 upvotes, #2 of 2024-02-21
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.