Zhaopeng Tu
Zhaopeng Tu on Hugging Face Daily Papers: 19 papers, 4 in the top 3 of their day, 438 upvotes.
- Too Good to be Bad: On the Failure of LLMs to Role-Play Villains 50 upvotes, #1 of 2025-11-10
- BatonVoice: An Operationalist Framework for Enhancing Controllable Speech Synthesis with Linguistic Intelligence from LLMs 3 upvotes, #24 of 2025-10-02
- SCAN: Self-Denoising Monte Carlo Annotation for Robust Process Reward Learning 0 upvotes, #41 of 2025-09-23
- CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models 28 upvotes, #5 of 2025-09-11
- RLVER: Reinforcement Learning with Verifiable Emotion Rewards for Empathetic Agents 31 upvotes, #7 of 2025-07-09
- DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning 15 upvotes, #24 of 2025-05-30
- Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms 14 upvotes, #29 of 2025-05-28
- Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training 9 upvotes, #23 of 2025-05-21
- Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models 25 upvotes, #3 of 2025-05-09
- SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning 15 upvotes, #5 of 2025-04-29
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning 11 upvotes, #17 of 2025-04-16
- Expanding RL with Verifiable Rewards Across Diverse Domains 17 upvotes, #11 of 2025-04-01
- Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs 51 upvotes, #2 of 2025-01-31
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs 27 upvotes, #5 of 2024-12-31
- Critical Tokens Matter: Token-Level Contrastive Estimation Enhence LLM's Reasoning Capability 47 upvotes, #2 of 2024-12-04
- Draft Model Knows When to Stop: A Self-Verification Length Policy for Speculative Decoding 6 upvotes, #17 of 2024-11-28
- Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training 5 upvotes, #14 of 2024-07-15
- Leveraging Word Guessing Games to Assess the Intelligence of Large Language Models 8 upvotes, #11 of 2023-11-01
- Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration 16 upvotes, #9 of 2023-06-16
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.