When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models
Jaejun Shim, HyunJin Kim, Young Jin Kim, JinYeong Bak
When2Think: Learning Difficulty-Aware Length Control for Efficient Hybrid Reasoning Models: 54 upvotes on Hugging Face Daily Papers, #10 of 24 papers on 2026-09-18. Day-by-day upvote history.
Large Reasoning Models (LRMs) often overthink easy problems and underthink hard ones, leading to inefficient computation allocation. Existing methods regulate generated computation or select between direct answering and explicit reasoning, but do not jointly control whether}to reason and how much computation to allocate within reasoning. We call the resulting difficulty-dependent loss in accuracy under computation reduction the efficiency tax. We propose When2Think, an RLVR-based post-training framework for instance-adaptive computation allocation. Its core mechanism, Instance-level Difficulty-Aware Control (IDAC), uses cached reference statistics of success and token cost to modulate a correctness-gated efficiency bonus based on generated token count. Importance sampling supports exploration of Think and NoThink, while Batch-Wise Standardization constructs standardized advantages for critic-free optimization. The framework requires neither a learned reward model nor a learned critic, and offline reference caching avoids online reference-model queries during policy updates. On AIME24, When2Think improves Pass@3 by 10.0 percentage points while reducing token usage by 27.9% relative to the backbone.
Paper page on Hugging Face · arXiv
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.