TL;DR: Too Long, Do Re-weighting for Effcient LLM Reasoning Compression
Zhongzhi Li, Xiao Liu, Zihao Tang, Lei Ji, PeijieWang, Haotian Xu, Xing W, Haizhen Huang, deng, Ying Nian Wu, Yeyun Gong, Zhijiang Guo, Xiao Liang, Fei Yin, Cheng-Lin Liu
TL;DR: Too Long, Do Re-weighting for Effcient LLM Reasoning Compression: 5 upvotes on Hugging Face Daily Papers, #35 of 50 papers on 2025-06-04. Day-by-day upvote history.
Large Language Models (LLMs) have recently achieved remarkable progress by leveraging Reinforcement Learning and extended Chain-of-Thought (CoT) techniques. However, the challenge of performing efficient language reasoning--especially during inference with extremely long outputs--has drawn increasing attention from the research community. In this work, we propose a dynamic ratio-based training pipeline that does not rely on sophisticated data annotations or interpolation between multiple models. We continuously balance the weights between the model's System-1 and System-2 data to eliminate redundant reasoning processes while preserving the model's reasoning capability. We validate our approach across models on DeepSeek-R1-Distill-7B and DeepSeek-R1-Distill-14B and on a diverse set of benchmarks with varying difficulty levels. Our method significantly reduces the number of output tokens by nearly 40% while maintaining the accuracy of the reasoning. Our code and data will be available soon.
Paper page on Hugging Face · arXiv
Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.