Cascade Speculative Drafting for Even Faster LLM Inference

Ziyi Chen, Xiaocong Yang, Jiacheng Lin, chenkai sun, Jie Huang, Kevin Chen-Chuan Chang

Cascade Speculative Drafting for Even Faster LLM Inference: 9 upvotes on Hugging Face Daily Papers, #12 of 17 papers on 2023-12-19. Day-by-day upvote history.

Speculative decoding enhances the efficiency of large language models (LLMs) by leveraging a draft model to draft for a larger target model to review. However, drafting in speculative decoding involves slow autoregressive generation and generating tokens of different importance with the same time allocation. These two inefficiencies lead to its suboptimal performance. To address this issue, we introduce Cascade Speculative Drafting (CS. Drafting), a novel approach that employs two types of cascades. The Vertical Cascade eliminates autoregressive generation from neural models. The Horizontal Cascade constitutes efficient time allocation in drafting with its optimality supported by our theoretical analysis. Combining both cascades, our CS. Drafting algorithm has achieved up to 72 percent additional speedup over speculative decoding in our experiments while keeping the same output distribution.

Paper page on Hugging Face · arXiv

Data: hysts-bot-data/daily-papers-stats and the Daily Papers API. Open data: tardellirs/paper-pulse-data. Sister project: Model Pulse, the download history of every model on the Hub.