📖 The AI Tool Bible
inference

Speculative Decoding

An inference speedup where a small draft model proposes several tokens ahead; a large model verifies them in parallel. Trades a bit of compute for ~2x throughput.

Related terms