New here? Speculative decoding in 4 beats
Skip to the run below if you know the ideaproblem
A big LLM writes one token per forward pass, and each pass mostly waits on reading billions of weights from memory. Writing 100 tokens takes 100 slow passes.
1 · draft
A small, fast model with the same tokenizer guesses the next γ tokens, one cheap pass per guess.
2 · verify
The big model checks all γ guesses in a single pass. Scoring γ+1 positions costs about the same as scoring 1, because the weights are read once.
3 · acceptfix
Keep the guesses the big model agrees with. At the first disagreement the big model supplies its own token instead. Every round adds 1 to γ+1 tokens, and with the sampling math it outputs exactly what the big model alone would.
Output, colored by where each token came from
Click any token to inspect the round that produced it
prompt
✓ draft token accepted
↺ correction (resampled by target after a rejection)
★ target token (bonus or first token)
Run something, or pick a recorded run.
Show the prompt tokens as the model sees them
Round inspector
Keys: ← → step · [ ] round · space play