Speculative Decoding Playground

New here? Speculative decoding in 4 beats

Skip to the run below if you know the idea
problem

A big LLM writes one token per forward pass, and each pass mostly waits on reading billions of weights from memory. Writing 100 tokens takes 100 slow passes.

1 · draft

A small, fast model with the same tokenizer guesses the next γ tokens, one cheap pass per guess.

2 · verify

The big model checks all γ guesses in a single pass. Scoring γ+1 positions costs about the same as scoring 1, because the weights are read once.

3 · acceptfix

Keep the guesses the big model agrees with. At the first disagreement the big model supplies its own token instead. Every round adds 1 to γ+1 tokens, and with the sampling math it outputs exactly what the big model alone would.

Output, colored by where each token came from

Click any token to inspect the round that produced it
prompt ✓ draft token accepted ↺ correction (resampled by target after a rejection) ★ target token (bonus or first token)
Run something, or pick a recorded run.
Show the prompt tokens as the model sees them

Round inspector

Keys: ← → step · [ ] round · space play

The code that just ran

speculative.py: a complete implementation in ~100 lines; highlighted lines follow the inspector

Every round at a glance

Tokens committed per round · click a bar to inspect it

Where the time went

Speculative vs. target-only output

Lab 1: why “accept with probability p/q” is exact

A 5-token toy vocabulary · drag the sliders

Lab 2: when does it actually get faster?

Leviathan et al. 2023, eq. 1 & thm 3.8

Pitfalls when you implement it yourself

Easy to get subtly wrong: the output still looks fine but is no longer the big model's