I wired speculative decoding into an inference engine I wrote myself. Across 18 test prompts the output matched the non-speculative run token for token. Correctness: perfect.
Throughput went from 92.0 tok/s to 65.26. That's 29% slower.
Same M1 Max, same model, and right next to it an open-source engine also running speculative decoding. Its acceptance rate is 0.38, lower than my 0.51. It runs at 133.4, which is 1.45x faster than my engine with speculation turned off.
It guesses worse and runs faster. Acceptance rate clearly isn't what decides this.
The part that stings: one division predicted this back in July. My first pass said "small win, +6.6%". That evening I found I'd used the wrong denominator, and the corrected answer was "you lose no matter how well you guess". In October I finally ran the full end-to-end loop, and the division was right.
If you just want the division, it's section 2. If you want to see how I turned a loss into a win on paper, that's section 3.
1. What speculative decoding is actually betting on
Short version: a cheap small model guesses the next token, and the big model checks the guess. If the guess is right, you get two tokens out of one round. If it's wrong, you get one, and correctness never suffers.
The model I use ships with a small multi-token prediction head (MTP) whose whole job is guessing the next token. One guess costs about 2.6 ms.
The bet hides in the "check" step. The big model has to compute two positions: one to confirm the current token and one to test the guess.
Decode, the phase where tokens come out one at a time, is almost entirely spent hauling weights out of memory while the compute units sit mostly idle. So the ideal verify step reads the weights once and computes both positions in the same pass, costing roughly what a single decode step costs. Then every correct guess is a free token.
If you can't compute both positions together, you're stuck running the forward pass twice. That's a very different deal, as the math shows.
2. The division
Tokens per round, divided by time per round, compared against plain decode:
speedup = (1 + p) × T_decode ÷ (T_verify + T_draft)
p is the acceptance rate, T_decode is the time per token without speculation, T_verify is one verify round, and T_draft is one guess from the small head. It's a single division with no fudge factors.
Plug in my numbers. Without speculation I get 92.0 tok/s, so T_decode = 1000 ÷ 92.0 = 10.87 ms. I eventually got verify down to 21.0 ms wall clock, draft is about 2.6 ms, p = 0.51:
(1 + 0.51) ÷ (21.0 + 2.6) ms = 1.51 ÷ 23.6 ≈ 0.064 token/ms ≈ 64 tok/s
measured: 65.26
That's within 2%.
2.1 A corollary that settles it on the spot
Suppose verify really is two forward passes, so T_verify = 2 × T_decode. Substitute:
speedup = (1 + p) × T_decode ÷ (2 × T_decode + T_draft)
< (1 + p) ÷ 2
≤ 1
p can't exceed 1. So single-token speculative decoding where verify costs two forward passes loses no matter how accurate the guesses are. My July ledger put p = 1 at 0.971x, which is exactly this ceiling minus the 0.643 ms draft.
My verify step takes 20.4 ms of GPU time, about 1.92x decode, so it's still effectively two passes. When I break it down by op type, every category comes in at 2.0x. None of them get amortized across the two positions. The dense weights really are being read twice.
2.2 The break-even line
Set speedup = 1 and solve:
T_verify ≤ (1 + p) × T_decode − T_draft
At p = 0.51: 1.51 × 10.87 − 2.6 ≈ 13.8 ms, about 1.27x decode.
So verifying two positions can cost at most 27% more than a single decode step just to break even. Mine costs 92% more. Parameter tweaks won't close that. It takes a different kernel structure.
3. How I turned a loss into a win in July
The first time I ran this math in July, I used 8.128 ms as the per-token time. That number came from summing per-block timings:
linear attention 4.356 + full attention 1.424 + MoE 1.266 + output head 0.902 + misc 0.18 ≈ 8.128 ms
Meanwhile the measured end-to-end speed of the whole graph was 91.4 tok/s, which is 10.941 ms per token.
I used 8.128 for the verify cost and 91.4 as the baseline. Same acceptance rate, 0.6457:
wrong: 1.6457 ÷ (2 × 8.128 + 0.643) = 1.6457 ÷ 16.899 ms ≈ 97.4 tok/s → 1.066x over 91.4
right: 1.6457 ÷ (2 × 10.941 + 0.643) = 1.6457 ÷ 22.525 ms ≈ 73.1 tok/s → 0.799x
The per-block sum missed 2.813 ms, which is 34.6% on top of the sum itself. All of it lives between blocks: dispatch overhead, encoder switches, barriers, CPU-side encoding.
A 34.6% error is exactly enough to turn a 20% loss into a 6.6% gain. Numerator and denominator were measured with two different rulers, so the conclusion came out backwards.
The rescue plan was a kernel that reads the weights once and computes two positions. I tried it on the output head first. The premise held: on a matrix that large, running it twice took 1.990x as long as once, so the weights really were read twice and in theory there was a full 2x of bandwidth to save.
Three implementations later, the best ran at 0.441x to 0.656x the speed of just running it twice. I ran controls for three hypotheses (register spilling, two input streams fighting each other, hitting the compute ceiling), and my own experiments ruled out all three. I still don't know the root cause. That's why verify was still 1.92x in October.
4. How to pick apart a speculative decoding benchmark that only reports acceptance rate
4.1 First, ask how many decode steps one verify round costs
Back to the engine that hits 133.4. Each round, its draft model guesses 7 tokens and the main model verifies 8 positions at once. A full round takes 30 to 33 ms, draft and verify included. Acceptance rate 0.38, averaging 3.1 to 4.9 tokens per round.
Same division: roughly 4 tokens ÷ 31 ms ≈ 129 tok/s, in the same range as the measured 133.4.
On my side, verifying 8 positions alone takes 51 ms, before counting the draft. One of its full rounds costs about 2.9 of my decode steps. Spread over 8 positions, that's 0.37 steps per position.
So it wins because verify is cheap, and guess accuracy has little to do with it. p is a term in the numerator; T_verify is the denominator. If the denominator is several times larger, no numerator can make up for it.
4.2 Then ask what sampling the acceptance rate was measured under
The official model card lists an acceptance rate of 88.58%. I chased that number for weeks before realizing it was measured with temperature sampling, and I run greedy decoding.
The two aren't comparable. Under temperature sampling, both sides have to randomly land on the same token. Under greedy, both just take the argmax. My own greedy measurement is 64.6%. Using 88.58% as the target was a measurement-definition error from the start.
4.3 Finally, ask how long the run was and what content it used
Same code: at 48 steps I measured 53.2% acceptance; at 128 steps it dropped to 34.7%. The confidence interval on the short run was ±14.3%, so the 53.2% was plain luck.
Content moves the numbers a lot too. The second drafting method I tried (DFlash2, which guesses a long run of tokens per round) accepted an average of 4.59 tokens per round on code prompts and only 0.5 to 1.12 on Chinese. An acceptance length measured only on code tells you almost nothing about chat in another language. That line ended up at about 19.5 ms per token, roughly 51 tok/s, 44% slower than no speculation. I killed it.
When you look at any speculative decoding benchmark, ask in this order: how many decode steps does one verify round cost; what sampling, length, and content was the acceptance rate measured on; is the headline number end-to-end or projected. If you can't answer one of those three, don't trust the number yet.
5. Where my numbers shouldn't be trusted
- One machine, one model. M1 Max, a 35B MoE model at 4-bit. 30 of its 40 layers use linear attention, where each position continues from the previous position's state. That part (53.2% of those layers' time) reads state rather than weights, so two positions can't share it. I haven't measured how much cheaper verify gets on a pure dense model.
- Small test set. 18 prompts, 6 each of English, Chinese, and code, 33-token prompts, 128 tokens generated each. The GPU cooled to 65°C before each run, and I took the median of the last 12. I didn't test long conversations.
- The 133.4 engine is a speed comparison only. It uses its own quantization, and the prefix that matches my output token for token is only 3 to 128 tokens long, so I can't say which is "more correct".
- 0.51 and 0.646 don't agree. End-to-end acceptance is 0.51; my July CPU reference implementation measured 0.646. The first version was around 0.11 because the draft head's KV cache never saw the prompt, and fixing that brought it to 0.51. I haven't tracked down the rest of the gap.
- The 64 in section 2 is an order-of-magnitude match. Verify went from 24.2 ms down to 21.0 ms while throughput climbed from 45.2 to 65.26, but I didn't pair them version by version. Only the final version lines up. Across the full ledger, predictions landed within ±5% of end-to-end measurements.
- "True batched verify could reach 1.38x" is a projection. It hasn't been measured.
- I don't know why the two-positions-at-once kernel is slow. Digging further needs GPU hardware counters, and my current measurement tooling can't get at them.
6. Before you wire up speculative decoding
- Measure whole-graph T_decode as 1000 ÷ measured tok/s. Don't use a sum of per-block timings.
- Measure one forward pass that computes 2 positions at once and divide by T_decode. If it's ≥ 2, stop: single-token speculation can't win.
- Measure acceptance rate p with your own sampling, your own content, and at least 128 steps.
- Plug into T_verify ≤ (1 + p) × T_decode − T_draft to get your break-even line.
- If break-even sits well below your current verify cost, write the batched verify kernel first and swap draft models later.
- If someone's benchmark only reports acceptance rate, act like you didn't see it.