For two months I had one number taped to the wall: 843.74.
That was the prefill speed (tok/s) of BaseRT, a closed-source inference engine for Macs, running Qwen3.6-35B-A3B on my M1 Max. Every time I finished an optimization round on my own engine, I compared it against that number.
Last night I re-measured. Same machine, same model. The fastest engine is now llama.cpp at 927.9. The engine that had been my yardstick for two months scores 280.7 in its latest release. That's a third of what it used to do.
Decode? 96.2 → 92.8. Only 3.6% slower. If you only ever look at generation speed, you would upgrade and never notice.
If you just want the arithmetic, skip to section 3. If you want to know why my "turn off a component" experiment turned nothing off, see section 4.
1. What happens when a yardstick sits on the shelf for two months
First, the terms. Prefill is the phase where the model reads your whole prompt before answering; it decides how long you stare at a blank screen after hitting Enter. Decode is the phase after that, where tokens come out one at a time. Below, pp512 is the throughput for reading a 512-token prompt, and tg128 is the throughput for generating 128 tokens in a row.
Everything ran the same night, sequentially, one engine at a time. Before measuring I paused macOS's photo analysis daemons (they eat CPU hard, and two days earlier they had already spoiled half of my measurements):
| Engine | Weights | pp512 | tg128 |
|---|---|---|---|
| llama.cpp b11382 | Q4_K_M (mixed 4/6-bit) | 927.9 | 61.3 ± 10.5 |
| BaseRT 0.2.0 (old) | official 4-bit bundle | 850.1 | 96.2 |
| my own engine | my own 4-bit format | 830.6 | 90.5 |
| Splash, community M1 port | official bundle | 783.7 (lower bound) | 81.2 / 125.0 |
| mlx-lm 0.32 | community 4-bit | 674.9 | 72.2 |
| BaseRT 0.2.6 (new) | official 4-bit bundle | 280.7 | 92.8 |
Two things are worth saying up front.
First, old BaseRT itself held steady. 843.74 in August, 850.1 now, less than 1% apart. The yardstick never broke. My assumption did: I took for granted that it was still the fastest. Back in July, llama.cpp was clearly behind on prefill on this machine. Over two months it caught up, and I never looked back once.
Second, this table isn't perfectly apples to apples. llama.cpp's Q4_K_M keeps some layers at 6 bits, mlx-lm uses a different 4-bit scheme, and my own engine uses my own quantization format and was timed in my own eval harness, not llama-bench. Trust the ordering; don't read much into the decimals.
2. Bisecting: which release fell off the cliff
The nice thing about BaseRT is that every version can be downloaded on its own, and old and new versions read the same model bundle. So I could run pp512 on the same weights file across releases, which takes the weights out of the picture entirely:
- 0.2.0: 850
- 0.2.1: 807
- 0.2.2: 807
- 0.2.3: 804
- 0.2.4: 279
- 0.2.5: 275
- 0.2.6: 279
One step from 804 to 279. The 850 → 807 drop earlier is real too, about 5%, but it's nothing next to this.
The 0.2.6 release notes say 0.2.4 and 0.2.5 produced garbled output for this model on M1, and 0.2.6 fixed that. So the likely story: 0.2.4 switched to a new compute path on M1, the first cut even got the math wrong, and later the output was fixed while the speed wasn't. That's an inference. I have no proof of which kernel it is.
So the next question: without reading any source, can you tell which part got slower?
3. One division: fixed overhead, or every token got pricier?
A slowdown comes in two shapes.
One is fixed overhead going up: an extra initialization, an extra sync, an extra buffer allocation. However long the prompt is, you pay the same few dozen milliseconds more.
The other is every token getting more expensive: some kernel that scales with token count got slower. The longer the prompt, the more you pay.
A single pp512 number can't tell them apart, because throughput = tokens ÷ time, and both shapes make it smaller. To separate them you need a few prompt lengths and a two-point fit:
time ≈ fixed overhead + per-token cost × tokens
per-token cost = (long-prompt time − short-prompt time) ÷ (long tokens − short tokens)
fixed overhead = time − per-token cost × tokens
throughput ceiling for long prompts ≈ 1000 ÷ per-token cost (ms)
It's a straight line from middle school. The slope is the price of one token; the intercept is the fixed overhead.
I measured four lengths on both 0.2.3 and 0.2.4 (milliseconds):
| prompt length | 0.2.3 | 0.2.4 | times slower |
|---|---|---|---|
| 32 | 129.8 | 223.9 | 1.72 |
| 128 | 257.5 | 610.2 | 2.37 |
| 512 | 636.4 | 1836.6 | 2.89 |
| 2048 | 2050.7 | 6869.0 | 3.35 |
Look at the right-hand column first. You can almost read the answer without doing any math: the longer the prompt, the bigger the slowdown. A fixed-overhead regression goes the other way. Those extra milliseconds get spread thinner over longer prompts, so the ratio drifts toward 1.
Now plug in the 512 and 2048 points:
- 0.2.3: per token (2050.7 − 636.4) ÷ 1536 = 0.92 ms, fixed overhead 636.4 − 0.92 × 512 ≈ 165 ms
- 0.2.4: per token (6869.0 − 1836.6) ÷ 1536 = 3.28 ms, fixed overhead 1836.6 − 3.28 × 512 ≈ 159 ms
Fixed overhead barely moved. Each token got 2.36 ms more expensive. All of the new version's slowdown is in the part that works per token.
The ceilings line up too. For 0.2.3, 1000 ÷ 0.92 ≈ 1086, and measured pp2048 has already climbed to 999. For 0.2.4, 1000 ÷ 3.28 ≈ 305, and measured pp2048 is 298, basically pinned against the ceiling. On this machine, the new version won't go past 305 tok/s no matter how long the prompt gets.
To be fair to the method: the line isn't accurate for short prompts. Fit through 512 and 2048, it predicts 194 ms for 0.2.3 at 32 tokens; the actual number is 130. At short lengths the GPU isn't saturated, the per-token cost is higher than on long prompts, and the curve bends. So fit over the range you actually care about. I care about prompts from a few hundred to a few thousand tokens, so I use 512 and 2048. Going segment by segment, 0.2.4's extra per-token cost lands between 2.2 and 2.7 ms. Same ballpark, same conclusion.
What I like about this division is that it works on any engine. No source, no profiler. You only need to control prompt length and read the elapsed time.
4. I turned off a component and not a single millisecond went away
Once you know it's "every token got pricier", the obvious next step is finding which component got pricier.
BaseRT has a set of BASERT_SKIP_* environment variables whose names suggest exactly that: skip routed MoE experts, skip attention, skip the shared expert, skip the LM head. Turn one off at a time, see how much time disappears, and whatever drops the most is your culprit.
Here's what I got (pp512, ms):
- 0.2.3, nothing skipped: 636.4
- 0.2.3, routed MoE experts skipped: 634.4
- 0.2.3, attention skipped: 636.4
A 35B MoE model, all routed experts skipped, and it saved 2 ms. That can't be right. The expert layers are most of this model; actually skipping them should cut the runtime by a big chunk.
So the correct reading is: these switches don't do anything in the bench binary. Maybe they're only wired up in server mode, maybe they were removed in some release and the docs never caught up. I nearly read "skipping MoE changes nothing" as "MoE isn't the bottleneck" and went digging in the wrong direction.
The rule is simple: if the thing you turned off should account for a big slice of time and the runtime barely changes, suspect the switch before you suspect the component. Before any ablation, use a component you know is heavy as a positive control and confirm the switch actually moves the runtime.
After that I tried 13 environment variables that steer kernel routing, one at a time on 0.2.4 and 0.2.6. The best one reached 283.6, under 2% above 279. The dense matmul kernels (types, shapes, call counts) were identical in the traces of both versions. What's left points at the new MoE or linear-attention path, and pinning down the exact kernel needs a Metal GPU capture. That's upstream's problem. It doesn't affect how I use it (I'm staying on 0.2.0), so I stopped there.
5. Three questions to ask about someone else's numbers
The same night I tripped over a few classic ways benchmark numbers mislead you.
One: prefill derived from time-to-first-token is only a lower bound. The Splash port has no bench subcommand, so I started its server, sent a 512-token prompt, measured time to first token (TTFT), and divided tokens by that. That interval also includes the HTTP round trip, tokenization, and the first decode step. So 783.7 only means "at least this fast".
Run section 3's division on it: TTFT is 649.4 ms at 509 tokens and 1692.3 ms at 1488 tokens, so 1.07 ms per token, about 107 ms fixed, and a long-prompt ceiling around 939. That's stitched together from one point in each of two runs, so treat it as a ballpark.
Do the same for BaseRT 0.2.3: pp512 is 804, but the ceiling is 1086. The two engines are 20 apart at pp512 and 150 apart at the ceiling. A throughput number at one prompt length is a single point on a curve, and rankings shift with length. If someone only reports pp512, you don't get their ceiling.
Two: decode speed with speculative decoding on depends on what text you feed it. Speculative decoding lets a small model guess the next few tokens and the big model verify them in one pass; every correct guess is free speed. This Splash build can't turn it off, and there's no switch in the source. Same engine, same model, 128 generated tokens each time:
- continuing random text: 81.2 tok/s
- writing normal prose: 125.0 tok/s
A 1.54× gap. Prose is easy to guess, random text isn't. That's all there is to it. If someone tells you their engine decodes at 125, ask what content they measured on, then ask whether speculative decoding was on. Comparing that directly against a plain 96.2 compares two different things.
Three: look at the error bars. llama.cpp's tg128 is 61.3 ± 10.5, a standard deviation of 17% of the mean, and I didn't track down why. With error bars that wide, whether 61 ranks above or below 72 doesn't really hold up.
6. What's wrong with this night's data
- One machine, one model. I've only seen the 0.2.4 regression on an M1 Max, and the release notes say the issue is M1-specific. M2 and later chips very likely don't have it, so don't use my numbers to talk anyone out of upgrading.
- 0.2.0 wasn't part of the bisect run. The 850.1 came from a separate run the same night, not back to back with the 0.2.1–0.2.6 sweep.
- Each version got one bisect run. The bench repeats internally and reported standard deviations all under 4, so I didn't run more. But run-level drift like temperature or background processes won't show up in a single run.
- My own engine was timed in a different harness. The 830.6 comes from my own eval setup (alternating with the previous version so neither gets a first-run advantage), not llama-bench or basert-bench. There may be a few percent of methodology difference between it and the rest of the table.
- The root cause of the regression isn't confirmed. "Switched to a new MoE or linear-attention path" is an inference, based on unchanged dense GEMM plus the release notes. There's no kernel-level evidence.
- The short-prompt region isn't a straight line. As noted in section 3, the division overestimates time below 128 tokens. The conclusion only holds from a few hundred tokens up.
7. Next time you upgrade an inference engine, test in this order
- Run before and after the upgrade, and report prefill and decode separately. If you only watch decode, this 3.6% change looks like noise and you miss that prefill fell to a third.
- Measure prefill at two or more lengths, picked from the range your real workload uses, e.g. 512 and 2048. Watch how the "times slower" ratio moves with length: rising means per-token cost grew, drifting toward 1 means fixed overhead grew.
- Compute per-token cost as (long time − short time) ÷ (long tokens − short tokens). 1000 divided by that is the engine's prefill ceiling on your machine.
- When comparing versions, pin the exact same weights file. Otherwise you can't tell whether the engine changed or the model bundle did.
- Before any ablation, use a component you know is heavy as a positive control, and confirm that turning it off actually changes the runtime.
- Whatever number you use as a yardstick, re-run every candidate against it every month or two. My yardstick never moved. Someone else quietly passed it while I wasn't looking.