At 23:15 that night I finished a rerun of the inference engine I've been writing and sent off a one-liner: "Ahead at 1K and 2K, basically a tie at 4K."
On a 4096-token prompt my engine did 979 tok/s. The competitor did 984. A 0.5% gap. Calling that a tie seemed fair.
An hour later I took it back. The competitor's 984 came from several runs back-to-back in the same process. Let it rest, fire a single request, and it does 1017–1053. My 979 was the opposite: the first shot from a freshly started process. Run mine back-to-back and it drops to around 850.
Cold versus cold, I was 4–7% behind. Hot versus hot, about 15% behind. I had compared my best shot against the competitor's throttled ones and manufactured a tie out of nothing.
The annoying part: the whole time, the OS kept reporting thermals as normal.
If you just want the conversion formula, jump to section 3. If you want to know why "just read less memory" went nowhere, read section 5.
1. First, make sure it really slows down
Quick context. This is a single-model inference engine I'm writing for Qwen3.6-35B-A3B on an M1 Max (64 GB, plugged in, Low Power Mode off). Everything below is prefill, the phase where the model reads the whole prompt before emitting its first token. Speed is a single division:
tok/s = 4096 ÷ prefill time (s)
4180 ms → 980 tok/s
4850 ms → 845 tok/s
Same process, same prompt, four shots in a row:
980 → 939 → 880 → 881
980 → 848 → 840 → 843 (another pass)
My first guess was leftover state, something like a KV cache from the previous request that never got cleared. Three controls ruled that out:
- Run the prompts in reverse order: whichever one goes first still lands around 980
- Raise the context window to 16384: still slows down
- Start a separate process for each prompt: all of them land at 978–982
So the slowdown depends only on which shot in the sequence it is. The prompt content and anything accumulated in the process don't matter.
Next I tried idle gaps between shots: 3 s, 15 s, 45 s.
3 s / 15 s / 45 s gap: all steady at 931–936
no gap: drops to 840–880
A 3-second rest and a 45-second rest look the same. That pattern looks like a budget. Once it's spent, you either rest to get some of it back or you stay throttled.
2. Thermals say normal, the clock is dropping
I sampled with mactop every 250 ms. On the first 4096 shot the GPU ran flat out at 1296 MHz, with power reading around 41 W. Once shots ran back-to-back, the clock fell to 1000–1130 MHz and power to 22–32 W.
Over the same window, the system's thermal_state stayed "nominal" and pmset never logged a thermal or performance warning.
Put plainly, the OS says everything is fine while the GPU is being throttled. If you only check the OS thermal state to decide whether you're being throttled, you will miss this one completely.
Then I brought the competitor in under the same conditions. Both warmed up first, then ran 4096 six times in a row in the same process. I took the median of the last 70% of samples:
| Competitor (BaseRT 0.2.0) | My engine | |
|---|---|---|
| Back-to-back 4096 | 993 / 1012 | 1st shot 980, then 841–866 |
| GPU clock | 1218–1246 MHz | 1109–1123 MHz |
| GPU power | 37–38 W | ~32 W |
| Memory bandwidth | ~29 GB/s | 49–51 GB/s |
| Rest-of-system power | 18.8 W | 24.4 W |
| Whole machine | ~60 W | ~60 W |
Both engines park the whole machine at about 60 W. Plugged in, this laptop behaves as if it has a 60 W ceiling.
The ceiling doesn't move, so the budget can only be split like this:
GPU share ≈ 60 W − memory/fabric side − small remainder
Competitor: 37.5 + 18.8 = 56.3
Me: 32 + 24.4 = 56.4
The remainder is about 3.6 W on both sides. My engine spends 24.4 − 18.8 = 5.6 W more on the memory side, and the GPU gets 37.5 − 32 = 5.5 W less. Watt for watt. Less power for the GPU means a lower clock: 1116 versus 1232, 9% lower.
My engine moves 70% more data per second through memory than the competitor does. On a cold start, with the GPU at full clock, that extra traffic is completely hidden. Once you hit the ceiling, it turns straight into lost clock speed.
3. A conversion you can compute yourself: sustained ≈ cold × clock ratio
If a piece of code needs a fixed number of GPU cycles, its speed scales with clock. I later did a cycle-level attribution, and the sustained slowdown came entirely from the lower clock. The cycle count didn't change. So you can write:
sustained speed ≈ cold speed × (sustained clock ÷ cold clock)
Plugging in the numbers:
My old build: 980 × 1116 ÷ 1296 = 844 measured 841–866
My new build: 1048 × 1162 ÷ 1296 = 940 measured 946
Competitor: 1040 × 1232 ÷ 1296 = 989 measured 993–1012
The first two rows are within 1%. The competitor's row is off by 0.4–2.3%, because I never measured its clock in the cold state. I assumed it also runs at the full 1296, and I took the median of its three single shots, 1040, as the cold speed.
The handy part is that it works in reverse. If all you have is a cold-start benchmark, glance at the clock under sustained load and you can estimate how far it will drop without running a long test. And if the estimate is way off from what you measure, something besides the clock is slowing you down. The cycle count changed, which is a different problem and needs a different investigation.
4. When you see a benchmark number, ask which shot it was
Back to that "tie." Neither measurement was wrong. The mistake was putting two different kinds of number side by side.
The competitor's bundled benchmark tool runs several repetitions back-to-back and reports the average, so its numbers lean hot by construction. My rerun that night started a fresh process for each prompt, so my numbers were cold by construction. Put the two together and both biases point in my favor: my side is the best shot, theirs is the throttled ones.
That same morning I had also built a comparison table of 13 engines. In it, my own 4096 number was only 742 and the competitor's was 984, and both came from back-to-back runs. I used that table to pick the next round's target, and never labeled how it was measured.
So for any inference benchmark on a laptop, an all-in-one, or a mini PC, ask three questions first:
- Is it the first shot from a fresh process, or repeated runs in the same process? If repeated, from which run onward?
- How long did it rest between runs?
- Is the number it's being compared against measured the same way?
On this machine, a "first shot from a cold start" number can run about 15% above back-to-back. Compare that against someone else's back-to-back number and "15% behind" turns into "a tie." That's exactly what I did.
Looking back, the ceiling had already shown itself the day before. I'd noticed that a long prompt queued behind other tests in the same process ran 4–15% slower, with the amount depending on the order. Comparing a build against itself could produce a verdict of −2.34%. My fix at the time was to give each long test its own process, which cut the noise to ±0.15%. That was most likely this same ceiling. I just didn't think of it that way then, and I never went back to verify it separately.
There's an older counterexample too. In July, on the same machine, I investigated why decode (the phase where the model emits tokens one at a time) got slower the longer it ran, and I suspected a power cap then as well. That time the data pointed the other way: my engine drew 20 W, less than the competitor's 25 W, and was also 20% slower. If it were a power ceiling, the higher-power engine should have been the one getting squeezed. So I ruled the power cap out. Same machine, different workload, opposite answer.
5. Saving bytes didn't help. Removing duplicate reads did
After the diagnosis my first instinct was obvious: too much memory traffic, so read less.
The first attempt was to split the long prompt into chunks, so the intermediate results would be small enough to stay in the chip's system level cache (SLC, a large cache shared by every unit on the chip) and never go out to DRAM:
| Chunk length | Data moved per prefill | Sustained speed |
|---|---|---|
| 4096 (no split) | 241 GB | baseline |
| 2048 | 191 GB | −1.1% |
| 1024 | 204 GB | −4.8% |
| 512 | 242 GB | −11.3% |
At 2048 per chunk, data moved dropped 21% and speed went down 1.1%. The finer the split, the slower it got, because every chunk carries a fixed overhead.
What actually worked was a different kind of change: find places where the same block of data gets read from memory separately by several compute units, and make it a single read that they all share.
- In the matrix multiplies, I changed the tile traversal order so the activation tile no longer gets re-read. Data moved per prefill went 238 → 175 GB, memory-side power 24.2 → 23.1 W, clock 1100 → 1140 MHz, speed +3.1%.
- In attention, each K/V block used to be read once by each of 8 query heads. Now each block gets staged once into on-chip shared memory and all 8 heads read it from there. The sustained cost of that step went from 753 ms / 34 GB to 395 ms / ~9 GB, speed +7.5%.
Both changes reduce data movement. So why did chunking fail while removing duplicate reads worked? My reading: the memory-side power number counts all traffic that leaves the GPU's own L2 cache, SLC hits included. Chunking only moved traffic from DRAM into the SLC, so that power line barely changed. Removing duplicate reads made that traffic disappear entirely. The explanation fits the data, but I have no counter that isolates SLC traffic, so it's an inference.
There's also a saturation point. Once data moved per prefill dropped below about 165 GB, further savings stopped raising the clock. Past that point the only lever left is cutting GPU cycles directly.
Putting it together:
Sustained 4096: 848 → 946 (+11.5%)
Cold 4096: 979 → 1048
Memory bandwidth: ~50 → 32–33 GB/s
GPU clock: ~1120 → 1158–1166 MHz
Memory side: 24.4 → 22.0–22.2 W
GPU: 31.5 → 34.4 W
Whole machine: unchanged, ~59.7 W
The ceiling didn't move an inch. Power just shifted from the memory side to the GPU. The competitor does 993–1012 sustained; I went from about 0.85× to about 0.94×, still 6% short. Before the changes my engine moved about 243 GB per prefill, measured, against roughly 119 GB for the competitor's entire prefill. There's more of that bill left to pay.
6. Which numbers here aren't solid
- One machine, one model, plugged in. I inferred the 60 W ceiling from both engines' whole-machine power settling around 60 W. I found no official documentation for it and didn't check on battery or on other M-series machines.
- All power and clock figures are mactop software readings at 250 ms intervals, with no external power meter.
- The competitor's sustained numbers come from its own bundled tool; mine come from my own test script. The timing methods aren't identical. The competitor's cold clock in section 3 is assumed.
- "Cold" isn't a single state either. The first shot from a fresh process reaches 980, but a 45-second rest only gets back to 933. My guess is that the GPU idles more fully during the few seconds a new process spends loading weights. Unverified.
- The back-to-back logs contain two outliers, one shot at 14.8 s and one at 8.9 s, with decode also several times slower at the same moments. I didn't track down the cause. They don't move the medians, but they show something else on this machine occasionally jumps in.
- "Median of the last 70% of samples" and "6 shots in one process, keep the last 3" are conventions I picked. Choose differently and the numbers shift a little.
- In section 5, "chunking just moved traffic into the SLC" is an inference, as noted there.
7. Measuring inference speed on a laptop: do it in this order
- Decide the measurement first. Do you care about the first shot, or about a resident process handling requests back-to-back? If it's the latter, measure back-to-back and throw away the first shot.
- Fire 5 or more shots in the same process and see whether run 2 onward drops. If it does, there's a ceiling, and you should stop quoting the first shot.
- Sample GPU clock and power while you measure. Don't trust the OS thermal state; this time it read nominal the entire run.
- Reconcile against "sustained ≈ cold × clock ratio." If it matches, only the clock slowed down. If it doesn't, go look at cycle counts.
- Before comparing with anyone else, confirm both numbers were measured the same way: which shot, how much rest, whose tool.
- Run long tests in their own process, or do a fixed warm-up before timing, so test order doesn't leak into the results.
- To go faster under the ceiling, first look for data that gets read several times. For byte-saving tricks like chunking or lower precision, check whether the traffic actually went away or just moved somewhere else.
- A hypothesis you ruled out last time deserves a fresh look under a different workload.