Back to blog

2026.10.05

12% behind llama.cpp, and four agents spent 8 hours to gain 0.66%: which operator does the gap between two engines actually live in?

My own inference engine does 830.6 tok/s prefill on an M1 Max; llama.cpp does 927.9. Four AI agents spent 8 hours tuning the MoE path and gained 0.66%. Then I stopped and spent one evening reconciling the two engines operator by operator. MoE turned out to be the one category where I was already faster. The whole gap was in dense projections and transposes. Eight more hours aimed at the right target gained 9.9%. This post covers how to reconcile two engines per operator: a replay-and-diff measurement, a coverage ratio, a bytes-read check for bandwidth-bound kernels, and why a standalone microbenchmark undercounts MoE by 60%.

推理引擎性能方法论Apple Silicon

Four AI agents, 8 hours, dozens of automated evaluation runs. Prefill throughput went up 0.66%.

Same M1 Max, same model. My own engine runs 830.6 tok/s. llama.cpp runs 927.9, about 12% faster. During those 8 hours, three of the four agents were working on the MoE experts, because my call from the previous round was "MoE is where the biggest gap is."

Then I stopped and spent one evening reconciling the two engines, operator category by operator category. MoE was the one place where I was faster: 245.9 ms versus 270.1.

Three agents had spent 8 hours optimizing something that was already ahead.

I switched targets based on the ledger, ran another 8 hours, and gained 9.9%, which puts me at 98.8% of llama.cpp.

If you just want the method, read sections 2 and 3. If you want to know how a standalone benchmark undercounts MoE by 60%, read section 4.

1. How I guessed wrong

Quick context. I'm writing an inference engine that serves exactly one model: Qwen3.6-35B-A3B. It's a mixture-of-experts model, meaning each layer contains 256 small "expert" networks and each token only wakes up 8 of them. Prefill is the phase where the model reads your whole prompt before generating anything. Every number below is for a 512-token prompt (pp512).

In the two rounds before this, I made two directional calls. Both were wrong:

  • Round 9, I said "the dense projections are saturated, leave them alone." The next round got +3.9% from exactly those projections.
  • Round 10, I said "MoE gate_up is furthest from peak, attack it." In round 11, three agents piled onto it. Every variant was flat or a regression, the worst at -40%.

Both calls had the same flaw: they were based on time shares inside my own engine, never on a same-method comparison against the competitor. I knew MoE took a lot of my time, so I figured it was worth optimizing. Taking a lot of time doesn't mean it's slower than the other guy's. The other guy might spend even more there.

Put plainly, I kept reading my own health report and never looked at theirs.

2. Step one: break the competitor down, and use coverage to check the breakdown

llama.cpp ships two tools that fit this job: one exports the operators and shapes that actually appear in a given compute graph, and one times each operator in isolation.

The method is two multiplications:

category time = per-call time × number of times that op appears in the graph
coverage = Σ(category times) ÷ measured end-to-end time

I didn't derive the counts from the model architecture. I counted nodes in the real graph: 3845 nodes, minus 1490 view nodes that just reinterpret memory without doing work, leaving 2355. Then I cross-checked against the architecture: 40 layers, 30 of them linear attention and 10 full attention, with separate gate and up MoE multiplies per layer, so 80 calls. It matched.

Summing it up:

all categories: 528.0 ms
llama-bench pp512 measured 927.87 tok/s → 512 ÷ 927.87 = 551.8 ms
coverage = 528.0 ÷ 551.8 = 95.7%

95.7% looks great. Dense projections at 209.6 ms are 38% of the total, MoE gate_up 113 ms, the linear attention block 76.5 ms.

I'm going to tear that pretty number apart in a minute. Here's why: coverage near 100% only tells you the sum lines up. It doesn't tell you each category is right. Two errors pointing in opposite directions cancel out and the total still looks great. Section 4 is exactly that.

3. Step two: replay each category inside the full graph, on both engines

Standalone operator timing has a basic problem. It pulls the op out and runs it alone, with no neighbors competing for resources and no concurrent overlap with other ops. Comparing that against my engine's in-graph timing is an apples-to-oranges comparison.

So for the real reconciliation I switched methods, and used the same one on both sides:

in-graph time of a category ≈ T(that category dispatched one extra time in place) − T(baseline)

In other words, during a full inference run, re-execute one category of ops in place, with a barrier on each side, and see how much the total grows. The growth is what that category costs in its real environment. The replayed writes go to a scratch buffer, so the actual output stays bit-identical.

The method itself needs checking. I ran three checks:

  • Replay everything at once: increment 605.4 ms, baseline 605.3 ms. Matches.
  • Replay each category separately and sum: 606.8 ms. Also matches.
  • Replay one category 1, 2, and 4 times: 162 ms per replay every time. Linear, so the replay itself isn't adding hidden overhead.

The ledger, sorted by gap:

Category My engine ms llama.cpp ms Gap
Dense projections 252.2 188.3 +63.9
Transposes around projections 25.3 0 +25.3
Attention 19.1 5.8 +13.3 (different scope, see section 6)
Linear-attention recurrence 53.5 55.5 −2.0
MoE experts 245.9 270.1 −24.2
Whole graph 605.3 553.5 +51.8

The MoE row is negative. I'm 24 ms faster there.

The bulk of the gap is in the dense projections (the ordinary matmuls around attention) plus a step llama.cpp simply doesn't have: I transpose the data before and after the projections, while llama.cpp consumes it in its native layout. That step alone adds 25.3 ms out of nowhere.

There's also a check here you can run yourself, to tell whether a matmul is bound by data movement:

if bandwidth-bound: time ∝ bytes of weights read
llama.cpp's dense weights are 8.5 bits; mine are 4.5 bits
byte ratio = 4.5 ÷ 8.5 = 0.53
if bandwidth-bound, I should take roughly 188.3 × 0.53 ≈ 100 ms
measured: 252.2 ms

Nearly half the bytes, 34% more time. Stop looking at bandwidth. The problem is compute or dequantization (unpacking 4-bit compressed weights into numbers the hardware can multiply). The estimate is rough, since 188.3 also includes a small slice of F32 router weights, but the gap is too wide for that to change the conclusion.

Along the way I caught something embarrassing. In the previous round, my verifier agent had written a per-kernel timing probe that missed the single largest kernel, 155.9 ms. That kernel binds its buffers first and sets its pipeline second, and setting the pipeline wiped the probe's record of the bindings, so it vanished from the log. A whole round of per-kernel timing was missing its biggest piece, and I never noticed.

4. Why the standalone benchmark undercounted MoE by 60%

Back to that 95.7% from section 2.

When llama.cpp's operator benchmark generates test inputs for the MoE op, it assigns every token a shuffle of experts 0 through 7. So 512 tokens only ever touch 8 experts, and each call reads just 4.5 MiB of weights. In real inference, each token picks 8 out of 256, and 512 tokens end up touching nearly all of them, reading 144 MiB per call.

Change only the expert distribution and re-time:

gate/up: 8 experts only 1447 µs → random 8 of 256: 2315 µs (+60%)
down:    8 experts only 1577 µs → random 8 of 256: 2391 µs (+52%)
MoE total: 174.9 ms → 280.9 ms
coverage: 95.7% → 114.9%

Coverage blows past 100%, because the opposite error is also present: in the full graph, small ops get fused or overlap with other work, so timing them alone and summing overcounts. One error low, one error high, and they happened to cancel at 95.7%.

So whenever someone benchmarks an MoE kernel in isolation, ask one question first: what does the expert distribution in the test input look like? If it only touches a handful of experts, the weights all sit in cache and you're measuring a speed that never happens in real inference.

The timing tools have their own traps. While I was at it, I tried several ways to time a single GPU kernel on the M1. The short version: the public API on this chip can only timestamp at the boundary of a whole batch of GPU work. To time one kernel, you have to put it in its own batch, and then it can no longer run concurrently with anything:

two kernels in the same concurrent batch: 21.00 ms (almost fully overlapped)
two kernels serial:                        39.59 ms
split into two separate batches:           37.82 ms

Split them up to measure, and the overlap disappears. You end up measuring a world that's slower than the real one. The OS profiler's per-shader timeline is sampled, not timestamped: run one kernel 6 times on its own and it records only 4 segments totaling 0.67 ms. Good enough to see who overlaps with whom. Not good enough to decide which version is faster.

5. Following the ledger: where the 9.9% came from

The ledger said: dense projections, transposes, scheduling. Leave MoE and linear attention alone this round. Six agents ran another 8 hours. What got merged, item by item:

  • New tile shape for dense projections, plus having projections write directly in the layout downstream ops want, and deleting a few dead transposes: about +5.2%
  • Have the last layer write only its KV cache. The last layer's q projection, attention, output projection, and MoE produce results nobody reads; only the K/V it writes into the cache are used later. llama.cpp already does this: about +2.3%
  • Turn the remaining ~110 transposes into tiled writes so memory writes coalesce: about +1%
  • Submit work to the GPU right after the first layer, fill embeddings in parallel, and a couple of other small things: about +0.5% each
baseline median 833.85 → winner 916.32 tok/s, +9.9%
vs llama.cpp: 916.32 ÷ 927.9 = 0.988

The most counterintuitive line in the round's write-up: copying llama.cpp's matmul tile shape outright was slower in every variant, -1.8% to -2.4%. So I computed the actual throughput of my biggest dense kernel:

970 GFLOP ÷ 155.9 ms ≈ 6.22 TFLOPS
llama.cpp's equivalent op, timed alone: about 6.2–6.4 TFLOPS

Even. That kernel was never worse than theirs. The "dense is 34% slower" in the ledger actually came from things outside the kernel: transposes, dead work, and the CPU not overlapping with the GPU.

I only knew to attack dense after doing the ledger, and I only learned the dense kernel was fine after attacking dense. Each layer I peeled back overturned the intuition from the layer before.

6. Which numbers here aren't solid

  • The replay adds a barrier on each side, so that category can no longer run concurrently with its neighbors. Each category's in-graph time is therefore an upper bound that includes a concurrency penalty. llama.cpp runs concurrently by default, so the penalty hits it harder, which means the +63.9 ms dense gap can only be understated.
  • When I replayed llama.cpp's small ops like add and norm, the whole graph got faster (1078 tok/s). Those nodes do in-place writes or fusion, and replaying them changes the results, so I didn't take numbers for those categories.
  • The attention row isn't apples to apples. My 19.1 ms includes q/k normalization, rotary position encoding, and splitting; llama.cpp's 5.8 ms is just the flash attention kernel itself. Not all of that 13 ms belongs to attention.
  • The quantization formats don't match. In this llama.cpp file the dense weights are 8-bit and only the routed experts are 4-bit; mine are all 4-bit. This reconciles two files, not two kernels at the same bit width.
  • llama.cpp's dense replay comes out at 188.3 ms, while its standalone op timings sum to about 231 ms, a 40+ ms difference. The standalone method may overestimate it. I haven't figured out why.
  • "Remove barriers and run projections concurrently": turning off concurrency in llama.cpp costs it 2.5%, so I assumed I had about that much to gain. Measured result: flat. My guess is every node in my graph is either big enough to saturate the GPU alone or under 0.05 ms, so there's nothing to hide behind. I don't have a GPU timeline to prove it.
  • While I was running the timing probes from section 4, the agents' automated evaluations were still running on the same M1. I contaminated 3 verdicts, including one that marked a good build as a regression. I added a noise gate to the evaluator afterward to throw out that window. The flip side: the 21.00 / 39.59 / 37.82 numbers in section 4 were also measured while evaluations were running, so the absolute milliseconds are contaminated and I only trust the ratios within a single run. The reconciliation in sections 2 and 3 was measured while no evaluations were running (background services were left on).

7. When two engines are N% apart, check in this order

  1. Stop. Until you've compared against the competitor using the same method, don't pick targets based on time shares inside your own engine.
  2. Export the competitor's compute graph, count each op category in the real graph, and cross-check the counts against the architecture.
  3. Compute coverage. Even near 100%, don't trust it yet. Look for two opposite errors canceling out.
  4. For standalone MoE benchmarks, check the expert distribution in the test input first. If it only touches a few experts, throw the number out.
  5. Get in-graph times on both sides with the same method: replay once in place, subtract baseline. Validate the method with "replay everything ≈ baseline" and "1/2/4 replays scale linearly."
  6. Sort by gap and only attack the top one or two. Categories where the competitor is slower than you: nobody touches them this round.
  7. When the competitor is faster on some category, do the byte math. If you read fewer bytes and are still slower, stop looking at bandwidth.
  8. Before copying the competitor's kernel shape, compute your own kernel's TFLOPS. If it's even, look outside the kernel: transposes, compute nobody reads, gaps between CPU and GPU.
  9. While your probes are measuring, don't run other evaluations on the same machine.