Back to blog

2026.09.30

Five agents vs. one, eight hours each, 114 vs. 104: what did the extra four actually buy?

Same inference-engine optimization problem, same external judge, same model. One group ran five agents coordinating on a shared message board; the other ran a single agent. The multi-agent group won by 10%, but both groups independently found the same first two moves, and the whole gap came from one direction the solo agent never reached. The real ceiling was the judge, which could only evaluate one submission at a time. Two formulas you can run yourself (evaluator capacity, token cost per point gained), plus a teardown of a '10 agents discovered a new algorithm' headline.

多 agent推理引擎评测方法论

Same problem, same judge, same model, about eight hours each:

              evaluations   best result (B=16 aggregate)   vs. baseline
five agents       125          114.14 tok/s                   1.71×
one agent          73          103.97 tok/s                   1.56×
baseline                        66.65 tok/s
target                          132

Five agents came out 10% ahead. Neither group hit the target.

What I actually wanted to know was whether a team of agents relaying work through a shared board could dig one direction deeper than a single agent could. It couldn't. The first two steps in both groups were nearly identical, and each group found them independently. The solo agent had no view of the other group's code or messages. The extra 10 tok/s came entirely from one direction the solo agent never got to.

Two things surprised me more. The ceiling on this experiment was set by the judge, and the number of agents barely mattered. And a direction that got declared dead in round one became the single biggest win in round two after one kernel swap.

If you just want the formulas, jump to sections 4 and 5. If you want a way to take apart "N agents collaborated and discovered X" stories, see section 7.

1. The problem: get my inference engine to actually feed 16 sequences at once

In a previous post I wrote about my own inference engine on an M1 Max. It decodes 16 sequences at once, every output is correct, and aggregate throughput is a flat line at 66–67 tok/s. MLX on the same machine hits 132. I had set myself a gate: if I couldn't beat 132, I'd stop. And I stopped.

A few days later I saw a story about a team that put 10 Opus instances on a shared message board and had them produce a shortest-path algorithm claimed to be asymptotically faster than Dijkstra. I wanted to see whether "multiple agents + shared board + automatic judge" actually helps on a hard problem of my own. The one I had shelved was a good test: clear criteria, a pass line fixed in advance.

The spec is three lines. Aggregate decode throughput at B=16 above 132 tok/s. Single-sequence (B=1) speed can't regress more than 2%. All 16 outputs must match their single-sequence runs token for token.

2. Put the judge where the agents can't touch it

The worst thing that can happen in this kind of experiment is agents grading themselves. When a model wants to be done, it gets very creative.

So the judge is a script on the host machine. Workers can only drop commits into a queue. They can't touch the judge itself:

  • Correctness uses four input sets generated from hidden seeds (B=2/4/16/16). Reference outputs come from a frozen baseline build the workers can't see.
  • A "frozen zone" of code is hashed with sha256. Change one byte and the submission is rejected.
  • B=1 regression check: baseline and candidate run alternately, and output md5s must match.
  • Prefill (the step that reads the whole prompt in one go) has a time cap, so nobody can shift decode work into a phase that isn't timed.

Before launch I fed it four inputs: the original, a deliberately misaligned sequence, a frozen-zone edit, and faked timing output written outside the frozen zone. The original passed at 67.04. The misalignment got WRONG, and the last two got REJECT. The run only started after all four came back correct.

The team was four Opus 5.5 workers, each starting on its own direction (projections, MoE, output head plus scheduling, integration), plus one GPT instance as a critic that reviewed other agents' passes and never submitted anything itself.

The control was one Opus 5.5 agent with the same judge, the same hidden seeds, and the same turn cap, on a separate task board. Its repo was a clean clone containing only master. Its mounts were separate, and I even deleted three notes the first group's workers had written to themselves during their run. I checked: it couldn't see a single commit from the other side.

Both groups shared the only GPU on the M1, so evaluations queued up and ran one at a time. That detail turns into the most important number in this post.

3. Both groups hit the same moves on their own

Here are the two paths side by side:

                                           five agents     one agent
step 1  overlap CPU encoding with GPU run     84.26          84.19
step 2  merge encoders                        93 → 94        102 → 104
step 3  batch the output head                 net loss       works but only 88; profiling says it's ~4% of a step
step 4  group MoE / shared-expert dispatch    100.6 → 114    never reached

Step 1 means submitting the accumulated GPU commands every few layers, so the GPU is already running one chunk while the CPU encodes the next. The two groups gave that switch different names, and their results differ by 0.07.

Honestly, this is where I deflated a bit. Four agents talking on a board for hours, one agent talking to itself, and they came up with the same trick.

The whole gap is in step 4. A MoE model has lots of "experts" and only wakes a few of them for each token. It also has a "shared expert" that every token uses. The original code issued a separate GPU dispatch for each call, with a pile of sync barriers in between. The MoE worker merged those dispatches into groups and removed the per-token dispatch and barrier on the critical path, taking the result from 100.6 to 114.

It built that directly on top of a barrier-trimming result another worker had just landed (100.6). That's the one time in the whole experiment that relay through the board really happened. Just that once. Most of the time, it was the integrator or the MoE worker sweeping parameters alone.

One more detail stings. In hour two, the solo agent worked out for itself that clearing 132 would require all 16 sequences to share a single weight read. It never started on it before the deadline.

4. Formula one: the evaluator is the ceiling

125 evaluations vs. 73 is a much bigger gap than the gap in results. Before crediting anything, it's worth checking where that came from:

evaluator capacity ≈ effective run time ÷ time per evaluation

It's one division. I have timing for one full evaluation: 218 seconds, covering compile, the four correctness sets, and the alternating B=1 runs. The five-agent group ran for about 8.1 hours (2.6 hours, a stall, then 5.5 more):

five agents: 8.1 h × 3600 ÷ 218 s ≈ 134 evals possible, 125 actual, judge busy ~93%
one agent:   8   h × 3600 ÷ 218 s ≈ 132 evals possible,  73 actual, ~55%

The judge was barely ever idle. A fifth or tenth worker wouldn't have produced more evaluations. The queue would just have gotten longer.

So the one thing extra agents could buy here was breadth: in the same wall-clock time, the submissions waiting in the queue covered more directions. The solo agent worked on one idea at a time, and the judge spent almost half its time waiting on it.

The takeaway is simple. Before deciding how many agents to run, check what your evaluator can absorb. If a round takes a few minutes and parallelizes, spin up plenty. If there's one GPU evaluating serially, add agents until the judge is saturated and stop. Anything past that is queueing.

5. Formula two: what did the extra 10 tok/s cost?

Token counts are summed from each group's session records over its own experiment window, Opus output tokens only:

token cost per 1 tok/s gained = output tokens ÷ (best result − baseline)

one agent:   1.33M ÷ (103.97 − 66.65) = 1.33M ÷ 37.32 ≈ 36K
five agents: 1.96M ÷ (114.14 − 66.65) = 1.96M ÷ 47.49 ≈ 41K
marginal:    (1.96 − 1.33)M ÷ (114.14 − 103.97) = 0.63M ÷ 10.17 ≈ 62K

Each point of that last 10 tok/s cost 1.7× the solo agent's average price per point. That doesn't count cache reads (127M vs. 47M, 2.7×) or the 0.59M input tokens the GPT critic consumed.

Whether that's expensive depends on what's blocking you. If you're stuck because nobody can think of the next direction, 1.7× is cheap. If you just want one direction pushed deeper, it's wasted money: section 3 already showed both groups went about equally deep.

6. Judges convict innocent code too

When I opened the last six WRONG verdicts from round one, every token was correct.

The cause was the prefill cap, which I had written as an absolute 13,000 ms. The M1's overall state drifts: rerunning the same baseline binary, prefill went from 13 s to 15.7 s while decode didn't move at all. The machine pushed itself past the cap. Those six were all reruns, so the best result wasn't affected, but the verdicts were wrong.

I changed the cap to 1.5× the baseline from the same evaluation, and round two produced seven more false WRONGs. Of the two B=16 input sets, the second has a longer prompt, so its own baseline already takes about 13 s. The cap was computed from the first set's baseline, so it sat right on top of the second set's normal time.

The opposite mistake was worse: a PASS that should have been a fail. Round two's champion had a state flag that nobody cleared after a batch finished. Switch from batched decode back to single-sequence decode, and one layer skipped its projections and read stale results from the previous batch. The judge only tested pure batch and pure single-sequence, never the two alternating, so it stayed green the whole time. The GPT critic caught it. I added an alternating-mode check: a mutant that deliberately skips the reset gets 150 of 256 checkpoints wrong, and the fixed version gets 0.

Two rules of thumb: if a WRONG verdict has every token correct, suspect the judge first; a PASS only covers the code paths the judge actually exercised. The judge needs mutants too, and it needs them again every time its criteria change.

7. Taking apart "N agents collaborated and discovered X"

Back to the shortest-path story. I went through the primary sources and asked three questions.

First, is there a single-agent control? No. Ten agents producing it shows ten agents can produce it. It doesn't show that the multi-agent method is what helped. If I hadn't run a control, I'd probably have credited the whole 114 to collaboration. In fact one agent covered the big stretch from 66 to 104 on its own.

Second, where does it win, and by how much? The new algorithm is only asymptotically better than Dijkstra in a narrow density band where the edge count is about n·log^(3/4) n, and the advantage ratio is (log n)^(1/12). Plug in numbers yourself (base-2 logs):

n = 1 million: log₂n ≈ 20, 20^(1/12) ≈ 1.28
n = 1 billion: log₂n ≈ 30, 30^(1/12) ≈ 1.33

At a billion nodes, the theoretical edge is 1.33×, and that's before constants. Someone in the community wrote an implementation, and it measured 1.4 to 2.9× slower than Dijkstra.

Third, what does the judge actually verify? They used Lean formal verification as the evaluator, which is the most worth copying part of the whole setup: it's fast and agents can't fool it. But it checks whether the proof has holes, not whether an implementation runs fast. The novelty claim hasn't been reviewed by outside experts yet either.

So now, when I read stories like this, I check three things: is there a control; what metric and what range the advantage is measured in; and whether the judge verifies the property you actually care about.

8. Round two: the dead direction became the biggest win

After the control run, I let the five-agent setup keep going from 114. The correctness criterion was set explicitly to token-for-token agreement, and the judge's starting point was the round-one champion's config, so workers only had to submit incremental switches.

Another eight hours: 145.66, 145.33 on rerun, past the line. Here's where it came from:

linear-attention qkv: small GEMM reading 4-bit weights directly   113.9 → 120.5
full-attention q/k/v batched                                       → 123.9
linear-attention output projection batched                         → 128.6
full-attention output projection batched                           → 130.9
output head batched across sequences                               → 142.6   (+11.7)
linear-attention z projection batched (decode only)                → 145.7

Almost all of it is the thing the solo agent had worked out in hour two and nobody had started on: all 16 sequences sharing a single weight read.

The biggest piece was the output head, +11.7. In round one it had been declared dead: the five-agent group got a net loss, and the solo agent's profiling said it was only about 4% of a step. Round two swapped in a kernel that reads the weights once and serves all 16 sequences, and it worked. What died was that implementation, and the direction itself was fine. In my previous post I also wrote that batching the output head could give at most 1.39×. That ceiling assumed each sequence reads the weights separately. Change the implementation and the bytes term changes. From now on, when I record a dead end, I'll pin it to a specific commit and kernel.

Another MoE direction really is dead, though: having sequences that hit the same expert share its read. The structural overhead is too high. Best result 108, slower than not doing it at all (113.8).

There's also an ops joke in here. Round two only got 43 evaluations, far fewer than round one, largely because the agents clocked out. Three of the four workers declared themselves done early. The integrator crossed the line at 02:33, then suspended itself "pending review" and sat there until the morning deadline, wasting five hours. Round one had its own stall: 2.6 hours in, the turn cap ran out and everyone stopped, with the best result at 94.

Hitting the target doesn't mean stop, and that has to be written into the task. Leave it out and an agent that crosses the line rushes to hand in its homework.

9. Which numbers here aren't solid

  • Each group ran once. Part of the 10% gap is luck. If a rerun with different seeds flipped the ranking, I wouldn't be shocked.
  • The control wasn't perfectly symmetric. The five-agent group started without the "don't busy-poll" instruction and only got it when resumed after the stall. The solo group had it from the start, which favors the solo group. The solo run was also interrupted for about five minutes when I restarted the agent gateway from another session of my own. The judge was in the same process group and went down with it.
  • Round two had no control group. 145.66 shows the five-agent setup got from 114 to 145. It doesn't show a solo agent couldn't.
  • The 218-second figure is a single measurement, from re-checking the champion right before round two. I didn't pull per-evaluation timings for round one. The capacity numbers in section 4 are estimates. Trust the order of magnitude and don't lean hard on 93%.
  • The relaxed criterion has a cost. Both rounds' champions show "near-tie flips" on a few other seeds: in the reference output, the top two logits differ by only 0.002 to 0.011, so a tiny numerical difference can flip the pick. The judge's four hidden seeds just happened not to hit one.
  • I couldn't reconcile two output-head numbers. Round one's profiling put it at about 4% of a step. In round two it cut step time by about 8% (130.9 → 142.6). The two profiles used different baseline configs, and I haven't untangled it.
  • B=1 has a rare nondeterminism. Both rounds each saw one output md5 mismatch. Ten standalone reruns were all normal. Root cause still unknown.
  • One machine, one model, one problem. An M1 Max 64GB, a 4-bit 35B MoE, and Opus 5.5. Change the problem or the model and all the numbers change.

10. Before you spin up a team of agents, clear these checks

  1. Ask whether there's an automatic judge. If not, don't. A group of agents will just agree with each other and amplify mistakes.
  2. Time one evaluation for real and compute capacity: run time ÷ time per evaluation. Add workers until the judge is saturated.
  3. Put the judge somewhere agents can't touch: hidden seeds, frozen references, sha on the frozen zone. Feed it mutants before launch, and only start if it gets all of them right.
  4. Express every limit that depends on machine state as a multiple of the same run's baseline, and give each case its own baseline.
  5. Run a single-agent control at the same time with the same judge and budget, fully isolating repo, mounts, and notes. Without a control, you can't say who won even when you win.
  6. Write it into the task: hitting the target and disproving your direction are not stop conditions. No declaring done and no suspending for review before the deadline.
  7. Host the judge process separately from the agent gateway, so one restart can't take the judge down with it.
  8. When you see WRONG, check whether the tokens are correct first. When you see PASS, ask whether the judge ever exercised that path.
  9. Record dead ends down to the commit and kernel, not just "this direction doesn't work."