Back to blog

2026.10.01

One renamed stat took the same model from 59% to 72%: what were my three AI game runs actually comparing?

I had a 35B local model play Reigns: Three Kingdoms and ran three live sessions comparing thinking, no thinking, and random choice. They came out indistinguishable. Only after turning the logs into 360 offline questions did I find that I had misnamed one of the four stats, so the model had been reasoning about the wrong quantity the whole time. Fixing that one name moved the same question set from 59.2% to 72.1%. Three calculations you can run yourself: how to split an accuracy gain by stat, where the question set tops out, and how many games a live comparison really needs.

LLM评测游戏方法论

Same 35B model, same 360 questions, same temperature, same prompt structure. I changed the name of one stat, and accuracy went from 59.2% to 72.1%.

I had picked that name myself. Before I caught it, I spent a full day and three live game sessions comparing "thinking on," "thinking off," and "pick at random" to see which played best. They were indistinguishable. For a while I blamed the model. In the end the cause was mine: I had mislabeled one of the game's stats, and the model had been reasoning from the wrong meaning since turn one.

If you just want the method, Section 5 has a sample-size formula you can run yourself. If you want to see exactly what the wrong name broke and how the error surfaced, start with Section 3.

1. The game and the three live runs

Reigns: Three Kingdoms is the Three Kingdoms-era entry in the Reigns series. The whole game is one move: a card shows an event, and you swipe left or right. Four stats run across the top. If any of them hits 0 or 100, your ruler dies and the next one takes over. Your score is how many days each ruler survives.

A 35B local model played it on a handheld. It saw the card text, both options, and the four stat values (I measured those from screen pixels), then chose left or right. A script handled screenshots, swiping, and the occasional duel minigame.

I named the stats from their icons: a wheat sheaf became "grain," a head became "people," a blade became "army." The fourth icon is a small tower. It looked like a storehouse to me, so I called it "treasury."

Then I ran three sessions. Only rulers who played to the end are counted:

Setting Complete reigns Mean days Danger turns, safer : worse
Thinking off 5 294 17:17
Random 4 315 11:11
Thinking on 3 397 9:11

A "danger turn" is my definition: some stat is at or below 20, or at or above 80, before the choice. Pick the right direction there and the ruler lives a while longer. Pick wrong and that's the end of that ruler.

With thinking off, the model went 17:17, which is exactly a coin flip. It could say "people at 9, extremely dangerous." But whenever neither option obviously touched that stat, it chose more or less at random. The random baseline averaged 315 days, 21 more than the model.

The thinking run looks best at 397 days, but it covers only 3 reigns, and in 135 of its 225 turns the reasoning hit the length cap and fell back to non-thinking mode. That run is contaminated, so it doesn't support any conclusion.

2. Live play is too slow, so get a faster evaluator

A live session makes a noisy ruler. Each reign lasts 30-odd turns, most of them irrelevant to survival. A 220-turn session takes 50 minutes, and only about three dozen of those turns tell you anything about directional judgment.

So I turned the existing logs into an offline question set. For every turn I had the stats before and after the choice. Subtract one from the other and you know which stats that option pushed up and which it pushed down. A question looks like this: here is the card text and both options; the option the player picked moves "people" and "treasury"; for each one, did it go up or down?

The filters were strict. At least one stat must move by 4 or more (smaller changes can't be told apart from pixel-reading jitter). Any jump of 45 or more is discarded, because that's usually a new reign or a misread. Repeated (card, option) pairs are merged, and pairs whose outcomes contradict each other are dropped. The result: 360 questions with 569 "up or down" calls.

Two baselines come first, because no score means anything without them. Coin flip: 50%. Always guessing each stat's most common direction: 53.8%.

Thinking off, original names: 59.2%. Above baseline, but not by much.

Split it by stat and the problem is obvious:

Stat Correct, original names Correct, fixed names
Stat 1 (wheat) 59/85 64/85
Stat 2 (people) 139/206 149/206
Stat 3 (army) 51/84 56/84
Stat 4 (tower) 88/194 141/194

The fourth stat was right 45.4% of the time, worse than a coin. On 194 calls that's statistically indistinguishable from guessing (p ≈ 0.22). Still, a 35B model reading plain event text and doing no better than chance on one stat is a red flag.

3. The tower was never a storehouse

A strategy guide, cross-checked against the in-game death messages, settles it. Left to right, the four stats are wealth, popularity, might, virtue:

  • Wealth: material resources, meaning food, timber, stone and coin. Too little and you starve; too much and others come for it.
  • Virtue: propriety and moral standing (mercy, restraint, observing the rites). Too high and you're too gentle to survive; too low and your followers abandon you.

When the fourth stat maxes out, the death message reads, roughly: "Your followers held so strictly to propriety that not one of them fought back." That death comes from too much virtue. Money has nothing to do with it.

Now look at how the model failed under the old name. Across the 194 fourth-stat calls, it guessed "down" 160 times. The true label was "up" in 104 cases, and it called 88 of those "down."

The pattern isn't hard to guess. Giving alms, granting pardons and keeping the rites all raise virtue. But the model had been told this stat was "treasury," so it was most likely thinking that doing good costs money and the treasury goes down. That's a perfectly reasonable inference. I had just handed it the wrong quantity to reason about.

The first stat was also misnamed: "grain" instead of "wealth." But those two mean nearly the same thing, so that stat barely suffered. A wrong name does real damage when it flips the meaning.

With the names fixed, same questions, same model: 72.1%. The fourth stat went from 88/194 to 141/194.

All three live sessions had run under the wrong name. I had handed the model a mislabeled manual and then tested whether it read better with glasses on.

4. What one word is worth: splitting the gain by stat

Total accuracy is a weighted average, so the gain splits exactly:

change in total accuracy = Σ (change in correct calls per stat) ÷ total calls

That's a subtraction and a division. Plug in the numbers:

Correct calls across four stats: 59+139+51+88 = 337 → 64+149+56+141 = 410
Total gain: (410 − 337) ÷ 569 = 73 ÷ 569 ≈ 12.8 points
From stat 4 alone: (141 − 88) ÷ 569 = 53 ÷ 569 ≈ 9.3 points

Of the 13 points, more than 9 came from that one word.

The split alone isn't enough, so I also ran a paired test. Both settings answered the same calls, so only the disagreements matter. On the fourth stat, 26 calls were right under the old name and wrong under the new one, and 79 went the other way: p ≈ 2×10⁻⁷. The other three stats split 5:10, 11:21 and 2:7, and none of those is significant. The numbers back up the claim that the rename's gain sits almost entirely on the fourth stat.

For reference, a small model with about 2B effective parameters scored 57.3% on the same set. Paired against the misnamed 35B (59.2%) it splits 167:156, p = 0.58, so they sit in the same tier. One wrong word dragged a 35B model down to small-model territory.

Compare that with turning on thinking. After the fix, thinking took the 35B from 72.1% to 77.2%, paired 77:48, p = 0.012, so it's a real gain. But median time per question went from 0.3 s to 12.8 s, more than forty times longer. The rename cost nothing and bought 13 points. Thinking cost 40x the time and bought 5.

Full ranking, for scale:

Model / setting Accuracy Median s/question
gpt-6-astra 83.7% 2.6 s
gpt-6-luna 77.7% 1.9 s
Local 35B, thinking on 77.2% 12.8 s
Local 35B, thinking off 72.1% 0.3 s
Local 35B, thinking off, wrong names 59.2% 0.3 s
~2B small model 57.3% 0.1 s

5. How many games does a live comparison need?

Looking back, the three live runs couldn't be separated partly because of the wrong name and partly because the samples were simply too small. You can work that out before you start.

The cheapest test for "which of two settings judges direction better" is a sign test: count the cases where A is right and B wrong, and the reverse. If one side truly wins a fraction q of those disagreements, getting to p below 0.05 takes roughly this many disagreements:

n ≈ (1.96 ÷ (2q − 1))²

That's the normal approximation. The exact binomial test needs a few more:

q = 0.65 → approx 43, exact 44
q = 0.60 → approx 96, exact 101
q = 0.55 → approx 384, exact 390

After the rename, a live session went 13:7 on danger turns, safer to worse. That looks good, but it's only 20 calls, p ≈ 0.26. Even if 65:35 is the true rate, you'd need 44 calls before it means anything.

Here's the production rate from live play: 220 turns in 50 minutes yielded 34 danger turns where the stat actually moved one way or the other, so about 40 per hour. If the true effect is 60:40, just showing it beats a coin flip takes about 100 of those: two and a half hours, for a single setting. Comparing two settings on mean days survived is slower still. My estimate was around 30 reigns per setting, which starts at 8 hours.

The question set gives you all 569 calls in one pass, and with thinking off a score comes back in a few minutes. It measures the same directional judgment, runs dozens of times faster, and supports paired tests. So the sensible order is to find the direction on the fast evaluator and use live play only for final confirmation.

6. How to read "the AI survived 300 days"

Game-agent writeups online usually look like this: "the model played a few games and averaged X days / reached a score of Y." After the tables above, ask three things first:

  1. What does random play get? In this game, random choices average 315 days and the 35B with thinking off averages 294. Without a random baseline, "300 days" tells you nothing.
  2. How many games? My "thinking on, 397 days" run is 3 reigns. That's a small fraction of what a sign test needs.
  3. Does the reported number track the skill being claimed? Mean survival is diluted by the many turns where nothing is at stake. If you want to measure judgment, count right and wrong directions on the turns that matter.

One more, the harsh one: when the score is low, check what you fed the model before you blame the model. I came close to publishing "in critical moments the 35B is no better than a coin flip." The model was fine. The manual I gave it was wrong.

7. What's wrong with this question set

To be upfront about it:

  • The labels are noisy. About 10% of calls (62 of 569) are ones where the two strongest models agree with each other and both disagree with the label. Spot checks mostly turn up counterintuitive game design. For example, "stop the food rations" actually dropped wealth by 22. So the ceiling for this set is roughly 90%, not 100%. The estimate is simple: ceiling ≈ 1 − (share of calls where the two strongest models agree and are both wrong). When merging duplicates, 6 of 366 groups had contradictory outcomes and were dropped, which is another sign of noise.
  • The questions are easier than live play. The set tells the model which stats will move and only asks for direction. In live play it has to read the screen and the preview markers itself. 72% here does not mean 72% of critical live choices go right.
  • The questions come from my own games. They only cover cards the model happened to draw. Events that never came up aren't in the set.
  • The wrong-name control was run for one setting only. Only the 35B with thinking off was tested under the original names. Every other model ran with the corrected names from the start. I haven't measured how much the wrong name would hurt a stronger model.
  • The live gain from the rename isn't established. After the fix, one live reign lasted 950 days, against a previous best of 600. That's tempting, but it isn't significant, and I'm not treating it as a result.

8. Next time a model plays any game with stats, do it in this order

  1. Look up each stat's official name and meaning, and cross-check a guide against in-game text. Don't guess from icons.
  2. Before scoring anything, set two baselines: random choice, and "always guess each stat's most common direction."
  3. Turn existing games into offline questions and break accuracy down by stat. If one stat scores below baseline, suspect your definition before you suspect the model.
  4. Compare two settings only with a paired test, counting only the calls where they disagree.
  5. Before a live run, estimate the disagreements you'll need with n ≈ (1.96 ÷ (2q − 1))², then divide by how many you produce per hour to get the hours it will take.
  6. Estimate the ceiling from the questions where two strong models agree and are both wrong.