Same model, same code, same 60 random seeds, nothing changed, run twice. Average round reached: 6.58 the first time, 6.32 the second.
A 0.26-round gap looks like nothing. But on one of those seeds, the first run made it to round 15 and the second died in round 3. The two runs split at step 7, in the shop, over whether to buy one joker.
Later I added one sentence to the prompt and the average jumped to 9.12. Paired over 60 seeds, 36 got further and 7 got less far. That effect is real. Now cut those same 60 games into three blocks of 20 in seed order: the first block gained only 0.65 rounds, p=0.30. Every experiment I had run before that used 20-game batches. Had I written that sentence a week earlier, I would most likely have thrown it out as "no effect."
If you want the formulas, go to Sections 3 and 5. If you want to know why running the unchanged config twice still isn't enough, Section 5.
1. A 35B model playing Balatro
Balatro is a poker-hands roguelike. Each round has a score target. You play poker hands for points, clear the target and move on, miss it and you're dead. Between rounds you visit a shop and buy jokers, which multiply your score. Late game runs entirely on jokers. Every three rounds make one "ante," and the targets climb hard as antes go up.
I had a 35B local model play it on a handheld. The setup is cheap: a game mod lists every legal action as candidates A, B, C, and so on, and the model reads exactly one token. Whichever letter gets the highest probability is the move. About 150 ms per step.
One metric: which round the run reaches. Two baselines: a greedy player that always plays the highest-estimated hand, and pure random.
First batch, 20 seeds: the model averaged 5.90 rounds, greedy 3.25, random 1.95. Paired by seed, the model went further than greedy in 16 games and less far in 0. That gap is big enough that you don't need statistics to see it.
The trouble started with the next step. Every change after that was worth a round or two, and you can't eyeball that.
2. The first "improvement," and the first backfire
The first change looked obviously right. My hand-written score estimate ignored jokers. With 5 jokers in hand, a pair of aces was estimated at 64 points and actually scored 704. The model reads one token, so it can't multiply jokers out in its head. It just picks according to my estimates, which meant it had been deciding with wrong numbers all along.
So I installed an open-source scoring simulator mod and let the game compute the real score of every possible play. I checked two hands: simulated 704, played 704; simulated 850, played 850.
Same 20 seeds again: average 5.90 → 7.00, up 1.1 rounds. Runs reaching ante 4 went from 4 to 6.
Looks good. Then count the pairs: 9 further, 4 less far, 7 tied, p=0.27. Not significant.
The second change was a shop prompt that explained interest and annotated each purchase with how much money you'd have left. Average went 7.00 → 6.60, 8 better and 9 worse. Breaking down the behavior, it clearly listened: the share of shop exits with under $5 dropped from 40% to 27%. But to bank interest, it cut joker purchases from 2.5 per game to 1.4. Jokers are where the points come from. It saved its money by giving up firepower.
Both times I was squinting at a 20-game average and guessing. That's when it hit me that I had no idea how big a difference 20 games could even detect.
3. Run the unchanged config first
This is an A/A test: both arms are identical, so any difference is pure noise.
I first turned on the game's fast-forward and headless modes, which cut a game from 109 seconds to about 35. Of that, the model's thinking took only 6 seconds. The rest was animation. Then: same config, 60 seeds, two runs.
| Run 1 | Run 2 | |
|---|---|---|
| Average round | 6.58 | 6.32 |
| Reached ante 4 | 13 games | 11 games |
| Paired | 12 seeds run 1 further | 15 seeds run 2 further |
33 seeds ended on the same round both times, p=0.70. No overall difference, as it should be.
The number that matters is the standard deviation of the per-seed difference between the two runs: 2.51 rounds. Same config, same seed, and a gap of two or three rounds in a single game is routine. The cause is run-to-run jitter in the 35B itself: on the same game state, its confidence can differ by up to 0.11 between runs. When two options are close, it picks A once and B the next time, and every card drawn after that is different. Only 11 of the 60 seeds went through both runs without a single divergent step. Among the 49 that did diverge, the median split came at step 9.
With that number you can do some arithmetic. The smallest improvement a paired comparison can reliably detect is roughly:
minimum detectable difference ≈ 2.8 × SD of paired differences ÷ √(number of seeds)
The 2.8 comes from the usual pair of thresholds, 95% confidence and 80% power. It's an approximation, so don't sweat the decimals. Solve for seeds instead:
seeds needed ≈ (2.8 × SD of paired differences ÷ difference you want to see)²
Plug in the A/A value of 2.51:
- 20 seeds: minimum detectable difference ≈ 2.8 × 2.51 ÷ 4.47 ≈ 1.57 rounds
- 60 seeds: ≈ 0.91 rounds
- To see +1 round you need about 49 seeds; for +2 rounds, about 12
Back to Section 2. In the exact-scoring run, the SD of paired differences over 20 seeds was 4.17, which means 20 games could detect at best a 2.6-round improvement. It gained 1.1, less than half the threshold. That p=0.27 means this batch couldn't answer the question. It never said the change didn't work.
4. One sentence, +2.5 rounds
With a ruler in hand, I went looking for the actual bottleneck before changing anything, instead of guessing.
Going through the logs of those 120 A/A games: at death the model held 2.0 jokers on average. Across all shop visits, it left the shop 654 times and bought a joker 247 times. The joker bar has 5 slots, and most of them sat empty for most of the game. Late game it didn't have the firepower, and once the targets climbed, it died.
So I added one sentence to the shop prompt: "You still have N empty joker slots. Empty slots are waste. Buy jokers first."
60 seeds:
| A/A run 1 | A/A run 2 | With the sentence | |
|---|---|---|---|
| Average round | 6.58 | 6.32 | 9.12 |
| Reached ante 4 | 13 | 11 | 34 |
| Paired against it (further:less far) | 36:7 | 35:8 | |
| p | 9e-6 | 4e-5 |
Jokers held at death went from 2.0 to 3.7. The mechanism and the result line up, so this one I can keep with confidence.
Then I did the split from the intro: cut the 60 games into three blocks of 20 in seed order and compared each block against A/A run 1.
| Block | Average gain | Further:less far | p |
|---|---|---|---|
| First 20 | +0.65 | 10:5 | 0.30 |
| Middle 20 | +3.25 | 13:1 | 0.002 |
| Last 20 | +3.70 | 13:1 | 0.002 |
The same real +2.5-round effect showed only a quarter of itself in the first block. Every batch I'd run before was 20 games, and with that kind of luck I would have written "no significant effect." Put bluntly: a 20-game experiment can miss the best change I made all week.
5. A/A noise is only a lower bound
I thought the 2.51 from Section 3 was good enough. The very next experiment proved me wrong.
The next change distilled human strategy guides into four rules in the system prompt: with no clear direction, play flushes and pairs; judge jokers by how many times they multiply your total score; build your economy in the first three antes; use planet cards early.
Against the previous version, 60 games: average 9.12 → 9.22, paired 25:25, p=1.
The behavior genuinely changed. Flush plays went from 11% to 34%, two-pair from 30% to 16%, median money on shop exit from $6 to $10. The result didn't move at all. Hand types and economy just aren't the bottleneck right now.
The problem is the SD of paired differences: 4.82. Nearly double the A/A value of 2.51. The joker-first run was 4.22.
Look at where the runs first diverge and the reason is obvious:
| Comparison | Seeds that never diverged | Median step of first divergence | SD of paired differences |
|---|---|---|---|
| A/A (nothing changed) | 11/60 | step 9 | 2.51 |
| Add "buy jokers first" | 0/60 | step 6 | 4.22 |
| Add strategy guide | 0/60 | step 2 | 4.82 |
In an A/A test, the only source of branching is model jitter. Many games go several steps before splitting, and after splitting they often end up in the same place: of the 49 seeds that diverged, 22 still died on the same round. Once you actually change the prompt, the two arms differ from the very first decision of the game, every card after that changes, and most of the benefit of pairing on the same seed is gone.
So planning sample size from the A/A SD is optimistic. Redo it with the 4.2 to 4.8 from real comparisons:
- 60 seeds: minimum detectable difference ≈ 1.5 to 1.7 rounds
- To see +1 round: 140 to 180 seeds
At about a minute per game with both arms run once each, that's roughly 4.7 to 6 hours. The correct reading of the guide run's "+0.10, p=1" is: if the true effect is smaller than 1.7 rounds, these 60 games can't see it.
That batch also produced the first full clear: one seed picked up a joker that adds 2 mult for every $5 you hold, which paired perfectly with the guide's save-money play. It won holding $206. One clear in 60 games. It looks great and tells you nothing statistically.
6. Reading "AI beat game X" or "one prompt tweak gained X%"
These are the four questions I ask now:
- How many games? One clear, one screen recording, is a sample size of 1. My clear was 1 of 60. Anyone could screenshot it and post "local model beats Balatro."
- Was it paired? Comparing on the same seed and starting conditions cuts noise a lot. If both arms ran their own random games and you compare averages, the number of games needed goes up further.
- Did they run an unchanged-vs-unchanged control? If not, they don't know how big their own noise is. In my A/A test, with nothing changed, a single game swung by 12 rounds.
- Did they report behavior or results? In the guide run, every behavior metric moved and flush rate tripled, yet results changed by zero. Reporting only "the model plays differently now" is reporting nothing.
Running Section 3's formula backwards is fast too. Someone says "average went up 1 round over 20 games." Take 2.5 as a lower bound on noise, and the minimum detectable difference at 20 games is already 1.6. Their number sits below the threshold, so you can treat it as if you never saw it.
7. What's wrong with these experiments
- The evaluator itself had a hole. Digging into causes of death later, I found a rule in my shop candidate generator: "when the joker bar is full, don't list shop jokers." There were 241 shop visits with a full bar, 122 of them with $25 or more in hand, and zero joker swaps. Every game ran on whatever 5 jokers it scraped together early. Every number above was measured inside that hole. The fixed version ran only 17 seeds before I stopped it, so its effect is still unknown.
- One model, one way of playing. All of it is the same 35B, Red Deck, lowest difficulty. Change the model or the difficulty and the noise changes; the SD has to be re-measured.
- "Round reached" is a coarse metric. It's an integer with a ceiling, and dying at 95% of the target counts the same as dying at 5%. Using "final score as a fraction of target" might shrink the SD. I didn't try.
- The formula is a normal approximation. Significance here was actually judged with a sign test, which only counts wins and losses, not margins. I didn't check case by case how far the two diverge. Treat the seed counts from the formula as orders of magnitude, not exact values.
- The three-block split is one split. The first 20 happened to be weak; cut differently and you'd get different numbers. It shows that 20 games can miss a real effect. It can't estimate how often that happens.
8. Next time you tune prompts for an AI agent, do it in this order
- Run the current config twice, unchanged, and compute the SD of the per-seed difference. That's your noise floor.
- Before any change, estimate seeds with (2.8 × SD ÷ difference you want to see)², multiply by game length, and see if you can afford it. For real changes, budget 1.5 to 2 times the A/A SD.
- Always pair on the same seed and only count who got further.
- Use existing logs to find the bottleneck first (mine was "only 2 jokers at death"), then write the one prompt line that targets it.
- Report behavior metrics and results separately. If behavior moved and results didn't, you changed something that isn't the bottleneck.
- After each experiment, re-audit the candidate generator for any class of action the model can never pick. In my 241 full-bar shop visits, the model couldn't have swapped a joker even if it wanted to.