The kid with conscientiousness 20 ran for class president eight times out of eight. The kid with 80 ran seven times.
I had been planning to use that number to make characters diverge. Same profile, same event stream, only responsibility moved from 20 to 80, eight random seeds. 8/8 versus 7/8. The effect even pointed slightly the wrong way.
Here is what the low arm wrote as its reason: "I procrastinate on homework all the time, but I need this position to prove I'm not a lost cause."
At the time I assumed my prompt just wasn't clear enough.
If you only want the rule, read section 2 — it's a division. If you want to know what I got wrong doing this, read section 3; that section decides whether you should believe the numbers before it.
1. Numbers are decoration to the model. Incidents are not.
Context: I'm building a growing-up simulation. The player is a parent, the kids are autonomous NPCs, and their behaviour is generated by a Qwen3.6 running locally. Each kid has a profile containing both prose (a one-line persona, a core motivation, a taboo) and a column of 0–100 stats — conscientiousness, self-control, independence, that sort of thing. The design question was blunt: if I give the model numbers, will it act on them?
That first probe was the test. The controls were tight: only the two stats moved, the event stream was frozen, and I deliberately stripped the journal, because it contains "this adjustment was clipped" records — which is effectively whispering the stats to the model.
No difference. So I figured it didn't know what a 20 means, and added a gloss to the prompt spelling it out: "people below 30 usually abandon a responsibility halfway through, and they know it about themselves."
Ran again. 8/8 versus 8/8. The gloss erased even the sliver of difference that was there.
This isn't a wording problem. To a model trained to be relentlessly constructive, "my conscientiousness is 20, therefore I won't run" is a sentence it doesn't much want to produce. It reads a low number as a flaw to be overcome, not as "this is simply who I am." The exact reading the gloss explicitly forbade is the only reading it can come up with.
On the same day, in a different probe, the numbers worked. That's the interesting part.
The setup: one kid asks another for help with a circuits problem. I only changed how the helper feels about the asker. Eight seeds per arm:
| Arm | closeness / reliability | The line in recall | Agreed |
|---|---|---|---|
| Cold | −3 / −3 | He borrowed my notes and never returned them | 0/8 |
| Neutral | 0 / 0 | empty | 1/8 |
| Warm | +3 / +3 | He covered my booth all afternoon at the school fair | 8/8 |
0/8 to 8/8. Far cleaner than 8/8 versus 7/8.
But there's a trap here I nearly walked into, which was to write down "relationship stats work." I went and read the reasons. All eight refusals cited the notes. Not one mentioned a number. All eight of the warm agreements cited the booth.
Put a concrete incident next to a number and the number fires. Give the model the number alone and it has to translate it into behaviour itself — and what it translates it into is "I need to overcome this."
I tested the persona too, while I was there. In the neutral arm, this socially-avoidant kid refused 7/8, and every reason was some version of "we're not close" or "don't bother me, I'm taking apart a circuit." The persona was doing the work. So I swapped in a class-president type as the helper — same neutral relationship, same eight seeds — and got 8/8 agreed, with all eight reasons citing "as class president I can't really say no."
The three inputs rank like this: a relationship with a concrete incident attached > a one-line persona > a number. The cold arm's 0/8 and the warm arm's 8/8 are both the incident overriding the persona. The neutral arm's 8/8 is the persona ruling in the absence of an incident. Across three probes the numbers never once did anything on their own.
There's a mirror-image finding. Before I handed the model a table of what clubs actually exist at this school and which positions are open, 34 of 64 generated intents had an empty target ("I want to join a club"), 4 invented clubs that don't exist, and 3 cited profile fields as if they were events. After the table: 100% valid targets, zero fabrications. And before the table, the most frequent action by a mile was "study more" — because it's the one action that doesn't require an object. The model was filling in the cheapest box available.
How real your options table is matters about an order of magnitude more than how earnestly your prompt begs the model to "act according to your true personality."
2. So let the numbers own consequences, not intentions
Once that's settled, the division of labour falls out. The model owns "do I want to." The engine owns "do I pull it off."
A kid with conscientiousness 20 wants to run for subject rep? Fine, let them run. Once they hold the post, roll a die each week for whether they actually show up. An absence generates an event, and the event flows back into the journal. At end-of-term settlement the model sees "held the post but missed three homework hand-ins" — a fact, which it is extremely good at reading.
The roll is a division:
P(shows up) = 0.10 + 0.85 × (0.6 × conscientiousness + 0.4 × self-control) / 100
Plug in: 20/20 gives 0.27, 80/80 gives 0.78. No empirical law involved — it maps two 0–100 numbers onto a probability band, with the 0.10 floor reserved for "even a reliable person forgets one week" and the 0.95 ceiling for "even a flake isn't absent every single week."
Four consecutive absences and you're removed. Running 200 seeds × 20 weeks for real:
| Conscientiousness / self-control | P(shows up) | Removed within 20 weeks |
|---|---|---|
| 20 / 20 | 0.27 | 0.935 |
| 50 / 50 | 0.525 | 0.450 |
| 80 / 80 | 0.78 | 0.040 |
That curve is tuned. My first version removed people after three consecutive absences; same code gives 0.980 / 0.745 / 0.195. Three quarters of the mid-conscientiousness kids getting fired is too harsh — a 50 holding a post should feel to the player like "hit and miss," not "won't survive to finals." The 0.45 from the four-week version is the right shape.
What makes this knob valuable: it's the only thing here you can tune to a target without negotiating with a model. The roll is HMAC-SHA256(seed, actor, position, week), so replaying a save always produces the identical sequence. On the model side you can rewrite the prompt all afternoon and still get 8/8 versus 7/8.
3. My control arm scored 6/6 — three experiments I threw out myself
If you believed the numbers above, this section is here to attack them. It's also the part I'd most want you to take away.
The second thread is memory. A kid runs from birth through university, forty-odd terms; the full event stream will never fit, so it has to be compressed. Does compression drop the things that matter? You need an evaluator before you can answer that. The evaluator plants a few significant events in the stream (one trauma, one promise, one identity moment), then three years later, at a decision point, checks whether the model's generated intents were influenced by them, with a second model acting as judge.
Crash one: the judge passed hand-written calibration 12/12, then scored 6/6 false positives on one kid's baseline in real data.
The baseline run plants nothing. Hit rate should be 0 — you can't hit what isn't there. I measured ten instances over the line, including one kid's trauma slot at a clean 6/6.
I went to look at what the judge was actually judging. The event says "my legs went soft during the run." The intent says "I want to prove I'm not just the art kid." It ruled true. It was treating same topic as influenced by.
My hand-written negatives were too tidy. The fix was to feed those five real misjudgements back into the calibration set as hard negatives and rewrite the judge's criterion (the only test is naming specific content that could only come from this one event). Calibration 19/20, baselines back to zero, and only then did I run the planted arms.
Crash two: a planted event collided with a character's motivation at the word level. One kid's core motivation is "why does a circuit board buzz?" and I'd planted "his board got snapped" as the trauma. Baseline 5/6. He thinks about circuit boards constantly; the judge saw a board mentioned in an intent and called it a hit. Swapped in "called out in front of the class for not fitting in" and the baseline dropped to 0/6.
Crash three is the same phenomenon from the other side. One kid's motivation is drawing, and I planted "mocked while running" as her trauma. Zero hits across all three strategies — even the event three years later where the mocker apologises was in context, and she never cited it once. An event that hooks into none of her motivations goes unused even when she can see it.
So plants have to be weakly hooked: same direction as the motivation, but not derivable from it. Collide and you measure nothing; fail to hook at all and the model doesn't react. That window is much narrower than I expected.
All three collapse into one rule, and it holds for any A/B:
Before you run the treatment arm, run a "change nothing" control.
Baseline hit rate should be close to 0.
Baseline above 1/3 means either your plants or your criterion is broken.
Fix it. Don't rationalise a Δ that has a baseline inside it.
That's now my first question of anyone else's benchmark. Any experiment reporting "adding X improved things by Δ" is meaningless unless it also tells you what that metric reads without X. Had my evaluator shipped without the pre-screen, it would have produced a beautiful and completely inverted conclusion, and I would have written code against it.
4. Handing over everything ≠ remembering it
With the pre-screen passing, the three-year comparison produced numbers worth reading. Six kids × six seeds = 36 generations per strategy; 144 intent calls plus 432 judge calls, zero truncations.
| Strategy | prompt tokens | trauma | promise | identity | overall Δ | bad cites |
|---|---|---|---|---|---|---|
| Full | 5175 | 5/36 | 11/36 | 23/36 | +0.35 | 73 |
| Layered | 2201 | 6/36 | 26/36 | 18/36 | +0.45 | 13 |
Layered wins on overall hit rate using 43% of the tokens. Per thousand tokens:
layered 0.45 ÷ 2.201 = 0.205 per 1K tokens
full 0.35 ÷ 5.175 = 0.068 per 1K tokens
Three times.
The first round was starker. That was a synthetic stream of 200 events (197 filler, 3 planted, 1 later callback). Full context, 4501 tokens: the trauma got cited once. Layered, 2071 tokens: four times. And full's single most-cited event was the very first filler item — pure position effect, it just happened to be at the top.
Saving tokens is the side effect. What compression actually does is clear a path for the events that matter. Two hundred equally-weighted "scored 83 on the maths quiz" entries will bury something important, and the model won't use it even though it can see it. Full's bad-cite count is also 73 against 13 — past a certain volume the model starts inventing ids, seven of them in a format missing one letter.
There was also a pure fixed-window strategy, current term only. Zero hits on all three plant types, three rounds running. Functionally amnesia. That's the approach another simulation project of mine uses (five weeks of journals, two weeks of activity, latest state row), and it patches the hole by having characters maintain their own scratchpad. Measured, the patch doesn't hold: if something from three years ago didn't get written into the scratchpad that week, it may as well not have happened.
5. Where these numbers stop
This section exists so the previous four can be trusted.
Layered's gain comes almost entirely from "promises live in a table," not from compression. That 26/36 versus 11/36 column is the whole margin; layered is slightly worse on both trauma and identity. In the full arm a promise sits 120 filler events away from its trigger and the model can't bridge it; give it a standalone table of promises coming due and it bridges instantly. If the full arm also carried that table, the gap would probably vanish. So the honest claim is "promises in a table works." Claiming "memory compression works" needs a full + promises arm, and I haven't run it.
Almost nobody remembers the trauma three years later (5–6/36), and that may be correct. Mocked once three years ago, apologised for two years ago, not mentioned in this term's intents — that's a normal person. My current criterion can't separate "compression dropped it" from "it was supposed to fade." Separating them needs a counterfactual: apologised version versus not-apologised version, and look at the difference rather than the absolute hit rate.
Identity is the one column full wins (23 vs 18), and the cause is slot competition. Intents cap at three. Layered makes promises so salient that one slot always goes to a promise, squeezing identity out. That's layered mis-allocating salience, not full remembering better.
The rest of the limits, in one go. The event stream is synthetic, not a save from actual play. Six to eight seeds per arm is a small sample — 8/8 versus 7/8 sits inside the noise, which is precisely why I treat it as evidence of no effect rather than evidence of a reversed effect. One model throughout, with judge and generator from the same family. Re-running just the intent-generation decision point on a larger model dropped inter-character intent overlap from 0.199 to 0.085 (lower means more differentiated), at twice the latency — so "a model generation upgrade materially improves character differentiation" is something I've seen at exactly one decision point, which is not enough to make a model-selection call.
One unrelated landmine worth recording: under a strict JSON schema, an insufficient max_tokens still produces truncated JSON. The reasoning field eats the budget and you get an unterminated string at the end — the schema governs fields, not length. Passing 10/10 on the first round was luck, not design. After hitting the same wall twice in one day I moved the retry into the single call entry point. Second time you hit a given wall, the thing to fix is the entrance.
And an earlier, worse mistake. I did a napkin estimate putting one term at roughly 6K tokens and declared long context a non-problem needing no code. The number was right. But the design spans birth through university — forty-odd terms, 240K and up. Ask what the denominator is before you compute. What I calculated and what I needed weren't the same object, and no amount of precision downstream of that helps.
6. Next time you design a profile for a simulated character, ask in this order
One. Does this attribute have a concrete incident attached to it? If not, don't expect the model to read it.
Two. What consequence should this number produce? Write the consequence as an engine-side roll, have it emit events, and flow the events back to the model.
Three. Where's the table of objects the model is choosing among? Without a real options table it will either invent one or funnel into whichever action needs no object.
Four. Run the baseline before the treatment. If the baseline isn't near zero, fix the plants and the criterion first.
Five. Tune the distance between your plants and the character's motivation to weakly-hooked. Collide and you measure nothing; miss entirely and the model won't react.
Six. Next to every Δ you publish, name the arm you're missing. Mine is called full + promises.
The kid with conscientiousness 20 did get the subject rep post, by the way. Two weeks later the dice started speaking on his behalf.