Back to blog

2026.10.08

6,950 Saved Lessons, 93.5% Never Read Once: Should My AI Assistant's Memory Get Better Retrieval, or Fewer Writes?

I built a lesson bank for my AI assistant. Every night a model distills 'what to do next time' from the day's conversations, and in four months it piled up 6,950 entries. I went through 962 retrieval calls: only 451 entries ever came back. The top 1% of entries took 47% of all hits. The confidence score the model gave each entry correlates with whether it ever gets used at r = -0.002. I was about to add the usage-feedback loop from the paper I'd copied. One division made me drop the plan. This post gives three formulas you can run on your own system: the read/write ratio, the number of searches you'd actually need, and an upper bound on hit rate when part of your logs are gone.

AI Agent记忆系统方法论评测

My AI assistant has a lesson bank. In four months it collected 6,950 entries.

6,499 of them have never been read back. Not once since the night they were written.

That's 93.5%.

And that's the generous version. As you'll see below, 29% of the retrieval logs are already gone, so the real number can only look worse.

The embarrassing part: right before I ran this count, I was about to add a feature. Every time an entry got used, bump a counter. Every time it turned out wrong, flag it. That's what the paper does, and I felt like I'd only built half of it. Once the numbers came in, I killed the whole plan. The reason is one division, and it's in section 3.

If you just want the yardstick, read sections 3 and 4. If you want to know where my own stats are shaky, section 6.

1. How the bank got filled

Every night, a large model reads through the day's conversations and pulls out two kinds of things. A strategy: what to do next time a similar situation shows up. A lesson: a "don't do that again" taken from something that went wrong. Each one is a sentence or two. The model also gives each entry a confidence score from 0 to 1, meaning roughly "how sure and how general is this." Then it goes into a vector store, mixed in with all my other notes.

The idea comes from Google's ReasoningBank paper (arXiv 2509.25140). The full loop in the paper has three steps: retrieve relevant experience, use it to guide the current task, then update the bank based on how the task went. I built only the very first slice of that: storage.

Once an entry is stored, there's exactly one way it can ever reach the assistant again. The assistant has to run a search on its own, and that search has to happen to surface it.

The extraction prompt is actually pretty restrained. It says "quality over quantity" in so many words, and it says "if there's nothing worth keeping, return an empty array; that is normal and encouraged." Even so, across 103 nights, the median night still produced 55 entries. Models have a very hard time not writing something down.

2. How I counted "was read"

The vector store doesn't log who looked at what. But every search the assistant runs leaves its raw results sitting in the conversation database. So the method is crude: pull every search call's output, parse it line by line, and match the first 120 characters of each result against the bank.

  • 962 search calls, from late May to October 7
  • 277 of them had lost their results. When a conversation gets long, old tool outputs get compressed down to a one-line summary and the original text is thrown away
  • The remaining 685 parsed cleanly and returned 1,044 results from this bank. All 1,044 matched exactly one entry

Result:

Total entries          6950
Returned at least once  451   (6.5%)
Never returned         6499   (93.5%)

My first reaction was to push back. Fresh entries haven't had time to get searched; lumping them in is unfair. Fine, split by age:

Time since written Entries Share ever read
Under 30 days 1360 1.3%
30–60 days 2257 1.9%
60–90 days 1978 9.2%
90+ days 1355 15.3%

So yes, it was unfair, but it doesn't change the conclusion. Entries older than three months are the best-case bucket, and 85% of those have still never been touched.

3. One division: reads can't keep up with writes

The loop I wanted to add was "count each time it's retrieved, flag it when it gets contradicted." A counter is only useful if it has something to count.

So how many entries does this bank get read back per month?

Max share of new entries that can be read ≤ entries read out this month ÷ entries written this month

September's numbers:

Written              1450 entries
Read-out slots        165 (the 123 parseable searches returned 165 results from this bank)
Ceiling              165 ÷ 1450 = 11.4%
Actual, deduplicated  73 distinct entries
                     73 ÷ 1450 = 5.0%

The 11.4% assumes every result read out is a different entry, and all of them are from this month. In practice most of what got read was old familiar faces. Deduplicated, it's 73 entries.

That's just a division. It has nothing to do with how good the retrieval algorithm is. Even a perfect feedback loop could only put a tick next to 73 entries that month. The other 95% of counters would stay at zero forever. Ranking or pruning on a field that's zero for nearly everything is picking winners out of noise.

Month by month, the gap is getting wider:

Month Written Read-out slots Distinct entries read
July 1693 543 225
August 2378 258 159
September 1450 165 73

Writes have been running an order of magnitude ahead of reads the whole time, and reads keep sliding. It's basically a warehouse with a loading dock and no exit.

Okay, so just search more?

The obvious objection: the real problem is too few searches. That can be computed too:

To have every new entry this month read at least once,
searches needed ≥ entries written per month ÷ average results from this bank per search

In September each search returned 1.34 results from this bank on average (165 ÷ 123):

1450 ÷ 1.34 ≈ 1082 searches/month, about 36 a day
Actual: 192 in September, about 6 a day

That's 5.6× more searching. And that's the fantasy case where every result is an entry nobody has seen yet. In reality the top 1% of entries soaks up 47% of hits (section 5), so dredging up the long tail through sheer search volume would take several times more again.

One more detail. This bank makes up 42.9% of everything in the vector store, but only 26.3% of the slots in search results. The ranker already prefers my other notes. Search more, and the extra results mostly come from those other notes.

So the bottleneck is on the write side. I'm not building the feedback loop. The next step is to tighten extraction first.

4. Two numbers that look useful and carry no information

The confidence score

Every entry got self-graded by the model at write time. The distribution: 3,113 entries at 0.8, 1,605 at 0.9, and 1,596 at a perfect 1.0. The model is pretty confident about its own writing.

So do high-confidence entries get used more? The correlation between confidence and "was ever read":

r = -0.002

Bucketed, it's easier to see:

Confidence Entries Share ever read
< 0.7 268 1.5%
0.7–0.8 860 5.9%
0.8–0.9 2664 7.7%
≥ 0.9 3158 6.0%

Restricting to entries at least 60 days old, to take the age bias out: the ≥0.9 bucket sits at 9.1%, actually lower than the 0.8–0.9 bucket at 14.9%. Correlation -0.088.

Put plainly, that field is decoration. At the moment a model writes a lesson down, it has no idea whether it will ever matter.

The yardstick: to find out whether a score field carries information, don't stare at how pretty its distribution looks. Bucket by score and check whether the downstream outcome rate moves monotonically with it. If it doesn't, it's decoration.

"Barely any near-duplicates"

I also checked a suspicion: is the bank stuffed with the same idea in different words? Using cosine similarity ≥ 0.9 as the duplicate threshold, I found only 29 pairs. Deduplication would keep 99.8% of entries.

At the time I wrote that into the conclusion: duplication isn't the problem.

Then I looked at the most-hit leaderboard. Three of the top ten describe the same thing: an open-source LLM gateway that estimates token counts before routing using a generic tokenizer, which badly overcounts Chinese text. Three phrasings, filed under three different categories, hit 18, 15, and 13 times. A 0.9 threshold only catches near-verbatim copies. Reword it and it slips through.

So that 29 only tells you verbatim duplicates are rare. How many entries say the same thing in different words, I never measured.

5. What the 6.5% that does get read looks like

The concentration is extreme:

Top 1% of entries (69) account for 47.1% of all hits
Of the 451 entries ever read, 244 were read exactly once
Only 110 were pulled up by 2 or more different conversations

The most-read entries look like this:

  • "Pick a test machine that reproduces the actual scenario you're making a claim about, not the convenient one": 18 hits
  • "To find an LLM endpoint's real context limit, run a needle-in-a-haystack test with a prompt close to the full window; a 200 from the health endpoint doesn't count": 17 hits
  • "When investigating how a self-hosted LLM endpoint really behaves, hit it directly with controlled comparisons: no parameter, then each explicit setting": 17 hits

What they share: they cut across projects, they work as a literal checklist step, and you can tell in one sentence whether you followed them. They read like a line on a preflight checklist, not like a war story.

Flip through the other 93.5% and a lot of it is "here's how parameter X got tuned that one time." True for that one time only.

That's where my plan comes from. The few dozen entries that actually get reused shouldn't sit in a vector store waiting for a lucky search. Promote them straight into an always-loaded checklist. Archive the rest by age.

6. Where this analysis is shaky

  • "Returned" isn't "used." An entry showing up in search results doesn't mean the assistant actually read it closely. I'm counting an upper bound. Real adoption can only be lower

  • 29% of the retrieval logs are gone. 277 searches had their results compressed away, so the hit count is a lower bound. You can compute a worst case:

    Hit-rate upper bound = (entries known to be read + lost searches × avg results per search) ÷ total entries
                         = (451 + 277 × 1.52) ÷ 6950
                         ≈ 12.5%
    

    That assumes every single lost result was an entry nobody had ever read before, which can't realistically be true. Even on that assumption, more than 87% have never been read

  • Entries were matched on the first 120 characters of text. All 1,044 matched uniquely, no ambiguity, but two entries with identical first 120 characters would be counted as one

  • One assistant, one person's habits. I don't search much; someone else's assistant might search a lot more. The formulas in section 3 work with your own numbers plugged in, but don't copy my values

  • The "search more" part is a projection. I never actually cranked retrieval up to 36 a day to test it

  • Duplication isn't properly measured. See section 4: a 0.9 threshold only catches verbatim copies

7. What to ask when someone says "my memory system has N entries"

You see this kind of number all the time: memory bank has X entries, extracted Y lessons, knowledge graph has Z nodes. All those numbers prove is that the write side is busy.

Three questions are enough:

  1. How many distinct entries get read back per month, and what's that as a fraction of monthly writes?
  2. How concentrated are the reads? What share do the top 1% take?
  3. Does the system's own score field (confidence, importance, whatever), once bucketed, have any relationship to whether an entry gets read?

My own answers: 5%, 47%, and no.

8. Checklist

Before you add any auto-writing memory layer to an AI assistant:

  1. Count the read side first: how many searches per month, and how many results from this layer each one returns on average
  2. Compute "monthly writes ÷ results per search" to get the searches you'd need, and compare to reality. If you're off by an order of magnitude, don't add more writes yet
  3. Don't let the model grade its own entries with a confidence score unless you plan to go back and check it against outcomes
  4. Once a month, bucket by entry age and compute read rates per bucket. Don't hide behind a single total
  5. Promote the few dozen entries at the top of the hit leaderboard into an always-loaded checklist
  6. Anything 90+ days old and never read goes to a cold archive. Recoverable. Don't just delete it
  7. Don't trust a single high similarity threshold for dedup. Look at the top of the hit leaderboard; reworded duplicates jump right out