Back to blog

2026.10.06

The Circuit Breaker Tripped 80 Times and the Board Still Said "Running": Who Was My Self-Healing Script Actually Saving?

Four AI workers started 1,652 times in one night, each living about 60 seconds, while the task board showed "running" the whole time. My diagnosis that morning: the failure counter was stuck between 10 and 12, never reaching the limit of 20, so the breaker never fired. Two days later I read the raw event table. The breaker had tripped 80 times, and every time my own revive script closed it again within about a minute and a half. This post covers how to work out from one counter snapshot how many times it was reset, why this kind of loop phase-locks to the cron interval, and why a revive script that silently exits looks exactly like a healthy one.

AI Agent可靠性方法论监控

Four AI workers, one night, 1,652 process launches in total. Each one lived for about 60 seconds, printed the same error, and exited.

For seven and a half hours, the task board showed all four cards as "running."

The next morning I tracked down the cause and wrote this conclusion into my own runbook: the failure counter was stuck at 10 to 12, it never reached the limit of 20, so the circuit breaker never tripped.

That conclusion was wrong. When I went back to the raw event table today, the breaker had actually tripped 80 times, 20 per card. Every single time, it was closed again about a minute and a half later. The thing doing it was a "revive script" I wrote myself.

If you just want the diagnostic, it's in section 2, and it's one subtraction. If you want to know why a script that silently exits is harder to catch than one that crashes, read section 4.

1. What happened that night

Quick context. I run a small team of AI agents that take turns optimizing an inference engine. Each agent is a card on a kanban board. A dispatcher scans the board every minute, and whenever a card is "ready" it launches a worker process to pick it up.

When I created the cards, I attached a skill to all four of them. A skill here is basically an operating manual the agent loads at startup. The problem: that skill only existed in my own profile. The workers run under separate profiles, and they couldn't find it. So every worker hit Unknown skill the moment it started, and exited.

The dispatcher's design is sound. After enough consecutive failures, stop retrying, mark the card "blocked," and wait for a human. That's a circuit breaker, same idea as the one in your fuse box: when something shorts, cut the power instead of letting it keep burning. These cards had a limit of 20.

By design, all four cards should have tripped within about 20 minutes, wasting 80 launches total, and then waited for me to fix them in the morning.

What I actually wasted was 1,652 launches.

2. The answer was in the first monitoring report: 110 = 5 × 20 + 10

At 3:37 a.m., a read-only monitor I'd left running filed its first report. It contained two numbers:

  • Failure counter on the board: 10
  • Crash count per card in the event log: 110

I read that as "the counter isn't going up," and so did the report itself. Looking back, those two numbers side by side were already the evidence. The limit is 20. How does a card crash 110 times without ever tripping?

Only if somebody reset the counter. How many times is one subtraction away:

times reset = (cumulative failures − current counter) ÷ breaker limit
            = (110 − 10) ÷ 20
            = 5

I checked the event table. Before 3:37, every card had tripped exactly 5 times: 01:51, 02:17, 02:40, 03:02, 03:24. Not one off.

Put simply, the counter snapshot was lying. The counter was cycling from 0 to 19 over and over, and whenever you looked, you saw some random point on that loop. Seeing 10 doesn't mean it's stuck at 10. It means you showed up halfway around.

So who was resetting it? The revive script. For long-running jobs like this I always set up a small script that runs every ten minutes or so, because workers tend to quit on their own before the deadline: they declare themselves done, or block themselves waiting for a review that isn't coming. The script's logic is dead simple. Before the deadline, if a card is "blocked," zero its failure counter, unblock it, and send it back to work.

It never asks why the card is blocked. A tripped breaker is also "blocked," so it gets rescued too.

In the event table, each of the 80 trips is followed by an unblock. The gap ranges from 54 seconds to 347 seconds, averaging 97.

I wrote a thing specifically to stop workers from slacking off, and it disabled the thing specifically meant to stop runaway loops. Each mechanism was fine on its own. Put them together and you've built a perpetual motion machine.

3. The loop locked onto the cron interval

One more detail I liked. The time from one trip to the next ranged from 1,267 to 1,567 seconds, averaging 1,333. You can predict that number ahead of time:

time to burn through the limit = breaker limit × dispatch interval = 20 × 60.3 s ≈ 1,206 s
revive script interval         = 660 s (measured over 86 runs: 655–664 s)
loop period ≈ ⌈1,206 ÷ 660⌉ × 660 = 2 × 660 = 1,320 s
measured: 1,333 s, about 1% off

Burning through 20 failures takes a little over 20 minutes, and the revive script shows up every 11 minutes. After a card gets rescued, the script's next visit finds it still burning. The visit after that finds it freshly tripped, so it gets rescued again. All four cards' unblock timestamps look like 01:57:38, 02:19:38, 02:41:42, 03:03:42: one tick every 22 minutes, like clockwork.

That period also sets your odds of catching the problem by glancing at the board:

P(you see "blocked" on a spot check) ≈ mean reset delay ÷ loop period = 97 ÷ 1,333 ≈ 7%

An independent number lines up with it. Over those 7.6 hours, each card had a live worker process 90.7% of the time (413 launches × ~60 s ÷ 27,446 s). The remaining 9-ish percent is blocked time plus gaps between processes.

In other words, any time you happen to look at the board, more than nine times out of ten you see "running" and a counter somewhere between 0 and 19. Everything looks fine, just a bit slow.

Without the revive script, the night would have cost 4 × 20 = 80 launches. With it, 1,652. Twenty times as much.

4. Flip it around: when the revive script itself dies, it's just as quiet

A day later, the same revive script taught me the opposite lesson.

Before round 13 kicked off, I copied the previous round's script and only swapped the round number. The deadline was a hard-coded number on its own line. I forgot to change it. It was still the previous round's deadline, which had already passed at 4:58 that morning. The very first thing the script does is "if past the deadline, exit." So every time it ran, it quietly called exit(0).

The script follows a convention: "print only when you did something; empty output means all good." So it ran every ten minutes or so, printed nothing every time, and looked perfectly healthy.

Meanwhile, six of the seven workers stopped between 10:48 and 12:25, and not one of them got pulled back. The last evaluation ran at 12:47. An 8-hour round, and by my estimate about two and a half hours of actual work. Final gain: 0.16%.

You can state this failure mode with a crude ratio:

information in empty output = P(empty | script healthy, nothing to do) ÷ P(empty | script dead)
                            = 1 ÷ 1
                            = 1

A likelihood ratio of 1 means that seeing empty output should not move your belief that the script is alive, not even slightly. I had made silence the signal for "healthy," and a dead script is silent too.

Section 2 was a self-healing script that rescued too much. Section 4 was a self-healing script that died. On the board, the two look identical: every card shows whatever status you'd expect, and nothing raises an alarm.

5. When someone says "we have auto-retry plus a circuit breaker, it's solid," ask two questions

Since these two incidents, I ask two things about any system that claims to recover from failures automatically:

First: after the breaker trips, who closes it again? If the answer is another piece of automation (a health check, a revive script, an orchestrator's restart policy), keep going: does that automation tell transient failures apart from deterministic ones? If the same error has repeated 20 times, the 21st attempt is very likely to fail the same way. If it can't tell the difference, the breaker is just a timer for the infinite loop.

Second: which number are you looking at when you say "solid"? If it's a point-in-time snapshot ("everything is running," "all failure counters are under the threshold"), it tells you nothing. Look at cumulative events instead: how many trips and how many resets over a window. Use the subtraction from section 2, (cumulative failures − current counter) ÷ limit. If it comes out above 0, something behind the scenes is zeroing the counter.

That night at 3 a.m. I had both numbers in front of me. I just never did the subtraction.

6. Where this post's evidence is soft

  • The revive script that ran that night was version 1. I upgraded it to version 2 the same afternoon, and the original file was overwritten. My claim that "it zeroed the counter whenever it saw a blocked card" rests on the event table (all 80 trips followed by an unblock, every unblock landing on a script run time) plus the fact that the previous round's script of the same kind contains exactly that logic. The evidence is strong, but I don't have the exact bytes of that night's version.
  • The wrong conclusion I mention in the opening sat in my runbook for two days. Its defensive advice was still right (don't attach skills to cards; within 3 minutes of creating cards, confirm a worker has actually stayed alive). Only the explanation of the mechanism was wrong. I've fixed it now.
  • The period formula in section 3 only locks to an integer multiple when the burn-through time and the script interval are in the same ballpark. If burn-through is much shorter than the script interval, the script rescues the card on every run, and the period is just the script interval. I have data from this one night only.
  • "About two and a half hours of actual work" in section 4 is an estimate from worker stop times and the last evaluation timestamp. I didn't sum each worker's working time exactly.
  • This whole post covers one incident, four cards, one root cause. So 7% and 90.7% describe that night, not a general law.

7. Add these to your self-healing script

  1. Before unblocking, read the last few failure messages. If the same error keeps repeating and no process lives past a minute, don't rescue it. Raise an alert instead.
  2. Tell "the worker stopped itself" apart from "the breaker tripped." Don't rescue the second kind by default.
  3. Don't have monitoring read the failure counter's current value. Count events instead: trips and unblocks over a window.
  4. Add a self-check: cumulative failures > breaker limit, yet the card isn't blocked → something is resetting it. Report that immediately.
  5. Make the self-healing script print a heartbeat line on every run, even when there's nothing to do. Have monitoring check the time of the last heartbeat. Don't treat empty output as healthy.
  6. Read parameters like the deadline from this round's config file, never from a copy of last round's script. After setting up the script, run it once by hand and print the deadline it read, so you can eyeball it.
  7. Within 3 minutes of creating the cards, confirm every worker has at least one process that stayed alive past 2 minutes. Until then, "running" on the board means nothing.