I had a suspicion that my AI partner lost the thread every time a conversation was resumed after a few days' pause. So I put it to work testing that suspicion — on 263 of our own working sessions from a single month. The suspicion didn't hold.
Instead, the review found a hole somewhere I wasn't looking at all: my own ideas. Roughly half the notions I had parked along the way never surfaced again.
Anecdotes didn't count as patterns
The setup itself was plainly boring: a month of sessions was distilled and trawled by a pack of extraction agents hunting for corrections, repeated requests, lost ideas, and resumptions. The rules were strict. A single episode didn't count as a pattern; a claim had to show up in at least two separate sessions before it was allowed to stand.
Those resumptions I was worried about? 37 clean restarts, one piece of work accidentally redone, zero outright failures. My sense that "something always got lost" turned out to be exactly that: a sense.
The ideas were another story.
Of roughly 95 traced ideas, about 30 percent got picked up again in the same session and 20 percent in a later one. The rest — about half — vanished without a trace. The common denominator wasn't the quality of the ideas but their address. Ideas that lived only in the chat died in the chat. Ideas that got one line in a project document survived.
263 sessions · ~95 traced ideas · ~50% never seen again · 37 clean resumes · 0 failures
So today's rule is banal: a mid-task idea gets one line in the project's status document — date, label, one sentence — and then back to work. At every resumption, the newest five get surfaced. There is no more ceremony to it than that.
Willpower doesn't scale
The review also measured another hobby-horse of mine: that the AI must restate the task in its own words before starting, so misunderstandings get caught while they're cheap. The rule was followed roughly 107 times and broken 18. The interesting part is where the breaks fell: all 18 came before the rule was built into the tooling as a fixed first step. After that: zero.
That lesson reaches beyond AI. Willpower doesn't scale. Tooling does.
So my self-image took a few scratches: the thing I feared worked fine, and the thing I blindly trusted was leaking. But that is the whole point of looking. A gut feeling is a hypothesis, not a report — and hypotheses get tested before you build your working day on top of them.