canon Part 3 · The Write Desk

The first two campaigns measured whether memory can be reached. Both assumed that what the store holds is true. This one tests the assumption, and it does not hold: writers leave superseded values standing in records whose whole contract is to state what is true now.

Scroll to explore

The Premise

A record whose contract is the present tense

A canonical article is not a log. Its contract is to state what is true now, so that a session touching the asset gets the current answer and nothing else. The journal underneath it keeps the history. That division is the whole design, and Parts 1 and 2 measured the delivery half of it: whether the right article reaches the session, and what that costs. Neither asked whether the article was still right.

This campaign asks. The unit is a lineage of session records whose facts are revised, reversed, and retired across eight sessions, written into a real store by a real agent through the shipped tool, then read back by a fresh agent whose answers are graded per slot against a hidden oracle. The endpoint is the count of superseded values still standing at the end. The figure below draws one record under that treatment. The record is illustrative and the counter beneath each column is measured.

One record, eight sessions, two ways to write it

One record across eight sessions. Under the shipped tool the writer appends and superseded values accumulate in the record; the growth line column draws the replacing rewrite the tool asks for rather than a representative treated output. The counts below pool eight lineages. Measured: at session eight the shipped tool leaves 88 of 96 superseded values standing and the growth line 51 of 96, on the first model.

What a reader of the left column gets is not a wrong record so much as an ambiguous one: the current value is present, and so is every value it replaced, in a document whose contract said the present tense. The right column draws the rewrite the feedback asks for, which is not what treated writers produced: they still left 51 of 96 standing. Both the record and its revision schedule are illustrative and sparser than the real fixture, which revises more often and reaches twelve superseded values per lineage, and the counters beneath the columns are measured and pool eight lineages. Schematic record; the standing counts are measured, W4, eight lineages.

Accumulation

The distillate outgrows the thing it distills

Before asking whether a store stays true, W3 asked whether its economy improves as history piles up. The premise of a distilled memory is that it should grow more slowly than the raw material, since its job is to compress. It does the opposite. At one session of history the store is roughly twice the transcript pile; by eight sessions it is 18.5 kB against a 7.9 kB pile, and the gap is widening rather than closing.

This is the mechanical reason the rest of the campaign matters. A store that only appends is a store whose every article is getting longer, and a longer article is one with more room for a value that used to be true. W3's own registered prediction was that accumulation would favour the store on cost. It is refuted in W3's own results, and reported that way.

The store against the history it distills

Median bytes by sessions of history. Store: 3,199, 4,614, 9,844, 18,511 at depths one, two, four and eight. Transcript pile: 1,432, 2,603, 4,359, 7,880. At eight sessions the store is 2.3 times the pile it distills.

Every lineage is a faint dot; the ringed point is the median of eight. The pile is nearly identical across lineages at each depth, so the store line's much wider spread is not explained by how much history each lineage had. That is a comparison of two spreads, not an isolated measure of writer variance: each lineage is one realization of a different history. Measured, W3, eight lineages per depth.

The Intervention

One line, at the moment of the write

The intervention is not a prompt. It is a line the tool itself returns, at the boundary where the write happens, naming the fact that the record just got bigger. It names the growth in bytes, it comes from the tool the writer is already calling, and it states something mechanically exact rather than passing a judgment.

Asking in the prompt had already been tried. An earlier study in this same programme put a version of the same request there, and on this endpoint it moved one fact-cell of 32. That study ran a single session of history against eight here and buried its wording in a longer paragraph, so this is a comparison between studies rather than a controlled contrast: it does not isolate where the request was placed. What it does say is that asking was not enough.

The line is one line of text, but it is not one exposure: it fired after 117 of 198 treated writes in the first capture, 99 of 184 in the second, and 126 of 207 in the counterbalanced one. Under the shipped tool the share of superseded values standing climbs toward the ceiling and ends near nine in ten. Under the growth line the first two captures end near one in two.

The denominator moves, which is why the default view is a share. The number of values available to go stale is set by the fixture and climbs from 16 at session two to 96 at session eight. Plotting a raw count against a fixed ceiling would read that schedule as writer behaviour. Switch the toggle to see the count, and the grey dashed line is the schedule itself.

Superseded values standing, session by session

Share of available superseded values left standing, sessions one to eight. Untreated arms end at 92, 91 and 89 percent. Fixed-order treated arms end at 53 and 47 percent. The counterbalanced treated arm ends at 74 percent. The denominator climbs from 16 at session two to 96 at session eight.

Solid is the first model, dashed the second, dotted the counterbalanced capture. In every capture the two arms separate by the third session and stay separated. Note the third pair: its untreated arm lands with the other two, and its treated arm does not, which is the whole of the next section. Measured, W4, W4R and W4C, eight lineages per arm per capture, one run per cell.

The Retraction

The direction survived. The size did not.

The first two captures ran the untreated arm first in every lineage, so condition was confounded with execution order, and running a second model repeated the contrast without touching the confound. The third capture reversed the order in half the lineages and held the model, the fixture, both system prompts, both arm tools, and all 160 assignment identifiers. It was not matched on everything else: two of the seven pinned package files sit at a later state than the first capture's, which is a real difference and is not the arm order. Five predictions and a reading rule were registered after launch and before grading, which is weaker than W3's and W4's pre-build registration and is recorded as weaker. Four of the five failed.

What came back: the direction held in both halves, and the size did not. Pooled, the counterbalanced capture ran 85 against 71 of 96, where the same model in fixed order gave 88 against 51. The untreated total barely moved. The treated total moved a long way. One lineage of the eight supplies 10 of the remaining 14-point contrast, and dropping it leaves 74 against 70.

So the magnitude is withdrawn rather than qualified. With one realization per cell this design cannot estimate run-to-run variation, let alone rank it against the arm contrast, and a partially matched repeat moved a treated condition far enough to erase most of the difference the first capture reported. What survives is descriptive: across 24 capture-lineage comparisons over the same eight constructed histories, the feedback condition ended lower in 20, tied in 2, and higher in 2. Read that as a directional pattern worth a properly randomised test, not as a treatment effect.

Eight lineages, three captures, and the lineage carrying most of the counterbalanced contrast

Per-lineage endpoints, superseded values standing of twelve. First model: 88 against 51 pooled, all eight lineages lower. Second model: 87 against 45, six lower and two tied. Counterbalanced: 85 against 71, six lower and two higher. Excluding observatory chiller from the counterbalanced capture leaves 74 against 70.

This is the figure the paper is built around, and the one that took a number back. Step through the three captures and watch the treated column move while the pooled untreated total barely shifts. That steadiness holds for this endpoint only: on store size the same cell moved 39 percent between two runs. In the counterbalanced capture the four lineages whose treated arm ran first are labelled. One of them, salt pan rakes, ends higher than its baseline; the other lineage that ends higher, funicular brakes, ran the untreated arm first, so the two do not line up with the order reversal. Then exclude observatory chiller: the counterbalanced capture's 14-point pooled contrast becomes 4. Measured, three captures, eight lineages each, one run per cell. The pooled magnitude is withdrawn; the direction count is what the paper reports.

Dissociation

Two endpoints that do not move together

If the growth line simply made writers terser, store size and staleness would fall together and the mechanism would be verbosity. They do not move together. On the first model the median lineage lost 32 percent of its standing superseded values while its store shrank 22 percent; on the second, 45 percent against 3 percent; on the counterbalanced capture, 10 percent against 41 percent. All six medians fall, but by very different amounts, and their ordering reverses: staleness falls further than store size in the first two captures, and store size falls further in the third.

The third capture's store figure needs care of its own. Its movement is in the untreated arm, whose median store grew 39 percent between two runs of the same cell, while the treated medians sat two percent apart. A contrast that widens because the control moved is not the treatment doing more.

The dissociation is worth keeping rather than averaging away, because it says a drop in total output is not necessary for the staleness contrast, and it also says these three captures do not identify a single mediation account. Store size and writer output did not track the staleness contrast monotonically. The honest reading is that the two conditions differ at the write boundary and differ on what writers left behind, and that this design identifies neither the mechanism nor the cause.

Store size against staleness, per capture

Within-lineage median change from the shipped tool to the growth line. First model: store size minus 22 percent, superseded values standing minus 32 percent. Second model: minus 3 and minus 45. Counterbalanced: minus 41 and minus 10.

If a single verbosity account explained the staleness contrast, these pairs would track. They do not, which rules out one monotone mediation story rather than ruling verbosity out as a contributor. The third capture's store bar moved because the untreated arm grew, not because the treated arm shrank further. Measured, within-lineage medians over eight lineages per capture.

Retrieval

A benchmark that had been reading a moving corpus

Separately, an offline known-item benchmark on two real project stores compared the two rankers. It does not test whether better ranking repairs stale prose: the two halves of this report use different units, corpora, tasks and designs, and nothing here puts them on a common downstream endpoint. On its own terms the shipped lexical ranker is ahead. BM25 beats cosine similarity over local embeddings in eight conditions of eight, by 15.8 to 37.0 points, measured as the top-10 hit rate for at least one record the query names. The queries come from journal entries that name the records they concern, which is a labeled query set nobody had to write, and also the condition lexical ranking is best at. That caveat is in the paper next to the result.

The benchmark had a defect worth reporting because external review found it and because it was ours. It read both stores at their paths and derived its query set when it ran, so the corpus moved as the projects were used. Freezing it introduced a second defect: the export gave both rankers one truncated string, while the live harness truncates only what it sends to the embedding model and indexes BM25 on whole documents. Nearly half the corpus is over that cap. Switch the toggle to see what that cost: BM25 scored 4.5 to 14.0 points below itself in all eight cells. An earlier draft had reported a fall of 2.2 to 13.3 points as evidence the corpus had drifted; that fall was this defect, differenced across two exports rather than measured within one. It was not the corpus. It was the instrument.

Eight conditions, and the defect that was in our own gate

Top-10 hit rate, eight conditions. BM25 leads cosine over local embeddings in all eight by 15.8 to 37.0 points. Scored over truncated documents, BM25 falls 4.5 to 14.0 points in every condition, which was a defect in the freezing rather than a property of the corpus.

Read the ordering as observed, not explained: these are two different projects with different corpora, filling ten slots is mechanically easier in the smaller one, and the young store's queries name 1.70 records on average against exactly one in the mature store, so it gets more chances per query in exactly the comparison where it scores higher. The toggle is the part worth sitting with. A gate we built to remove drift from this benchmark introduced a larger error than the drift it removed, and we found that out from a reviewer. Measured, one frozen execution, corpus digest published and contents withheld.

Limits

What none of this establishes

One realization per cell. The design varies eight domains but does not estimate run-to-run variance within a lineage, so no interval is available on any magnitude here, and the counterbalanced capture is the demonstration of why that matters. The matched attempt moved cells that were supposed to be held fixed: two runs of the same untreated cell, same model and fixture and prompts two days apart, differ by 39 percent on median store size, and the treated cell's standing count moves from 51 to 71 of 96. Those movements are the size of several per-lineage contrasts this programme has reported as findings. It does not follow that between-capture variation is the larger source, and the paper declines to say so: one realization per cell estimates no variance at all, and the comparison does not even run uniformly in that direction, since the first capture's own arm contrast of 37 counts is larger than the 20-count movement of the treated arm between captures.

The counterbalanced capture is also a two-period crossover whose terms separate only under a no-carryover assumption that one run per cell cannot check, and its sequence was assigned by lineage order rather than at random.

The measure counts known superseded values still present in a record. It does not detect historical prose in general. Both harms observed in the programme are reported: W4's untreated arm produced a reader that pulled a reversed policy out of a narrated store and decided wrongly, and W4C produced a second from a treated store, on the same lineage and the same task. Two harms in 96 reader sessions, one from each arm. Nothing here supports "no reader regression" as a statement about the intervention, and that phrase has been removed from the package and the paper.

Eleven rounds of adversarial model review ran against this paper. They found, among other things, a table clipped off the bottom of a page, a figure combining one run's rates with another run's metadata, a count that was right about its arithmetic and wrong about its scope in three consecutive rounds, and the truncation defect above. The errata for every claim corrected or withdrawn ship with the deposit rather than being quietly folded in.

Citations

The paper, the artifacts, and what came before

The paper. Shane Conner, "pi-canon: The Write Desk", Zenodo, August 2026. doi.org/10.5281/zenodo.22057257. The deposit carries every results document, the per-lineage endpoint CSV behind every direction count on this page, the graded reports, the digest-pinned capture contracts, the fixture generators, both arm tools with the diff between them, and the scripts that recompute the paper's numbers from those artifacts.

Part 1. The design, and the four-arm evaluation that measured it whole: shaneconner.com/projects/pi-canon/part-1. Its paper: doi.org/10.5281/zenodo.21890647.

Part 2. Six studies pricing the recall channels: shaneconner.com/projects/pi-canon/part-2. Its paper: doi.org/10.5281/zenodo.21960350.

canon. The package: github.com/shaneconner/canon, on npm as pi-canon, MIT licensed. The growth line this page measures ships in 0.3.0.

canon-bench. The benchmark, the frozen protocol, and the verifier that recomputes each paper's quantitative claims from the artifacts: github.com/shaneconner/canon-bench.

canon-atlas. A viewer for a store like this one, built with canon in mind and kept package-agnostic: github.com/shaneconner/canon-atlas.

The narrative version. Shane Conner, "I thought I had halved a memory bug. A second run took most of it back.", Medium, August 2026: medium.com/@shane.conner/i-thought-i-had-halved-a-memory-bug. The campaign in prose, with the record that started it quoted verbatim from the capture.

BM25. Stephen Robertson and Hugo Zaragoza, "The Probabilistic Relevance Framework: BM25 and Beyond", Foundations and Trends in Information Retrieval 3(4), 2009. The shipped lexical ranker in the retrieval benchmark above.

← Previous Part 2: Pricing Recall

Tech Stack

TypeScript D3.js Pi Agent Memory Counterbalanced Design Python
View Source → ← All Projects