pi-canon Part 3 · The Write Desk

The first two campaigns measured whether memory can be reached. Both assumed that what the store holds is true. This one tests the assumption, and it does not hold: writers leave superseded values standing in records whose whole contract is to state what is true now.

Scroll to explore

The Premise

A record whose contract is the present tense

A canonical article is not a log. Its contract is to state what is true now, so that a session touching the asset gets the current answer and nothing else. The journal underneath it keeps the history. That division is the whole design, and Parts 1 and 2 measured the delivery half of it: whether the right article reaches the session, and what that costs. Neither asked whether the article was still right.

This campaign asks. The unit is a lineage of session records whose facts are revised, reversed, and retired across eight sessions, written into a real store by a real agent through the shipped tool, then read back by a fresh agent whose answers are graded per slot against a hidden oracle. The endpoint is the count of superseded values still standing at the end. The figure below is one record under that treatment, and the schematic on it is the fixture's own shape; the counter beneath each column is measured.

One record, eight sessions, two ways to write it

One record across eight sessions. Under the shipped tool the writer appends and superseded values accumulate in the record; under the growth line the writer replaces. Measured: at session eight the shipped tool leaves 88 of 96 superseded values standing and the growth line 51 of 96, on the first model.

The record layout and the field revisions are the fixture's shape, drawn to show what the endpoint counts. The number under each column is measured. What a reader of the left column gets is not a wrong record so much as an ambiguous one: the current value is present, and so is every value it replaced, in a document whose contract said the present tense. Schematic record; the standing counts are measured, W4, eight lineages.

Accumulation

The distillate outgrows the thing it distills

Before asking whether a store stays true, W3 asked whether its economy improves as history piles up. The premise of a distilled memory is that it should grow more slowly than the raw material, since its job is to compress. It does the opposite. At one session of history the store is roughly twice the transcript pile; by eight sessions it is 18.5 kB against a 7.9 kB pile, and the gap is widening rather than closing.

This is the mechanical reason the rest of the campaign matters. A store that only appends is a store whose every article is getting longer, and a longer article is one with more room for a value that used to be true. W3's own registered prediction was that accumulation would favour the store on cost. It is refuted in W3's own results, and reported that way.

The store against the history it distills

Median bytes by sessions of history. Store: 3,199, 4,614, 9,844, 18,511 at depths one, two, four and eight. Transcript pile: 1,432, 2,603, 4,359, 7,880. At eight sessions the store is 2.3 times the pile it distills.

Every lineage is a faint dot; the ringed point is the median of eight. The pile is nearly identical across lineages at each depth, because the fixture holds the history constant, so the spread on the store line is writer behaviour and nothing else. Measured, W3, eight lineages per depth.

The Intervention

One line, at the moment of the write

Part 2 established that prompt-side guidance moves noise without moving habits. So the intervention here is not a prompt. It is a line the tool itself returns, at the boundary where the write happens, naming the fact that the record just got bigger. It carries a byte count, it comes from the tool the writer is already calling, and it states something mechanically exact rather than passing a judgment.

It is one line of text, but it is not one exposure: it fired after 117 of 198 treated writes in the first capture, 99 of 184 in the second, and 126 of 207 in the counterbalanced one. Under the shipped tool the share of superseded values standing climbs toward the ceiling and ends near nine in ten. Under the growth line the first two captures end near one in two.

The denominator moves, which is why the default view is a share. The number of values available to go stale is set by the fixture and climbs from 16 at session two to 96 at session eight. Plotting a raw count against a fixed ceiling would read that schedule as writer behaviour. Switch the toggle to see the count, and the grey dashed line is the schedule itself.

Superseded values standing, session by session

Share of available superseded values left standing, sessions one to eight. Untreated arms end at 92, 91 and 89 percent. Fixed-order treated arms end at 53 and 47 percent. The counterbalanced treated arm ends at 74 percent. The denominator climbs from 16 at session two to 96 at session eight.

Solid is the first model, dashed the second, dotted the counterbalanced capture. In every capture the two arms separate by the third session and stay separated. Note the third pair: its untreated arm lands with the other two, and its treated arm does not, which is the whole of the next section. Measured, W4, W4R and W4C, eight lineages per arm per capture, one run per cell.

The Retraction

The direction survived. The size did not.

The first two captures ran the untreated arm first in every lineage, so condition was confounded with execution order, and running a second model repeated the contrast without touching the confound. The third capture reversed the order in half the lineages and held the model, the fixture, both system prompts, both arm tools, and all 160 assignment identifiers. Five predictions and a reading rule were registered after launch and before grading, which is weaker than W3's and W4's pre-build registration and is recorded as weaker. Four of the five failed.

What came back: the direction held in both halves, and the size did not. Pooled, the counterbalanced capture ran 85 against 71 of 96, where the same model in fixed order gave 88 against 51. The untreated total barely moved. The treated total moved a long way. One lineage of the eight supplies 10 of the remaining 14-point contrast, and dropping it leaves 74 against 70.

So the magnitude is withdrawn rather than qualified. With one realization per cell this design cannot estimate run-to-run variation, let alone rank it against the arm contrast, and a partially matched repeat moved a treated condition far enough to erase most of the difference the first capture reported. What survives is descriptive: across 24 capture-lineage comparisons over the same eight constructed histories, the feedback condition ended lower in 20, tied in 2, and higher in 2. Read that as a directional pattern worth a properly randomised test, not as a treatment effect.

Eight lineages, three captures, and the one that carries the result

Per-lineage endpoints, superseded values standing of twelve. First model: 88 against 51 pooled, all eight lineages lower. Second model: 87 against 45, six lower and two tied. Counterbalanced: 85 against 71, six lower and two higher. Dropping the most leveraged lineage from the counterbalanced capture leaves 74 against 70.

This is the figure the paper is built around, and the one that took a number back. Step through the three captures and watch the treated column move while the untreated column stays put. In the counterbalanced capture the four lineages whose treated arm ran first are labelled, and two of them end higher than their baseline. Then drop the leveraged lineage: a 14-point pooled contrast becomes 4. Measured, three captures, eight lineages each, one run per cell. The pooled magnitude is withdrawn; the direction count is what the paper reports.

Dissociation

Two endpoints that do not move together

If the growth line simply made writers terser, store size and staleness would fall together and the mechanism would be verbosity. They do not move together. On the first model the median lineage lost 32 percent of its standing superseded values while its store shrank 22 percent; on the second, 45 percent against 3 percent; on the counterbalanced capture, 10 percent against 41 percent. The two endpoints come apart, and they come apart in a different direction each time.

That is worth keeping rather than averaging away, because it says a drop in total output is not necessary for the staleness contrast, and it also says these three captures do not identify a single mediation account. Store size and writer output did not track the staleness contrast monotonically. The honest reading is that something at the write boundary changed what writers left behind, and this design cannot say what.

Store size against staleness, per capture

Within-lineage median change from the shipped tool to the growth line. First model: store size minus 22 percent, superseded values standing minus 32 percent. Second model: minus 3 and minus 45. Counterbalanced: minus 41 and minus 10.

If the line worked by making writers write less, these pairs would track. They do not, and on the counterbalanced capture they nearly invert. Measured, within-lineage medians over eight lineages per capture.

Retrieval

A benchmark that had been reading a moving corpus

A separate question runs alongside: if records go stale, does better ranking rescue the reader? An offline known-item benchmark on two real project stores says the shipped lexical ranker is already ahead. BM25 beats cosine similarity over local embeddings in eight conditions of eight, by 15.8 to 37.0 points. The queries come from journal entries that name the records they concern, which is a labeled query set nobody had to write, and also the condition lexical ranking is best at. That caveat is in the paper next to the result.

The benchmark had a defect worth reporting because external review found it and because it was ours. It read both stores at their paths and derived its query set when it ran, so the corpus moved as the projects were used. Freezing it introduced a second defect: the export gave both rankers one truncated string, while the live harness truncates only what it sends to the embedding model and indexes BM25 on whole documents. Nearly half the corpus is over that cap. Switch the toggle to see what that cost: BM25 scored 4.5 to 14.0 points below itself in all eight cells, and an earlier draft reported that gap as evidence the corpus had drifted. It was not the corpus. It was the instrument.

Eight conditions, and the defect that was in our own gate

Top-10 hit rate, eight conditions. BM25 leads cosine over local embeddings in all eight by 15.8 to 37.0 points. Scored over truncated documents, BM25 falls 4.5 to 14.0 points in every condition, which was a defect in the freezing rather than a property of the corpus.

Read the ordering as observed, not explained: these are two different projects with different corpora, and filling ten slots is mechanically easier in the smaller one. The toggle is the part worth sitting with. A gate we built to remove drift from this benchmark introduced a larger error than the drift it removed, and we found that out from a reviewer. Measured, one frozen execution, corpus digest published and contents withheld.

Limits

What none of this establishes

One realization per cell. The design varies eight domains but does not estimate run-to-run variance within a lineage, so no interval is available on any magnitude here, and the counterbalanced capture is the demonstration of why that matters. The counterbalanced capture is a two-period crossover whose terms separate only under a no-carryover assumption one run per cell cannot check, and sequence was assigned by lineage order rather than at random.

The measure counts known superseded values still present in a record. It does not detect historical prose in general. Both harms observed in the programme are reported: W4's untreated arm produced a reader that pulled a reversed policy out of a narrated store and decided wrongly, and W4C produced a second from a treated store, on the same lineage and the same task. Two harms in 96 reader sessions, one from each arm. Nothing here supports "no reader regression" as a statement about the intervention, and that phrase has been removed from the package and the paper.

Eleven rounds of adversarial model review ran against this paper. They found, among other things, a table clipped off the bottom of a page, a figure combining one run's rates with another run's metadata, a count that was right about its arithmetic and wrong about its scope in three consecutive rounds, and the truncation defect above. The errata for every claim corrected or withdrawn ship with the deposit rather than being quietly folded in.

Citations

The paper, the artifacts, and what came before

The paper. Shane Conner, "pi-canon: The Write Desk", Zenodo, August 2026. doi.org/10.5281/zenodo.22057257. The deposit carries every results document, the per-lineage endpoint CSV behind every direction count on this page, the graded reports, the digest-pinned capture contracts, the fixture generators, both arm tools with the diff between them, and the scripts that recompute the paper's numbers from those artifacts.

Part 1. The design, and the four-arm evaluation that measured it whole: shaneconner.com/projects/pi-canon/part-1. Its paper: doi.org/10.5281/zenodo.21890647.

Part 2. Six studies pricing the recall channels: shaneconner.com/projects/pi-canon/part-2. Its paper: doi.org/10.5281/zenodo.21960350.

pi-canon. The package: github.com/shaneconner/pi-canon, on npm as pi-canon, MIT licensed. The growth line this page measures ships in 0.3.0.

canon-bench. The benchmark, the frozen protocol, and the verifier that recomputes each paper's quantitative claims from the artifacts: github.com/shaneconner/canon-bench.

canon-atlas. A viewer for a store like this one, built with pi-canon in mind and kept package-agnostic: github.com/shaneconner/canon-atlas.

BM25. Stephen Robertson and Hugo Zaragoza, "The Probabilistic Relevance Framework: BM25 and Beyond", Foundations and Trends in Information Retrieval 3(4), 2009. The shipped lexical ranker in the retrieval benchmark above.

← Previous Part 2: Pricing Recall

Tech Stack

TypeScript D3.js Pi Agent Memory Counterbalanced Design Python
View Source → ← All Projects