pi-canon Part 2 · Pricing Recall
Part 1 built the memory and measured it whole. Part 2 takes it apart: six preregistered studies, 679 graded cells, and the recall channels priced: touch and recommendation directly, search bounded through the arm that carried it. Some of the design survived its own numbers. Some of it did not.
Scroll to explore
The Question
One end-to-end number, and the refusal it earned
Part 1 ended on 19 of 20 traps avoided against 8 of 20 for a floor with no memory, and it was careful to say what that number was not: package-level by construction, unable to attribute the result to the journal, the spine, or surfacing separately. The last section of that page named an ablation as the experiment the claim needed. This page is that experiment, run until the parts had prices.
The frame that survived contact with the data is not components but channels. Stored knowledge is inert until something carries it back into a session, and this design has exactly three carriers. Touch is unsolicited: a tool call touches a governed asset and the governing article's one-line capsule arrives, unasked. Recommendation is unsolicited too, and reaches further: a retriever ranks the articles that touch cannot reach, the off-spine and declared-rule residue, against the session's active context, and a suggestion rides only when its score stands out. Search is the agent's own act, and it is the only channel that reaches the journal. Two channels push, one pulls, and they are separated here because they fail separately. The studies priced touch and recommendation directly; search has no study of its own, and what this page can honestly say about it comes from the arm that carried it.
The arms that recur through the studies are worth naming once, because every chart below leans on them. Bare is no memory at all. Flat is the same distilled knowledge handed to the session as one plain document. Journal is the project's full history behind search. Pull is articles behind addresses the session must ask for. Push is pull plus surfacing on touch. Flat, pull and push carry the same distilled content compiled three ways; bare and journal are the floors on either side of it, nothing and everything.
Everything on this page comes from the 0.2 paper and its deposit: six preregistered studies of 75, 165, 105, 124, 120 and 90 graded cells, a 128-session surface experiment, and a 52-turn replay on a real store. A verifier shipped with the deposit recomputes 135 headline numbers from the raw artifacts. The worker throughout is one model at one setting, and section by section this page says what that buys and what it does not.
Three ways knowledge gets back in, and what each one measured
The three recall channels between the store and a session. Touch: a tool call touches a governed asset and the article's capsule arrives, unasked; it carried 4.87 of 5 decisions against bare's 0.00 in study 3. Recommendation: a retriever ranks the articles touch cannot reach and a suggestion rides only past the standout gate at 1.4; it added +0.87 to +1.31 rule facts over a search-carrying arm in study 4, and the shipped cutoff cut its volume nine-fold with no detectable recall change in study 5. Search: agent-solicited, the only channel that reaches the journal; no study of its own, and the search-carrying arm scored 0.33 of 2 rule facts in study 4. Discovery is the split beneath all three: 15 of 15 pull cells that asked for the decisive address scored, 0 of 30 that did not ask.
The figure from the paper, with the channels put under your cursor. Select a channel to isolate its path and its price; the prices are the studies' headline results and the sections below carry their qualifications. The line at the foot is the reason the unsolicited channels exist at all: sessions ask only about what they know exists, and the sharpest measurement of that is study 2's pull arm, where every cell that asked for the decisive address scored and no cell that failed to ask did. Structure is the package's; every number is measured, and each is qualified in its own section below.
Addressing
Where reading the whole thing stops working
The cheapest objection to this design is that none of it is necessary: distill the knowledge, hand the session the file, done. Study 1 could not reject that objection, and the honest reading is that at 692 bytes it never could have: the flat comparator sat at ceiling on the four soundly graded decisions, leaving addressing no visible benefit to pay for. What study 1 did establish is the price of the opposite habit, treating history as memory. The arm that answered from the full journal burned a median of 340,119 tokens per session against the flat document's 21,309, for no more correctness. History is for appending, not for reading.
Study 2 made size the factor: the same distilled content, compiled at 0.7 KB, 52 KB and 259 KB, behind flat, pull and push access. The curves cross between 52 KB and 259 KB. At small, flat ties the addressed arms on correctness and wins on cost, 23,553 fewer tokens than pull. At large, the flat document scores 1.47 of 5 while the same content behind addresses scores 4.93, at a third fewer tokens than flat spends failing.
Why flat fails at scale took three corrections to state honestly, and the mechanism is not the one usually reached for. The answer strings still arrive: count them and every flat-large session got all five. What almost never arrives is the answer scoped to the thing it governs. This corpus asks the same question many times for different subsystems, so at 259 KB the document holds 201 occurrences of the answer pattern and only 69 of them govern the asset under test. The session's read window covered a fraction of the file with the governing sections past the cut, and the rival occurrences it did read arrive stripped of their scope. Strict scoped delivery came to 0.33 of 5; fourteen of fifteen cells got none of the five governed facts and still scored 1.21 on average, which is what guessing from plausible rivals looks like. An address is the unit that survives this: fetched by name, arriving with its scope attached.
The same study put a number on discovery, and it is the cleanest number in the programme. Pull requires the session to ask; 15 of 15 cells that asked for the decisive address scored, 0 of 30 that did not ask scored. Asking was perfectly predictive, and the gap between push and pull grew with store size, +0.07 to +0.87. Surfacing on touch is not a refinement of addressing. It is the thing that makes addressing work when the store is too big to already know.
The crossover, and the mechanism under it
Graded decisions of 5 by memory size and access mode, 15 cells per point. Flat: 5.00, 4.93, 1.47 across 0.7 KB, 52 KB, 259 KB. Pull: 4.87, 4.07, 4.07. Push: 4.93, 4.73, 4.93. The overlay shows flat-large's delivery split: answer strings 5.00 of 5, governed facts 0.33 of 5, fourteen of fifteen cells receiving none.
The same distilled content, compiled at three sizes. Flat ties or wins until the store outgrows one read, then collapses; push barely moves across a 370-fold size change, because an address costs the same whatever surrounds it. Switch to what flat-large delivered and the mechanism is on the chart rather than in a footnote: every session received all five answer strings, almost none received a fact scoped to the asset under test, and the two compose through one read window and 201 scope-stripped rivals. The crossover point is environment-specific; what generalizes is the shape, and the study's own report says exactly that. Measured, study 2, 165 cells; delivery split recomputed by the deposit's verifier.
Touch
Same store, same answers, three ways to reach them
Study 3 is the floor the touch channel stands on, and it is built from stores that all contain the answer. Every decision graded in it is answerable from every store in the study; what varies is whether anything makes the answer findable. No address and no surfacing lands at 0.67 of 5. One flat document lands at 2.53. Addressed and surfaced on touch lands at 4.87 against a bare floor of 0.00, with fewer median tokens than the flat document spends, 62,750 against 81,480, and an eighth of what the journal arm reads.
The comparison is a triangulation across arms built for other contrasts, matched on content rather than on bytes, and it is labeled post hoc exactly as the paper labels it. What it is for is the regime question: the difference between 0.67 and 4.87 is not storage, since the knowledge is present in both stores. It is representation, addressing and delivery, moving together. The addressed arms are told at session start that a store exists and where it lives; a flat file's presence tells a session nothing about what is inside.
Touch also had to survive subtraction. When the 2x2 further down found that presenting the memory costs first-pass correctness, the response was to delete surface pieces, and the deletions were priced here: study 3 detected no loss from removing the orientation line, and the whole result held with the reduced surface. What touch buys survived; what the wrapping cost did not.
The answer is present in every store
Three reachability regimes over content-matched stores, graded decisions of 5. No address, no surfacing: 0.67. One flat document: 2.53. Addressed, surfaced on touch: 4.87, against a bare floor of 0.00. The addressed arm also read fewer median tokens than the flat document, 62,750 against 81,480.
The answer is present in every store; what differs is whether anything makes it findable. The rows are drawn from arms built for other contrasts, matched on content rather than bytes, so this is a triangulation and it is labeled post hoc. The bottom row is the shipped configuration, and it is also the cheapest of the three on context read. Measured, study 3, 105 cells; the arm behind the top row was mislabeled in its run config and disclosed as such in the paper.
Recommendation
The channel that pays in exactly one place
Recommendation is the programme's longest arc, and it opens with the channel losing. In study 3 the ranked residue had nothing to add: the arm carrying it ran beside the touch channel, which was already delivering the decisive articles, and the residue itself was filler drawn from the same value pool as the graded decisions. On decisions it was a null against push, and on the end-to-end check, an endpoint added post hoc, exploratory and uncorrected, it was the only memory arm to lose ground, at 27% more tokens. The channel shipped off, with that reason on the record.
Study 4 gave the residue something worth finding, and the same mechanism roughly quadrupled what the session walked away with. The store held declared rules governing no asset, the class of knowledge touch structurally cannot reach. Every recommendation arm beat push with search, +0.87 to +1.31 rule facts of 2. The preregistered claim was that the best of six arms beats push, and p=0.0002 is that selected best-arm contrast, unadjusted; all six beat push, the weakest at p=0.0015, and a Bonferroni correction across the six leaves the best contrast significant. Push carries the search action, so this is a beat over a search-carrying arm; five of its fifteen sessions did search and recovered the origin rule, so the contrast says recommendation increased rule-fact recall beyond that arm, not that search was unused or failed.
The same study killed the channel's original gate. The cutoff was an absolute score threshold, and an absolute score cannot ship: the retriever normalizes against a query-dependent ceiling, so the same quality of match scored 0.25 in study 3, 0.03 in study 4, and 0.30 in a synthetic probe. The shipped value would have been wrong for the next corpus and could silently deliver nothing. Its replacement is a ratio within the query, how far the top article stands out above the rest, and study 5 priced it on the capsule arms: at the shipped 1.4, the measured difference from the uncut channel was minus 0.07 rule facts at p=1.0, an n=15 null rather than an equivalence claim, while suggestions fell from 26.4 per session to 3.0 and the open rate rose from 0.17 to 0.82. Set the cutoff past the decisive reach, at 2.0, and the capsule arm goes silent, 0 rankings passed of 141, and hands back everything it had won, minus 0.73 at p=0.0015; the name arm at the same level passed one ranking of 132.
The arm that closes the argument is the decoy: the channel pointed at a store that genuinely cannot answer. It passed 0 of 139 rankings, suggested nothing, leaked nothing. A ratio cutoff is a relevance test, and a store with nothing relevant never stands out. That silence, more than any delivery number, is the case that the cutoff belongs on whenever the channel is.
Two cutoffs, one shape: volume collapses before recall detectably moves
Study 4's absolute threshold at 0, 0.035, 0.045: suggestions 28, 10, 3.6 per session; open rate 0.27, 0.41, 0.85; rule facts 1.58, 1.38, 1.62 of 2. Study 5's standout ratio over its capsule arms at 1 (off), 1.4, 2.0: suggestions 26.4, 3.0, 0; open rate 0.17, 0.82, silent; rule facts 1.73, 1.67, 1.00.
Raising the cutoff collapses volume and raises the open rate without detectably moving recall, until the cutoff overshoots the decisive reach and silences the channel; both columns chart the studies' capsule arms, and the silence at 2.0 is the capsule arm's, 0 of 141 rankings, while the name arm at the same level passed one of 132. The left column's knob is the one that cannot ship, because its scale does not survive a change of corpus; the right column's is the one that does, and the shape replicates across both. The zero-threshold rule value on the left is the corrected n=12 figure the paper reports. Measured, studies 4 and 5; study 4's threshold levels are independent sessions, not a within-session sweep.
The shipped default is 1.4, selected by three single-cell smoke tests at 1.0, 1.4 and 2.0 that were spent and reported before the protocol froze; the 120 cells then priced that frozen operating point rather than choosing it. The figure below is the argument in its rawest form: every ranking study 5's free-running arms performed, drawn at its measured standout ratio, with the cutoff in your hands. The decisive moments sit apart from the crowd; everything else sits at or below 1.33, with a median of 1.08. A cutoff below the gap passes noise. A cutoff above it silences the channel. The band between is wide, which is why one frozen constant worked at all.
Every ranking the free channel performed. Set the cutoff yourself.
Every ranking performed by study 5's free-running arms, 276 rankings across 30 sessions, drawn at its measured standout ratio. The decisive rankings, one per session, take exactly two measured ratios: 1.8147 in 25 sessions and 1.8280 in 5; everything else sits at or below 1.33, median 1.08. At the shipped cutoff of 1.4, the arms that ran gated passed exactly 15 rankings in 15 cells each; at 2.0 the high-capsule arm passed 0 of 141. The decoy arm's 139 rankings topped out at a ratio of 1.23, so at 1.4 it stayed silent.
Each mark is one ranking from the arms where the channel ran free, at the standout ratio the trace recorded. Drag the cutoff, or use the slider, and the counts recompute from these measured values: rankings passing, and sessions that would have received at least one suggestion. Two honesty notes belong here rather than in fine print. Moving the cutoff over these rankings answers what this set of rankings would have passed, not what a session run at that cutoff would have done; sessions unfold differently once suggestions arrive, which is why the measured anchors from the arms that actually ran gated are printed beside the line. And the gap the cutoff sits in is a property of this corpus and this retriever; the paper prices the constant, not a law. Turn on the decoy overlay for the other half of the argument: a store with nothing relevant to say never crosses the line, at any setting worth shipping. Recomputed live from the per-ranking trace in the deposit, runs standout, kit/trace.jsonl.
The Surface
Presenting the memory costs more than anyone designed for
The most uncomfortable result in the programme has nothing to do with recall. A development observation suggested that merely presenting the memory system, before it stores or retrieves anything, was costing first-pass correctness: sessions with an inert memory surface shipped worse first attempts at the same task. The 2x2 below was frozen to localize it. Two factors, each inert: the tool schema, a description the model rereads every turn, and the orientation line, one steered sentence telling it to read before working. Four arms, 32 sessions each, one plant task, first-pass correctness graded deterministically.
With neither piece, 25 of 32 sessions shipped a correct first pass. With both, 8 of 32. Each factor's main effect is negative and significant, the tool schema at −1.05 log odds, the orientation line at −1.32, and the interaction straddles zero, so additive and worse-than-additive cannot be told apart at this size. One thing the design cannot say, and the paper's external review made it say so out loud: which piece costs more. That comparison reduces to 19 of 32 against 21 of 32, and it resolves nothing.
What followed was subtraction, executed by the numbers. The orientation line was removed, priced by study 3 detecting no loss without it. The check-back, a session-end nudge to revisit the store, was removed when study 6 measured it as a null at 0.96 power. The capsule on surfaced lines stayed, and the paper is explicit that this one is a judgment call rather than a verdict: one exploratory result favors capsules, a bare address spends a read round trip to teach what the capsule teaches on the line, and the data left it open. The caveat that cannot be subtracted: studies 4 through 6 ran with the old surface. Their internal contrasts are unaffected, and their absolute levels carry it.
Two inert pieces, four measured cells
The 2x2 surface experiment, 128 plant sessions. Neither piece: 25 of 32 correct first passes. Tool schema alone: 21 of 32. Orientation line alone: 19 of 32. Both: 8 of 32. Tool main effect minus 1.05 log odds, p=0.0074; orientation minus 1.32, p=0.0008; interaction straddles zero. Which piece costs more is unresolved: that comparison reduces to 19 of 32 against 21 of 32.
Flip the two pieces on and off; every state you can reach is a measured cell, 32 sessions each, nothing interpolated. Both pieces are shams, a schema with an inert implementation and a line pointing at a store that persists nothing, so what is being priced is presentation itself. The one reading the chart refuses to support is a ranking between the two pieces; the deletion decisions rested on study 3 and study 6 pricing each removal, not on this comparison. One plant task, measured 128 times: this powers the result up, it does not generalize it. Measured, surface experiment, chain 06.
The End-to-End Check
The rerun that had to disown its own recall numbers
Study 6 reran Part 1's chain shape under a frozen protocol, 30 repetitions, three arms, and its recall story begins with a contamination. The earlier harness left the prior sessions' transcripts on disk at recall time, and the bare arm, with an empty store, read at ceiling: 14 to 15 of 15. That number was not memory. It was the agent finding the transcripts. Study 6 removed them and forked recall from a hash-verified pre-probe snapshot, and the floor told the truth: bare fell to 2, 19, 0 and 1 of 30 across the four facts while the memory arms held 29 to 30 of 30. The one fact bare still gets, scheduler ids, is the one the checked-out code answers by itself.
The trap replicated at double the prior repetitions: a session two days later attributes a planted sweep to its actual job in 0 of 30 bare cells against 19 to 21 of 30 with memory, and every arm's own current task sat at 29 to 30 of 30 with every pristine check green. Memory arms do the work and carry the knowledge; bare does the work and structurally cannot.
The check-back died here, and the manner of its death is the method in miniature. The preregistered endpoint came back a risk difference of exactly 0.000 with an interval of −0.223 to +0.223, at 0.96 power for the effect it was designed to find. The post hoc trace made the null legible: in the nine canon plants that shipped broken, the nudge was delivered in eight, and the number of sessions that then touched the flagged file was zero. Delivered, and not acted on. The feature was removed the same day, and the removal is the study's contribution: a null at power, converted into one less thing every session pays for.
What the recall floor reads, before and after the fork
Recall share judged correct, two harness regimes. Transcripts in the work tree: bare reads 14 to 15 of 15 on every fact despite an empty store. Transcripts removed, recall forked from the pre-probe snapshot: bare collapses to 2, 19, 0, 1 of 30 while canonbase holds 29 to 30 and canon 29 of 30. Scheduler ids is the one fact the checked-out code answers by itself.
The toggle is the finding. On the left setting, the floor arm is reading the prior sessions' transcripts off disk, and a recall benchmark that leaves them there is measuring archaeology rather than memory. On the right, with recall forked from a hash-verified snapshot, the floor answers only what the project state itself can answer, and the gap it opens against the memory arms is the estimand the protocol froze. The correction is disclosed in the paper rather than smoothed over, because the contaminated number was ours. Measured; left panel from the prior harness at n=15, right panel study 6 at n=30 per arm.
The Replay
A store it was never tuned on
Every constant above was priced on one corpus family, so before freezing, the shipped retriever was replayed over a real working store: 33 articles of live project memory, 52 real session turns, trace-level capture. The first thing the replay found was a failure. A store that small and saturated lets ranking find a regime the priced corpus never enters: once the still-eligible residue is a tail of near-zero scores, the standout ratio is a ratio over junk, and it explodes. A 4.79 standout on a 0.101-score match, the ratio anti-correlating with relevance, the whole 33-article store riding over the session down to scores of 0.002.
The fix is a consumed floor: while the cutoff is active, the crowd floors at the strongest already-delivered article the same query raises, a bar that moves with the query and the store rather than sitting at any fixed score. Re-run on the same trace, the rides drop from 33 to 15, the minimum observed ride score rises from 0.002 to 0.075, and the strongest ride improves, 0.184 to 0.264, because an article not spent on a junk query is still available for the query that deserved it. The bound on all of this is stated in the section title the paper gives it: none of these 52 turns needed history, so the replay measures the false-positive side only. It shows the drained-store failure going away on the same trace; it cannot show that surfacing helps when history is needed. Whether value ports is the next study, not this one.
Every article that rode, on the same 52 turns
The portability replay, 33-article store, 52 real turns. Unguarded: all 33 articles ride, scores 0.002 to 0.184. Guarded by the consumed floor on the same trace: 15 rides, scores 0.075 to 0.264; the strongest ride improves because the store was not spent before the query that deserved it.
Each dot is one ranked article that rode out to the session. The top row is the drained-store failure: with no floor, the ratio test degenerates on a saturated store and everything eventually stands out over junk. The bottom row is the same trace with the consumed floor, which is the shape that shipped. Same turns, same store, one guard apart; this is a same-trace comparison, so it can show the drained-store failure going away and cannot show transfer of value. Measured, portability replay, both traces preserved in the deposit.
The Verdicts
What shipped, and why, by the numbers that decided it
The 0.2.3 configuration is the studies' verdicts, executed. Four kinds of verdict, and every entry below names its evidence.
On by default
- Addressing. It is what survives scale: the crossover, study 2, with the delivery mechanism counted rather than assumed.
- Surfacing on touch. It is what makes addressing discoverable: asking predicted scoring exactly, and push over pull grew with store size.
- The append-only journal. Never read whole, never surfaced unsolicited; study 1's 340,119-token median is the price of treating history as memory, and search is the only channel that reaches it.
Functional, off by default
- Recommendation. It measurably paid exactly where study 4 put it, and it cost where study 3 put it, a store with nothing to add, on an endpoint that study added post hoc. The package cannot know which store it faces, enabling the channel is one option away, and it needs a retriever the package deliberately does not choose for you.
- The standout gate, when enabled. The ratio at its measured 1.4 with the drained-store guard; the absolute threshold it replaced cannot ship at all.
Removed by their own measurements
- The orientation line. Each surface piece cost first-pass correctness in the 2x2, and study 3 detected no loss from removing the line.
- The check-back. A null at 0.96 power, with the post hoc trace showing the nudge delivered and the file never opened.
Kept as a judgment call, labeled as one
- Capsules on surfaced lines. One exploratory result favors them, minus 0.53 rule facts for bare addresses where the channel ran free, and the design argument is that a capsule teaches on the line while an address spends a read round trip to teach the same thing. The data left it open; the paper says so.
The Journey
Each study broke something in the one before it
Read in order, the eight runs are not a suite; they are a sequence, and the sequence is the finding. Study 1's null said the comparator was at ceiling, so study 2 made size the factor and found the crossover, and its review corrected the delivery instrument twice on the way. Study 3 shipped the recommendation channel off on a post hoc cost signal, mislabeled its own decoy, and saturated four cost measures, which taught the ceiling lesson the 2x2 then exploited to price the surface itself. Study 4 gave the channel something worth finding and paid for it: the run withdrew its own naming comparison through a re-offer defect, exposed why absolute cutoffs cannot ship, and lost a cell to a directory collision its correction later traced. Study 5 replaced the threshold with the standout ratio, sized its own internal null contrast, and rebuilt the decoy into a real control. The replay then found the regime the priced corpus never enters, and the drained-store guard closed it with one mechanism verified on the same trace. The three motivation runs before study 6 produced a chain-validity fix, the contamination discovery, and two observational hints; study 6 then quantified the contamination, replicated the trap at doubled repetitions, and returned the check-back null, and the day it reported, the check-back came out of the package. Every one of those breaks is in the record with a date on it.
The programme, card by card
Eight runs in order, with headline verdicts. Study 1, 75 cells: beats searching by +2.07 of 5; ties flat, comparator at ceiling. Study 2, 165 cells: the curves cross, flat 1.47 against addresses 4.93 at 259 KB; asked 15 of 15 scored, not asked 0 of 30. Study 3, 105 cells: push 4.87 against bare 0.00; the channel adds nothing here and ships off; decoy mislabeled. The 2x2, 128 sessions: first pass falls 25 to 8 of 32 with both surface pieces. Study 4, 124 of 135: +0.87 to +1.31 rule facts over an arm carrying search; an absolute cutoff cannot ship. Study 5, 120 cells: at the ratio's shipped 1.4, minus 0.07 rule facts against uncut at p=1.0, at a ninth of the volume; the rebuilt decoy stays silent, 0 of 139. The replay, 52 real turns: a drained store lets all 33 articles ride, and the consumed floor fixes it on the same trace. Study 6, 90 cells: the check-back is null at 0.96 power and is removed; recall is decontaminated, bare falls to 2, 19, 0, 1 of 30 while memory holds 29 to 30.
A filled square is a measured result; an open circle is a null the design bounds; a crossed square is a defect found in one of our own instruments, and the map needs that glyph four times, which is the point. Select any card for the fuller story, including what broke and what the break taught the next run. The replay card's bound is the no-need limitation above; the 2x2's caveat, that studies 4 through 6 ran with the surface it priced, rides on its card. Every number is the paper's; the dispositions live in the claims map shipped with the deposit.
The Fine Print
What none of this establishes, in the same breath
One worker model, one setting, throughout: every number above is a property of that pairing, and eligibility-style results move when the worker does. One corpus family built to have scoped rivals, and one task shape, sessions making graded decisions against planted knowledge; the crossover point and the standout gap are properties of that corpus, and the paper claims the shapes, not the constants. The 2x2 grades one plant task 128 times, which powers it without generalizing it. Studies 4 through 6 ran with the surface the 2x2 later priced. Study 4's attrition is accounted rather than summarized: 124 preserved of 135 launched, the loss differential over six arms, one grade dropped with its observation withdrawn. And the replay bounds the false-positive side only.
The review behind the numbers is the same discipline Part 1 used, grown teeth: five rounds of model review over the evidence base, two over the conversion, an external reviewer model over the built PDF, and verification rounds over every revision, each round certified by the next rather than by itself, every finding and disposition recorded in the claims map beside what replaced it. All of it is fact-checking by language models against the artifacts, a way of catching our own errors before they publish. None of it is peer review, and it substitutes for none. Nothing here is replicated elsewhere.
The deposit is the arbiter. Beside the paper: the per-cell artifact trail for all 679 study cells and the 128 surface sessions, worker transcripts and work trees and surfacing traces included, the frozen protocols with their amendment histories, the sealed corpora, both replay traces, and a claims map recording every review round. A verifier recomputes 135 headline numbers from those files under an identity header binding the run to the release. If a check fails, the paper and the deposit disagree, and the deposit is right.
Citations
The paper, the artifacts, and the work this stands on
The paper. Shane Conner, "pi-canon: Pricing Recall in Long-Term Memory for AI Agents", Zenodo, August 2026. doi.org/10.5281/zenodo.21960351. The deposit beside the PDF carries the artifact trail for every run above, the figure sources, the claims-to-evidence map, and the verifier that recomputes 135 headline claims from the files.
Part 1. The design, and the four-arm evaluation that measured it whole: shaneconner.com/projects/pi-canon. Its paper: doi.org/10.5281/zenodo.21890648.
pi-canon. The package: github.com/shaneconner/pi-canon, on npm as pi-canon, MIT licensed. The release this page describes is 0.2.3.
pi-fold. The companion package on the working tier, separate and optional in both directions: shaneconner.com/projects/pi-fold.
BM25. Stephen Robertson and Hugo Zaragoza, "The Probabilistic Relevance Framework: BM25 and Beyond", Foundations and Trends in Information Retrieval 3(4), 2009. The reference retriever behind the recommendation studies, and the normalization behavior that killed the absolute threshold.
The AGENTS.md convention. agents.md. The one-file shape the flat arms compile to.
Lost in the Middle. Nelson Liu and colleagues, 2023. arXiv:2307.03172. Positional degradation in long contexts; study 2's flat-large failure is a harder one, sections never entering context at all.