What this run does and does not show
On the 99 current-state probes of the pooled run (3 depth buckets x 3 seeds), naive similarity
retrieval answered 60% correctly and asserted a stale value on most misses. Sapience (L1+L2
consolidation) answered 100%. McNemar on the paired probes: 40 flips in Sapience's
favor, 0 against, p<0.0001.
The honest caveat, stated up front: a steelman baseline that keeps similarity retrieval but
re-orders the retrieved snippets by write order recovers to 98%, statistical parity with
Sapience on this probe set (2 flips, p=0.50). On this benchmark the failure is not retrieval reach,
it is that similarity ranking discards update order. Sapience resolves supersession at write time
instead of hoping the prompt order saves it; the recall-witness callouts above show the baseline
usually retrieved the current value and still picked a stale one.
Probe selection for this page: current-state probes where the naive arm was wrong with a stale
assertion and Sapience was correct, restricted to rows whose recorded answer strings are complete
(a few rows store truncated text and were excluded), chosen for topic legibility and depth variety.
That is 8 of the 40 such probes; it is a demonstration gallery, not the statistic. The statistic is
the paired numbers above, computed over all 99 probes.
From the pre-registered run of 2026-07-02; full methodology in the paper.
Source: research/memory-experiments/l2_synthesis/runs/grid_wiki/ (results) and
scenarios/grid_wiki/ (probe definitions). Reader: claude-haiku-4-5, both arms, matched evidence
budget (1,400 tokens). Corpus: real MediaWiki revision snapshots, revision timestamps stripped; the club chains shown here carry no in-text date signal (27 of 99 probes in the full run retain natural in-text date or ordinal signals, reported as a separate stratum in the analysis). Total run cost $3.60. Zero empty or failed rows in
any arm.