One row per claim. Each carries its number, its n and seeds, the caveats that travel with it, the result path, and the date it was last verified against that path. The short version is the evidence hub; this is the version for diligence.
Where an older page and this ledger disagree, the ledger wins and the caveat says "supersedes". Every BABILong row states that the retrieval pipeline is task-adapted, in the same sense that the fine-tuned specialists on that leaderboard are task-trained; row 3 is our own audit of how much that adaptation contributes. Numbers that do not appear here are listed at the end with the reason: superseded, retired, or retracted.
| # | Benchmark | Claim | Number | n, seeds | Caveats | Result path | Verified |
|---|---|---|---|---|---|---|---|
| 1 | BABILong qa3 ladder, official HF cells | Multi-hop accuracy holds from 128K to 10M tokens on the leaderboard's official HuggingFace cells | 86 / 87 / 87 / 84 / 86 at 128K / 256K / 512K / 1M / 10M, official scorer | n=100 per cell; seed 42 canonical (per-task predictions in PR #19); seeds 123 and 456 supplementary | Task-adapted retrieval pipeline (row 3 audits the contribution). Supersedes the August 5 page's 84.0 at 1M under the harmonized protocol and the data room's 86.0 at 128K: same cells, now scored by the official scorer. | research/olympiad/leaderboard_submission_2026_08_25/ (official_metric_report.json; seeds_supplementary/; PR #19) | 2026-08-25 |
| 2 | BABILong avg(qa1-qa5) at 10M | #1 overall on BABILong at 10M tokens | 97.2 official scorer (97.5 = 3-seed judge figure); best published 76.6 (ARMT, fine-tuned on these tasks) | n=100 per task per seed; 3 seeds, zero variance (deterministic), so effective n=100 per task | Scoped AT 10M, never "every length". qa1, qa4 and qa5 are extraction-solvable (naive solvers reach 100 without a language model); the reasoning load sits in qa2 (34 to 100 over the naive floor) and qa3 (32 to 87.7, 3-seed figure; the seed-42 official cell is 86, row 1). 97.2 is the headline on every surface; 97.5 appears only labelled as the judge figure. Task-adapted pipeline (row 3). | research/olympiad/leaderboard_submission_2026_08_25/ (official_metric_report.json; SKEPTIC_VERDICT.md) | 2026-08-25 |
| 3 | BABILong zero-adaptation audit at 10M | What the task adaptation contributes: the shipped generic pipeline, zero task adaptation, same tasks, same scale | qa1 4.7 / qa2 4.3 / qa3 6.3 / qa4 45.7 / qa5 62.7 (24.7 average) | pooled 3 seeds (42 / 123 / 456), n=300 per task; $3.58 total | qa1-qa3 sit below the 16.7% guess floor via abstention. Retrieval-grounded qa1 and qa2 = 0/300: the bAbI template grammar is invisible to semantic similarity. qa5 carries ~10.7pp of example-echo luck. The shipped LLM-lasso ingest was not exercised (verbatim-store, benchmark-standard ingest). This row is the disclosure that travels with rows 1, 2, 4, 8 and 14: those pipelines are task-adapted; this one is not. | research/olympiad/_local_runs/avg10m_ext_2026_08_25/zero_adaptation/ (zero_adaptation_summary.json); prereg docs/babilong-zero-adaptation-prereg-2026-08-25.md; audit post | 2026-08-26 |
| 4 | BABILong qa3, 30M and 100M stretch | Flat from 2M to 100M tokens on qa3 (multi-hop) | 2M 81.0 / 5M 80.0 / 10M 80.0 / 30M 80.0 [75.1, 84.1] / 100M 79.7 [74.8, 83.8], official scorer | 3 seeds, n=300 per scale; ~685 reader tokens per question at every scale; ~$3 total | Generator-built corpora, item-paired to the 4K-1M ladder, PG-19 train-pool filler (seeded). Near-deterministic, so 3 seeds are not 900 independent items. Task-adapted pipeline (row 3). Two qa3-at-10M values exist on purpose: 86 in row 1 (official HF cells) and 80.0 here (paired generator corpora); same task, same scorer, different corpora. Firsts: first published multi-hop (qa3) results beyond 10M (ARMT published qa1 at 50M, 2024); first BABILong/RULER-class reasoning measurements at 100M (Magic.dev HashHop, in-house 2024, unreleased model). Supersedes the August 5 page's judge-scored "~82% at 2M, 5M and 10M". | harness research/olympiad/harness/run_babilong_qa3_100m_stretch_2026_08_25.py; run family /Volumes/X10 Pro/qa3-100m-stretch-2026-08-25/ (Labs); figure figures/fig-babilong-100m-ladder-2026-08-27.png | 2026-08-27 |
| 5 | RULER variable tracking, 1M to 10M | State tracking holds from 1M to 10M tokens | 24/24 at 1M; 16/16 at 2M, 5M and 10M; exact match | 2 seeds; corpora sha-pinned | ~260 tokens read per query; the 2M-10M cells cost $0.34 total. The structural chain-walk does the work; the claim is scoped to that mechanism. No frontier comparison number on this row: the June frontier cells are retired pending a clean re-run, and a SOTA claim waits for publication with the full protocol and harness. Supersedes the August 5 page's 1M-only framing (n=24, 3 seeds) and its frontier comparison. | research/ruler_fullsuite/ (runs/phase1/sap_ladder/; synth_vt_ladder_2026_08_21.py) | 2026-08-24 |
| 6 | MuSiQue, beyond the window | Iterative retrieval over the store holds where a truncated full-context reader collapses | 3-seed ladder 43 to 28 EM (Sapience, iterative) vs 19 to 0 (truncated full-context reader) | n=30 items; 3 seeds; one reader (deepseek-v3.2); official exact match, no LLM judge | A gentle decline, not a plateau. FiD / cross-encoder steelman (2026-08-09): ties in-window (47.8 vs 44.4 pooled EM, not significant), never exceeds iterative, and retrieves MORE gold (support-recall 0.94 / 0.92 / 0.90 vs 0.76 / 0.67 / 0.63), so the residual is reader composition, not retrieval. 5-hop is unsolved. Supersedes the August 5 page's "+33 to 37pp" over-window framing. | research/musique_beyond_window_repro/runs_fid/ (results_fid.jsonl; analysis_fid.json) | 2026-08-24 |
| 7 | Supersession (currency) | When facts change, the old version cannot outrank the new | 98% current-answer accuracy on 108 update probes (2 misses), +70.4 pooled over similarity retrieval; structure-aware vs similarity retrieval +67 to +78 points on date-obscured chains | 3 seeds; discordant-pair McNemar | Hand-ordering steelman recovers parity, which isolates write order as the signal similarity retrieval discards. The +67 to +78 is a range of per-cell deltas, not a confidence interval (skeptic correction R-1 on the 2026-08-12 clean re-run: binding-cell range +66.7 to +74.1, pooled +70.4, Newcombe 95% [+59.3, +78.4]). The 0/81 figure is the WEIGHT-LAYER count (L3 replay with the supersession filter: 0 of 81 stale facts persist vs 14 of 81 unfiltered; KO e2ee87cc, skeptic verdict docs/akshat-three-curricula-skeptic-verdict-2026-08-27.md) and is not a retrieval number. Real-Wikipedia currency numbers on the August 5 page (+40.4pp; 100 / 98 / 59.6) are not carried onto this ledger. | research/memory-experiments/l2_synthesis/runs/currency_2x2_2026_08_12/ (docs/currency-clean-rerun-skeptic-verdict-2026-08-12.md); docs/akshat-three-curricula-skeptic-verdict-2026-08-27.md (L3 replay, B_super arm) | 2026-08-24 |
| 8 | HippoRAG-2 at 1M | The strongest published structured-memory baseline, on identical items | 15.5 (HippoRAG-2) vs 84.0 (Sapience) on the identical 100 BABILong qa3 items at 1M; ~2,500x our token footprint | n=100 items | Framing: HippoRAG-2 "does not transfer to this regime", never "weak". Our side is the task-adapted pipeline (row 3). | docs/hipporag-baseline-final-2026-08-06.md (run: baseline-arm/_runs/hipporag_qa3_1M_2026_08_02/, report_hipporag2.json) | 2026-08-24 |
| 9 | Published null | Within a single window, single shot, static corpus: a tie | +0.00pp, p=1.0, vs a frozen frontier model with strong retrieval | pre-registered; losing condition published before the run | The one regime where structure has nothing to add; reported as it landed. | pre-registration and post-mortem in the data room (on request) | 2026-08-24 |
| 10 | Repo-Evolution (coding currency) | What is current in an evolving codebase | 88% (frontier reader) / 82% (8B reader) with Sapience vs Mem0 51%, grep 29%, notes file 25%, flat RAG 18% | same model, same questions; gold derived mechanically from git, no LLM judge | Skeptic-audited 2026-07-30; live on the coding evidence page. Supersedes the August 5 page's "+30 to 65pp" and "26 of 27 identical" framing of the same result. | research/repo_evolution/; coding evidence page | 2026-08-27 |
| 11 | Repo-Task pilot | Task-level coding pilot: two observations, no result yet | 11.4K vs 97.5K tokens; 133K vs >300K (censored) | two single observations (n=2); split sign | NO completion rates, NO effect sizes. Agents consulted memory voluntarily in ~0-8% of episodes. A powered run is in progress; nothing on this row is a result. | research/repo_task/ (pilots A4 / A5) | 2026-08-27 |
| 12 | Weight path (lab) | Retention on the weight path | adapter-trained retention 47 to 75 points above the no-adapter control; matched-or-better acquisition | as in the figure of record | NOT shipped: what ships is the store, retrieval and consolidation over a frozen model. The "50 to 75" phrasing is retired. The +6.01pp weight-consolidation figure is a different measurement (capability uplift on held-out tasks, row 17) and is never summed with or compared to this retention delta. | papers/sapience-frontier/figures/fig_l3_retention_2026-08-09.py (.png) | 2026-08-24 |
| 13 | Query latency at 1M | Per-query latency stays ~1.4s as stored history grows | ~1.4s per query vs ~10s frontier-alone, warm cache, at 1M (~7x; ~20x cold-prefill as a parenthetical only) | n=8 per cell | Older, single study (2026-06-14). Stated as "~7x", never rounded up. | KO b51b4646 (skeptic-approved 2026-06-14) | 2026-06-14 (older, single study) |
| 14 | Matched pairs, 1M and 10M | Same frozen model, with and without Sapience | 1M pair (BABILong qa3): 34.4% alone vs 66.7% with Sapience (+32, McNemar p=1.5e-5); 10M pair: 45.0% non-abstention with Sapience (27 of 60, 3 seeds; 44.4% pooled) vs 0 of 20 for the same model given its full 1M window (12.5% pooled, the correct answers being abstentions); 0% where the evidence lies beyond the readable window (window arithmetic) | 1M: n=90 item-paired, both arms Sonnet 4.6; 10M: internal 10M corpus, held-in cells, matched protocol; non-abstention n=60 (n=72 pooled), 3 seeds; the full-1M-window run is one deterministic seed | The 10M caption reads "internal 10M corpus, held-in cells, matched protocol"; it never carries a benchmark name (Ruling A). Approved runs 2026-06-19 (close-out) and 2026-06-21 (full window). Attestation: check_calibration_leakage.py --source beam10m exit 0 (run 2026-08-27 on studio-1 over 113 jsonl files under research/olympiad/_local_runs/beam_forward_growth_2026_06_16/, incl. held-in seeds 42/123/456). Task-adapted pipeline on the 1M pair (row 3). Supersedes the data room's +37.8pp (n=45) and the August 5 page's +32.2pp framing of the 1M pair. | research/repo_evolution/fig_public_pairs_2026-07-30.png; 1M pooled analysis research/olympiad/_local_runs/qa3_1m_stage2_pooled_analysis_2026_08_03.json | 2026-07-30 (figure); attestation pending |
| 15 | Token economics | Reads scale with the query, not the history | a few hundred tokens per answer instead of hundreds of thousands; ~1,800x cheaper per query on the licensed figure | money-plot figure of record | Licensed shape from the features canon. The ~961x lookup-probe figure is a different measurement (store-growth crossover, row 18) and coexists with this row under its own scope; the other older-page variants (~304,000 tokens re-sent past the window; $0.007 vs $13 per query at 1M) are not carried forward; the money plot is the single figure of record for the per-query ratio. | figures/fig-scale-v5.1-frontier-to-10M-2026-08-28.png (v4 superseded) | 2026-08-27 (canon) |
| 16 | Modern RAG, run by us, on the same 10M cells | The strongest retrieval baseline a critic would build, on the leaderboard's official qa3 cells at 10M | hybrid BM25 + dense + reciprocal-rank fusion (cross-encoder variant) at a 2-4K-token budget: 11.3 to 12.3% on qa3 at 10M, official scorer (9.7 to 11.0 strict judge); at Sapience's own token budget, 3%; reads 4.9x the tokens Sapience reads per question | same official HF 10M cells as row 1; $0.021 | Pre-registered with its kill condition. About the old leaderboard RAG entry (11). The Sapience side of the comparison is row 1 (task-adapted pipeline, row 3). Skeptic verdict KO 87c5e4fa (approved 2026-08-27): "3 reader replicates", never "3 independent seeds"; strict-judge values quoted alongside official; achieved budgets ~3.4K of 4K disclosed; scope retrieve-then-read at 4K or less (FiD-fusion at 10M and a wide top-150 arm are future work). | research/qa3_steelman_rag_2026_08_25/ (run family /Volumes/X10 Pro/qa3_steelman_rag/); prereg docs/qa3-10m-steelman-rag-prereg-2026-08-25.md | 2026-08-27 |
| 17 | Consolidation into weights (DREAMS) | Weight consolidation adds capability on held-out tasks | +6.01pp held-out accuracy from training-time consolidation (LoRA, into weights); 95% CI +3.16 to +8.87 | 10 of 10 seeds; replicated on 5 fresh seeds | DREAMS +6.01pp = capability uplift on held-out tasks (what weight consolidation adds); L3 47 to 75 points (row 12) = retention vs sequential SFT (what replay prevents losing). Different measurements, never summed or compared. Within-domain replay in DREAMS hurts (-6.7 to -8.9pp) while L3 replays same-domain content to protect it; the uplift-vs-retention framing resolves the apparent contradiction. Identical inference in all arms, so the lift is not test-time compute. Lab result, not shipped. | Discovery by Dreaming, arXiv 2607.16256; consolidation ruling KO 897fe616 (2026-09-02) | 2026-09-02 |
| 18 | Store-growth crossover (lookup probes) | Retrieval over a persistent store holds where dumping the store into context collapses | store-in-context 78.3 / 41.7 / 13.3 as the store grows past the window vs 100% [94, 100] retrieving from the store; ~961x fewer tokens per answer at the largest store | 3 seeds; lookup probes | Lookup probes only; the structure-specific gains are a separate study. A different measurement from row 15's ~1,800x per query at 1M (qa3 payload vs raw window); the two coexist under different scopes and are never combined. | KO 45fce2e5 (skeptic NUMBERS APPROVED; re-passed 2026-09-02, verdict item 19) | 2026-09-02 |
| 19 | NoCha-style, real novels beyond the window | Claim verification over public-domain novels once the book exceeds the reader budget | 38.6% structured retrieval vs 20.1% positional truncation, pair accuracy; item-level +18.5pp, cluster-bootstrap 95% CI [6.9, 29.6], paired-t p=0.0018 | 3 seeds are retrieval-order reshuffles over the same 63 book-pairs, so effective n=63 (never 189); same reader (gemini-2.5-flash) and the same 24K budget in both arms | "NoCha-style": public-domain Gutenberg classics that sit in pretraining, not the real 2023+ NoCha set. Beyond-window regime only (24K reader budget vs 67K to 251K-token books); no cross-regime pooling. Single-claim TRUE/FALSE verification scored as pair accuracy. Scope is retrieval, not multi-hop or consolidation reasoning. The pooled-189 McNemar p and the single-seed 46% are never cited. | research/olympiad/results/nocha_beyond_window_2026_06_27/; docs/nocha-panel-2026-06-26.md ("MULTI-SEED COMPLETE"); skeptic KO eb9fcdcc (2026-07-03, deck-eligible, scoped); locked aggregation KO 1e6c892b (2026-06-30) | 2026-09-02 (consolidation ruling) |
| Methodology | The official BABILong scorer penalizes chain-of-thought preambles | Official first-sentence substring metric scores CoT preambles as wrong: the language model alone under load emits multi-sentence outputs 0/24/65/43% of the time at 4K/64K/128K/1M, so its official score reads 74.0/45.5/18.5/26.7-31.1 vs dual-judge 74.0/55.9/60.0/31.1-34.4; every disagreement is a CoT cutoff, none in the other direction. Sapience is identical under both scorers (83/84/86 vs 84/84/86) | re-score of persisted predictions, 2026-08-27 | Consequence: cortex-only arms are reported under the dual judge in any figure that includes them; official-scorer comparisons that would flatter our arm are not used. Gemini "16.7% n=6" at 1M retired (3 items attempted twice; no clean multi-vendor 1M cell exists) | research/olympiad/_local_runs/official_rescore_cortex_2026_08_27/rescore_summary.json; KOs 5ce6d3f1, 57cef0d4 | 2026-08-27 | |
| BEAM-public (10M tier, full 200) | Frozen-config baseline, published either way | 0.401 mean graded score (0.365 on the clean held-out subset) vs Hindsight re-judged under the same pinned judge: 0.677 (0.662 clean) | n=595/600 scored, 3 seeds (0.392 / 0.410 / 0.400), bootstrap CI about +/-0.06 | Dev conversations 1 and 10 sit inside the headline; matched pinned judge gemini-3.5-flash; 5 content-correlated wedge rows excluded; Hindsight's own leaderboard figures (64.1 / 73.5) are under a different judge and are not the comparison; the query-time v2 arm and later versions are reported as they land. Task-adapted note does not apply (this is the shipped frozen configuration) | studio-1:/Users/oz/X10Pro/beam-public/runs_full200/seed{42,123,456}/; re-judge research/beam-public/runs/rejudge_hindsight/rows.jsonl (0.6769 recompute of record); KO e8c629b8 | 2026-08-27 | |
| Consolidation tier (write-time aggregation gate) | Exact aggregates materialized at write time beat a compute-matched counter | +22.8pp (95% CI +8.1 to +38.2) over a compute-matched language-model counter emitting the identical sentence, on covered aggregation questions | Paper v4.3, result (3); single reader at a single scale; synthetic instrument | The same contrast runs -3.3pp on trend questions and is exactly neutral on single-episode controls; broader semantic and schema consolidation is implemented but unestablished; the earlier 26.7 to 66.7 synthesis figure is retired (see below) and is never cited | Sapience frontier paper v4.3, Zenodo 21937795; v4.1 precursor +11.1pp with skeptic verdict docs/l2-functional-skeptic-verdict-2026-08-10.md; KO eff68f7f | 2026-08-14 | |
| Discovery channel: source lead time | Cross-domain sources sit surfaceable for years before the field imports them | 29 of 155 realized imports had their true source in a skimmable top-40 slate more than a year early; median lead time 12.5 years across those 29. In the per-case lag study, a median 72% of the delay from problem to solution was reducible in principle | 155 realized imports, abstract-level ranks (a floor); lag study per-case median reducible fraction 0.72, n=48 | Floor framing: these ranks use abstract-level distillation, so a full-text pass can only add targets. The famous cases (Boole to Shannon, 83 years) are documented history, separate from the measured set. Within a window, dense retrieval out-recovers this channel; its value is the stratum retrieval cannot enter. Targets are a held-out future, so a solve cannot be memorization | figures/fig-reachback-2026-08-09.py over lt_leadtime.json; lag study KO eabcb6e2 (deploy #50, skeptic-approved); foresight slate on the archived research overview | 2026-08-09 | |
| Frontier reference cell: Fable 5 alone at about 660K (BABILong qa3) | A frontier model alone, at its effective window, on the same task | 64.0% (32 of 50), Wilson 50.1 to 75.9 | n=50 of a planned 60 (one stopped shard, low bias risk); single pass; dual judge unanimous | Claude Code harness transport is part of the cell; a baseline reference star, not an architecture headline; this stays the paper's Fable anchor and no true-1M Fable cell exists (Oliver 2026-08-11) | fable660k_cc_harness; skeptic verdict 2026-06-12; KOs 07d4dfc0, dfcab5b5 | 2026-06-12 |
Two qa3 values at 10M appear because they come from different corpora: row 1 scores the leaderboard's official HuggingFace cells (86 at 10M, seed 42; 87.7 as the 3-seed figure); row 4 scores generator-built corpora item-paired to the 4K-1M ladder (80.0 at 10M). Same task, same official scorer, different corpora. Row 3 is the audit of what the task adaptation contributes and is the disclosure for every other BABILong row.
Numbers that appear on the three older pages (the August 5 technical notes, the partners benchmarks page, the data-room benchmarks page) and are not rows above. Each with its reason. Nothing is deleted from the archive; these are simply no longer cited.