Sapience Labs · evidence · ledger · last verified 2026-09-03

Benchmark evidence ledger

One row per claim. Each carries its number, its n and seeds, the caveats that travel with it, the result path, and the date it was last verified against that path. The short version is the evidence hub; this is the version for diligence.

Where an older page and this ledger disagree, the ledger wins and the caveat says "supersedes". Every BABILong row states that the retrieval pipeline is task-adapted, in the same sense that the fine-tuned specialists on that leaderboard are task-trained; row 3 is our own audit of how much that adaptation contributes. Numbers that do not appear here are listed at the end with the reason: superseded, retired, or retracted.

The ledger

#BenchmarkClaimNumbern, seedsCaveatsResult pathVerified
1BABILong qa3 ladder, official HF cellsMulti-hop accuracy holds from 128K to 10M tokens on the leaderboard's official HuggingFace cells86 / 87 / 87 / 84 / 86 at 128K / 256K / 512K / 1M / 10M, official scorern=100 per cell; seed 42 canonical (per-task predictions in PR #19); seeds 123 and 456 supplementaryTask-adapted retrieval pipeline (row 3 audits the contribution). Supersedes the August 5 page's 84.0 at 1M under the harmonized protocol and the data room's 86.0 at 128K: same cells, now scored by the official scorer.research/olympiad/leaderboard_submission_2026_08_25/ (official_metric_report.json; seeds_supplementary/; PR #19)2026-08-25
2BABILong avg(qa1-qa5) at 10M#1 overall on BABILong at 10M tokens97.2 official scorer (97.5 = 3-seed judge figure); best published 76.6 (ARMT, fine-tuned on these tasks)n=100 per task per seed; 3 seeds, zero variance (deterministic), so effective n=100 per taskScoped AT 10M, never "every length". qa1, qa4 and qa5 are extraction-solvable (naive solvers reach 100 without a language model); the reasoning load sits in qa2 (34 to 100 over the naive floor) and qa3 (32 to 87.7, 3-seed figure; the seed-42 official cell is 86, row 1). 97.2 is the headline on every surface; 97.5 appears only labelled as the judge figure. Task-adapted pipeline (row 3).research/olympiad/leaderboard_submission_2026_08_25/ (official_metric_report.json; SKEPTIC_VERDICT.md)2026-08-25
3BABILong zero-adaptation audit at 10MWhat the task adaptation contributes: the shipped generic pipeline, zero task adaptation, same tasks, same scaleqa1 4.7 / qa2 4.3 / qa3 6.3 / qa4 45.7 / qa5 62.7 (24.7 average)pooled 3 seeds (42 / 123 / 456), n=300 per task; $3.58 totalqa1-qa3 sit below the 16.7% guess floor via abstention. Retrieval-grounded qa1 and qa2 = 0/300: the bAbI template grammar is invisible to semantic similarity. qa5 carries ~10.7pp of example-echo luck. The shipped LLM-lasso ingest was not exercised (verbatim-store, benchmark-standard ingest). This row is the disclosure that travels with rows 1, 2, 4, 8 and 14: those pipelines are task-adapted; this one is not.research/olympiad/_local_runs/avg10m_ext_2026_08_25/zero_adaptation/ (zero_adaptation_summary.json); prereg docs/babilong-zero-adaptation-prereg-2026-08-25.md; audit post2026-08-26
4BABILong qa3, 30M and 100M stretchFlat from 2M to 100M tokens on qa3 (multi-hop)2M 81.0 / 5M 80.0 / 10M 80.0 / 30M 80.0 [75.1, 84.1] / 100M 79.7 [74.8, 83.8], official scorer3 seeds, n=300 per scale; ~685 reader tokens per question at every scale; ~$3 totalGenerator-built corpora, item-paired to the 4K-1M ladder, PG-19 train-pool filler (seeded). Near-deterministic, so 3 seeds are not 900 independent items. Task-adapted pipeline (row 3). Two qa3-at-10M values exist on purpose: 86 in row 1 (official HF cells) and 80.0 here (paired generator corpora); same task, same scorer, different corpora. Firsts: first published multi-hop (qa3) results beyond 10M (ARMT published qa1 at 50M, 2024); first BABILong/RULER-class reasoning measurements at 100M (Magic.dev HashHop, in-house 2024, unreleased model). Supersedes the August 5 page's judge-scored "~82% at 2M, 5M and 10M".harness research/olympiad/harness/run_babilong_qa3_100m_stretch_2026_08_25.py; run family /Volumes/X10 Pro/qa3-100m-stretch-2026-08-25/ (Labs); figure figures/fig-babilong-100m-ladder-2026-08-27.png2026-08-27
5RULER variable tracking, 1M to 10MState tracking holds from 1M to 10M tokens24/24 at 1M; 16/16 at 2M, 5M and 10M; exact match2 seeds; corpora sha-pinned~260 tokens read per query; the 2M-10M cells cost $0.34 total. The structural chain-walk does the work; the claim is scoped to that mechanism. No frontier comparison number on this row: the June frontier cells are retired pending a clean re-run, and a SOTA claim waits for publication with the full protocol and harness. Supersedes the August 5 page's 1M-only framing (n=24, 3 seeds) and its frontier comparison.research/ruler_fullsuite/ (runs/phase1/sap_ladder/; synth_vt_ladder_2026_08_21.py)2026-08-24
6MuSiQue, beyond the windowIterative retrieval over the store holds where a truncated full-context reader collapses3-seed ladder 43 to 28 EM (Sapience, iterative) vs 19 to 0 (truncated full-context reader)n=30 items; 3 seeds; one reader (deepseek-v3.2); official exact match, no LLM judgeA gentle decline, not a plateau. FiD / cross-encoder steelman (2026-08-09): ties in-window (47.8 vs 44.4 pooled EM, not significant), never exceeds iterative, and retrieves MORE gold (support-recall 0.94 / 0.92 / 0.90 vs 0.76 / 0.67 / 0.63), so the residual is reader composition, not retrieval. 5-hop is unsolved. Supersedes the August 5 page's "+33 to 37pp" over-window framing.research/musique_beyond_window_repro/runs_fid/ (results_fid.jsonl; analysis_fid.json)2026-08-24
7Supersession (currency)When facts change, the old version cannot outrank the new98% current-answer accuracy on 108 update probes (2 misses), +70.4 pooled over similarity retrieval; structure-aware vs similarity retrieval +67 to +78 points on date-obscured chains3 seeds; discordant-pair McNemarHand-ordering steelman recovers parity, which isolates write order as the signal similarity retrieval discards. The +67 to +78 is a range of per-cell deltas, not a confidence interval (skeptic correction R-1 on the 2026-08-12 clean re-run: binding-cell range +66.7 to +74.1, pooled +70.4, Newcombe 95% [+59.3, +78.4]). The 0/81 figure is the WEIGHT-LAYER count (L3 replay with the supersession filter: 0 of 81 stale facts persist vs 14 of 81 unfiltered; KO e2ee87cc, skeptic verdict docs/akshat-three-curricula-skeptic-verdict-2026-08-27.md) and is not a retrieval number. Real-Wikipedia currency numbers on the August 5 page (+40.4pp; 100 / 98 / 59.6) are not carried onto this ledger.research/memory-experiments/l2_synthesis/runs/currency_2x2_2026_08_12/ (docs/currency-clean-rerun-skeptic-verdict-2026-08-12.md); docs/akshat-three-curricula-skeptic-verdict-2026-08-27.md (L3 replay, B_super arm)2026-08-24
8HippoRAG-2 at 1MThe strongest published structured-memory baseline, on identical items15.5 (HippoRAG-2) vs 84.0 (Sapience) on the identical 100 BABILong qa3 items at 1M; ~2,500x our token footprintn=100 itemsFraming: HippoRAG-2 "does not transfer to this regime", never "weak". Our side is the task-adapted pipeline (row 3).docs/hipporag-baseline-final-2026-08-06.md (run: baseline-arm/_runs/hipporag_qa3_1M_2026_08_02/, report_hipporag2.json)2026-08-24
9Published nullWithin a single window, single shot, static corpus: a tie+0.00pp, p=1.0, vs a frozen frontier model with strong retrievalpre-registered; losing condition published before the runThe one regime where structure has nothing to add; reported as it landed.pre-registration and post-mortem in the data room (on request)2026-08-24
10Repo-Evolution (coding currency)What is current in an evolving codebase88% (frontier reader) / 82% (8B reader) with Sapience vs Mem0 51%, grep 29%, notes file 25%, flat RAG 18%same model, same questions; gold derived mechanically from git, no LLM judgeSkeptic-audited 2026-07-30; live on the coding evidence page. Supersedes the August 5 page's "+30 to 65pp" and "26 of 27 identical" framing of the same result.research/repo_evolution/; coding evidence page2026-08-27
11Repo-Task pilotTask-level coding pilot: two observations, no result yet11.4K vs 97.5K tokens; 133K vs >300K (censored)two single observations (n=2); split signNO completion rates, NO effect sizes. Agents consulted memory voluntarily in ~0-8% of episodes. A powered run is in progress; nothing on this row is a result.research/repo_task/ (pilots A4 / A5)2026-08-27
12Weight path (lab)Retention on the weight pathadapter-trained retention 47 to 75 points above the no-adapter control; matched-or-better acquisitionas in the figure of recordNOT shipped: what ships is the store, retrieval and consolidation over a frozen model. The "50 to 75" phrasing is retired. The +6.01pp weight-consolidation figure is a different measurement (capability uplift on held-out tasks, row 17) and is never summed with or compared to this retention delta.papers/sapience-frontier/figures/fig_l3_retention_2026-08-09.py (.png)2026-08-24
13Query latency at 1MPer-query latency stays ~1.4s as stored history grows~1.4s per query vs ~10s frontier-alone, warm cache, at 1M (~7x; ~20x cold-prefill as a parenthetical only)n=8 per cellOlder, single study (2026-06-14). Stated as "~7x", never rounded up.KO b51b4646 (skeptic-approved 2026-06-14)2026-06-14 (older, single study)
14Matched pairs, 1M and 10MSame frozen model, with and without Sapience1M pair (BABILong qa3): 34.4% alone vs 66.7% with Sapience (+32, McNemar p=1.5e-5); 10M pair: 45.0% non-abstention with Sapience (27 of 60, 3 seeds; 44.4% pooled) vs 0 of 20 for the same model given its full 1M window (12.5% pooled, the correct answers being abstentions); 0% where the evidence lies beyond the readable window (window arithmetic)1M: n=90 item-paired, both arms Sonnet 4.6; 10M: internal 10M corpus, held-in cells, matched protocol; non-abstention n=60 (n=72 pooled), 3 seeds; the full-1M-window run is one deterministic seedThe 10M caption reads "internal 10M corpus, held-in cells, matched protocol"; it never carries a benchmark name (Ruling A). Approved runs 2026-06-19 (close-out) and 2026-06-21 (full window). Attestation: check_calibration_leakage.py --source beam10m exit 0 (run 2026-08-27 on studio-1 over 113 jsonl files under research/olympiad/_local_runs/beam_forward_growth_2026_06_16/, incl. held-in seeds 42/123/456). Task-adapted pipeline on the 1M pair (row 3). Supersedes the data room's +37.8pp (n=45) and the August 5 page's +32.2pp framing of the 1M pair.research/repo_evolution/fig_public_pairs_2026-07-30.png; 1M pooled analysis research/olympiad/_local_runs/qa3_1m_stage2_pooled_analysis_2026_08_03.json2026-07-30 (figure); attestation pending
15Token economicsReads scale with the query, not the historya few hundred tokens per answer instead of hundreds of thousands; ~1,800x cheaper per query on the licensed figuremoney-plot figure of recordLicensed shape from the features canon. The ~961x lookup-probe figure is a different measurement (store-growth crossover, row 18) and coexists with this row under its own scope; the other older-page variants (~304,000 tokens re-sent past the window; $0.007 vs $13 per query at 1M) are not carried forward; the money plot is the single figure of record for the per-query ratio.figures/fig-scale-v5.1-frontier-to-10M-2026-08-28.png (v4 superseded)2026-08-27 (canon)
16Modern RAG, run by us, on the same 10M cellsThe strongest retrieval baseline a critic would build, on the leaderboard's official qa3 cells at 10Mhybrid BM25 + dense + reciprocal-rank fusion (cross-encoder variant) at a 2-4K-token budget: 11.3 to 12.3% on qa3 at 10M, official scorer (9.7 to 11.0 strict judge); at Sapience's own token budget, 3%; reads 4.9x the tokens Sapience reads per questionsame official HF 10M cells as row 1; $0.021Pre-registered with its kill condition. About the old leaderboard RAG entry (11). The Sapience side of the comparison is row 1 (task-adapted pipeline, row 3). Skeptic verdict KO 87c5e4fa (approved 2026-08-27): "3 reader replicates", never "3 independent seeds"; strict-judge values quoted alongside official; achieved budgets ~3.4K of 4K disclosed; scope retrieve-then-read at 4K or less (FiD-fusion at 10M and a wide top-150 arm are future work).research/qa3_steelman_rag_2026_08_25/ (run family /Volumes/X10 Pro/qa3_steelman_rag/); prereg docs/qa3-10m-steelman-rag-prereg-2026-08-25.md2026-08-27
17Consolidation into weights (DREAMS)Weight consolidation adds capability on held-out tasks+6.01pp held-out accuracy from training-time consolidation (LoRA, into weights); 95% CI +3.16 to +8.8710 of 10 seeds; replicated on 5 fresh seedsDREAMS +6.01pp = capability uplift on held-out tasks (what weight consolidation adds); L3 47 to 75 points (row 12) = retention vs sequential SFT (what replay prevents losing). Different measurements, never summed or compared. Within-domain replay in DREAMS hurts (-6.7 to -8.9pp) while L3 replays same-domain content to protect it; the uplift-vs-retention framing resolves the apparent contradiction. Identical inference in all arms, so the lift is not test-time compute. Lab result, not shipped.Discovery by Dreaming, arXiv 2607.16256; consolidation ruling KO 897fe616 (2026-09-02)2026-09-02
18Store-growth crossover (lookup probes)Retrieval over a persistent store holds where dumping the store into context collapsesstore-in-context 78.3 / 41.7 / 13.3 as the store grows past the window vs 100% [94, 100] retrieving from the store; ~961x fewer tokens per answer at the largest store3 seeds; lookup probesLookup probes only; the structure-specific gains are a separate study. A different measurement from row 15's ~1,800x per query at 1M (qa3 payload vs raw window); the two coexist under different scopes and are never combined.KO 45fce2e5 (skeptic NUMBERS APPROVED; re-passed 2026-09-02, verdict item 19)2026-09-02
19NoCha-style, real novels beyond the windowClaim verification over public-domain novels once the book exceeds the reader budget38.6% structured retrieval vs 20.1% positional truncation, pair accuracy; item-level +18.5pp, cluster-bootstrap 95% CI [6.9, 29.6], paired-t p=0.00183 seeds are retrieval-order reshuffles over the same 63 book-pairs, so effective n=63 (never 189); same reader (gemini-2.5-flash) and the same 24K budget in both arms"NoCha-style": public-domain Gutenberg classics that sit in pretraining, not the real 2023+ NoCha set. Beyond-window regime only (24K reader budget vs 67K to 251K-token books); no cross-regime pooling. Single-claim TRUE/FALSE verification scored as pair accuracy. Scope is retrieval, not multi-hop or consolidation reasoning. The pooled-189 McNemar p and the single-seed 46% are never cited.research/olympiad/results/nocha_beyond_window_2026_06_27/; docs/nocha-panel-2026-06-26.md ("MULTI-SEED COMPLETE"); skeptic KO eb9fcdcc (2026-07-03, deck-eligible, scoped); locked aggregation KO 1e6c892b (2026-06-30)2026-09-02 (consolidation ruling)
MethodologyThe official BABILong scorer penalizes chain-of-thought preamblesOfficial first-sentence substring metric scores CoT preambles as wrong: the language model alone under load emits multi-sentence outputs 0/24/65/43% of the time at 4K/64K/128K/1M, so its official score reads 74.0/45.5/18.5/26.7-31.1 vs dual-judge 74.0/55.9/60.0/31.1-34.4; every disagreement is a CoT cutoff, none in the other direction. Sapience is identical under both scorers (83/84/86 vs 84/84/86)re-score of persisted predictions, 2026-08-27Consequence: cortex-only arms are reported under the dual judge in any figure that includes them; official-scorer comparisons that would flatter our arm are not used. Gemini "16.7% n=6" at 1M retired (3 items attempted twice; no clean multi-vendor 1M cell exists)research/olympiad/_local_runs/official_rescore_cortex_2026_08_27/rescore_summary.json; KOs 5ce6d3f1, 57cef0d42026-08-27
BEAM-public (10M tier, full 200)Frozen-config baseline, published either way0.401 mean graded score (0.365 on the clean held-out subset) vs Hindsight re-judged under the same pinned judge: 0.677 (0.662 clean)n=595/600 scored, 3 seeds (0.392 / 0.410 / 0.400), bootstrap CI about +/-0.06Dev conversations 1 and 10 sit inside the headline; matched pinned judge gemini-3.5-flash; 5 content-correlated wedge rows excluded; Hindsight's own leaderboard figures (64.1 / 73.5) are under a different judge and are not the comparison; the query-time v2 arm and later versions are reported as they land. Task-adapted note does not apply (this is the shipped frozen configuration)studio-1:/Users/oz/X10Pro/beam-public/runs_full200/seed{42,123,456}/; re-judge research/beam-public/runs/rejudge_hindsight/rows.jsonl (0.6769 recompute of record); KO e8c629b82026-08-27
Consolidation tier (write-time aggregation gate)Exact aggregates materialized at write time beat a compute-matched counter+22.8pp (95% CI +8.1 to +38.2) over a compute-matched language-model counter emitting the identical sentence, on covered aggregation questionsPaper v4.3, result (3); single reader at a single scale; synthetic instrumentThe same contrast runs -3.3pp on trend questions and is exactly neutral on single-episode controls; broader semantic and schema consolidation is implemented but unestablished; the earlier 26.7 to 66.7 synthesis figure is retired (see below) and is never citedSapience frontier paper v4.3, Zenodo 21937795; v4.1 precursor +11.1pp with skeptic verdict docs/l2-functional-skeptic-verdict-2026-08-10.md; KO eff68f7f2026-08-14
Discovery channel: source lead timeCross-domain sources sit surfaceable for years before the field imports them29 of 155 realized imports had their true source in a skimmable top-40 slate more than a year early; median lead time 12.5 years across those 29. In the per-case lag study, a median 72% of the delay from problem to solution was reducible in principle155 realized imports, abstract-level ranks (a floor); lag study per-case median reducible fraction 0.72, n=48Floor framing: these ranks use abstract-level distillation, so a full-text pass can only add targets. The famous cases (Boole to Shannon, 83 years) are documented history, separate from the measured set. Within a window, dense retrieval out-recovers this channel; its value is the stratum retrieval cannot enter. Targets are a held-out future, so a solve cannot be memorizationfigures/fig-reachback-2026-08-09.py over lt_leadtime.json; lag study KO eabcb6e2 (deploy #50, skeptic-approved); foresight slate on the archived research overview2026-08-09
Frontier reference cell: Fable 5 alone at about 660K (BABILong qa3)A frontier model alone, at its effective window, on the same task64.0% (32 of 50), Wilson 50.1 to 75.9n=50 of a planned 60 (one stopped shard, low bias risk); single pass; dual judge unanimousClaude Code harness transport is part of the cell; a baseline reference star, not an architecture headline; this stays the paper's Fable anchor and no true-1M Fable cell exists (Oliver 2026-08-11)fable660k_cc_harness; skeptic verdict 2026-06-12; KOs 07d4dfc0, dfcab5b52026-06-12

Two qa3 values at 10M appear because they come from different corpora: row 1 scores the leaderboard's official HuggingFace cells (86 at 10M, seed 42; 87.7 as the 3-seed figure); row 4 scores generator-built corpora item-paired to the 4K-1M ladder (80.0 at 10M). Same task, same official scorer, different corpora. Row 3 is the audit of what the task adaptation contributes and is the disclosure for every other BABILong row.

Retired or not carried forward

Numbers that appear on the three older pages (the August 5 technical notes, the partners benchmarks page, the data-room benchmarks page) and are not rows above. Each with its reason. Nothing is deleted from the archive; these are simply no longer cited.