Sapience Labs · evidence · last verified 2026-09-03

Results and their provenance

Every number on this page links to the artifact it comes from, with the date it was last checked against it. The protocols were pre-registered.

What we claim

Every number, with its n, seeds, caveats and result path, is one row of the benchmark ledger. Nothing on this page appears without a row there.

How it learns

How the system learns: at capture, overnight, and into the weights in the lab; the model underneath stays frozen and swappable
The system learns at three timescales. At capture, the write gate decides what is worth keeping and each claim keeps its reasoning. Overnight, consolidation compacts the day into structure and writes it back to the store; the morning briefing surfaces what changed. In the lab, replay into adapters carries the store into the weights with superseded facts excluded; that path is a lab result, not what ships. The model underneath stays frozen and swappable.

At capture, work is distilled into knowledge objects: claims with their reasoning, evidence and provenance. Encoding is gated by prediction error: what contradicts or extends the store writes strongly; the redundant fades. When a new claim lands on an entity that already carries one, the two are linked at write time and the old claim is marked superseded. At read time, retrieval walks the chain to its head, so a superseded value cannot outrank its replacement no matter how similar it is to the question. Offline, consolidation compacts specifics into structure: fast capture, slow integration.

What one query touches
How it reads: what one query touches is a few hundred tokens of evidence, not the history.

Which number is which

Six BABILong figures appear across our pages. All six are real and scoped; this is the scope of each, and the one that is the headline.

NumberWhat it measures, and its scopeLedgerRole
97.2BABILong, average of qa1 to qa5 at 10M tokens, official scorer (97.5 is the 3-seed judge figure); best published 76.6. At 10M only, never "every length"; task-adapted pipeline, with the audit published beside it.row 2the headline
24.7Zero-adaptation audit: the shipped generic pipeline, no task adaptation, average at 10M (qa1 4.7 / qa2 4.3 / qa3 6.3 / qa4 45.7 / qa5 62.7); pooled 3 seeds, n=300 per task. What the task adaptation contributes.row 3published audit
80.0qa3 (multi-hop) flat from 2M to 100M: 81.0 / 80.0 / 80.0 / 80.0 / 79.7, official scorer, 3 seeds, n=300 per scale, about 685 tokens read per question.row 4the scale claim
87.0qa3 ladder mid-cells on the official HuggingFace cells: 86 / 87 / 87 / 84 / 86 at 128K / 256K / 512K / 1M / 10M; n=100 per cell, seed 42 canonical, single seed.row 1ledger only
66.7 vs 34.4Matched-reader pair at 1M (qa3): the same frozen Sonnet 4.6 with and without Sapience, +32 points, n=90 item-paired, McNemar p=1.5e-5.row 14the 1M pair
68.9 vs 31.1An earlier n=45 half of that same pair (+37.8), superseded by the n=90 pooled figure.retired listarchive only

Two qa3 values at 10M exist on purpose: 86 in row 1 (official HuggingFace cells) and 80.0 in row 4 (generator-built corpora item-paired to the 4K-1M ladder); same task, same official scorer, different corpora. Every BABILong number here comes from a task-adapted retrieval pipeline, in the same sense that the fine-tuned entries on that leaderboard are task-trained; the 24.7 audit measures how much that adaptation contributes.

The figure

Split-axis: the language model alone collapses past 128K; with Sapience flat to 100M
BABILong qa3, 4K to 100M. Left panel 4K-1M under the dual judge (all series); right panel 2M-100M under the official scorer. The language model alone falls past 128K; with Sapience the line holds near 80 out to 100M. The cortex-only arms are shown under the dual judge because the official first-sentence metric scores chain-of-thought preambles as wrong.

Three documents

What it is for

We build the third component of the AI stack - beside the model and the harness, with its own state and its own learning rules; the model stays swappable, and the system learns while it works. AI today generates. What it cannot yet do is connect. Sapience is the connective layer: one living memory that connects your sessions to each other, your AIs to each other, you to your AIs - and, with consent, you to other people.

Most network products are worth nothing until the network shows up. Sapience is at its most useful on day one, alone - the memory and long-context reasoning now carry public leaderboard receipts. The network isn't the price of admission; it's the compounding on top.

And the network starts at two, not at a million: on day one you connect your Sapience with your coworkers, your partner, your friends - one sentence, and one yes - so the first network you need is one you already have.

And this isn't several products stapled together. The same knowledge object that gives one person's AI memory - a claim with its evidence, its reasoning, and a name attached - is exactly the object that lets two people's AIs connect. Matching, consent, and attribution are properties of the substrate, not features bolted beside it. Build the memory and the network is latent in it.

Ultimately, the network is the thing we are building.

Papers and archive

The architecture paper (Zenodo DOI) is linked from spnc.ai; the discovery paper for the science-of-science line is at partners.coretx.ai/public/papers-disc-stack. Older one-pagers, overviews and benchmark pages are kept in place and stamped with what superseded them: archive.