SapienceLabs

Research

A brain-inspired AI architecture that learns with you.

Sapience sits behind the AI you already use and keeps learning from your work. What it does for you, and how it measures against the strongest models, with the figures.

What it does for you

One Sapience, across every AI you use.

It learns the way a brain does: what you establish today is still there next month, current, and in every tool.

  • Works today

    Learns with you, across every AI you use.

    Claude, ChatGPT, Cursor or Codex: your sessions share the same Sapience. What you worked out in one is there in the next, and the answer reflects what is current, not what used to be.

  • Works today

    Stays yours. It never trains anyone else's model.

    Your record is kept separately from the model, so the model underneath is a part you can swap. Nothing leaves by default, and nothing of yours trains a foreign model.

  • Works today

    Share with the people you choose, with your name on it.

    Share a slice of your work in a sentence and a colleague's AI answers with it, credited to you. What you granted, you can withdraw.

  • Starting now

    Introduces you to people whose work fits yours.

    Somewhere in a lab you have never heard of, someone is working on the other half of your problem. Sapience is beginning to find those pairings across fields, with both names on the connection.

Measured against the frontier

The results, with figures.

The same model with and without Sapience, and Sapience against the strongest published models. Lab results are marked as such.

1 · Remembers at scale

Accuracy against stored history, 4K to one billion tokens, against frontier models.

BABILong multi-hop accuracy against tokens of stored history, 4K to one billion, log axis: Sapience holds between 79.7 and 87 at every length, 82 at one billion; the models on their own, reading the whole history, fall away as it grows; a model with a 1M window at chance, 14.3, on the billion-token items.

Sapience holds between 79.7 and 87 at every length from 4K, 82 at one billion; the models on their own, reading the whole history, fall away as it grows.

BABILong qa3. The official cells run to 10M; 30M to 1B use our corpora with repeating PG-19 book filler. The model reads a few hundred tokens per answer.

n=100 items per cell, three replicate runs; 95 percent interval 74.3 to 89.0 at one billion. Official BABILong cells to 10M; our corpora beyond, about 6 points under the official cells at equal length.

2 · OpenAI's own long-context test

OpenAI MRCR v2, 8-needle, 512K to 1M tokens.

OpenAI MRCR v2, 8-needle items of 512K to 1M tokens: Sapience 98.4 against vendor-published values, with reader input cost per answer beside it

Sapience ties GPT-6 Astra, 98.4 against 96.3, reading about fifty times fewer tokens per answer.

OpenAI's own grader. Astra's number sits inside our interval, so the two tie and no ordering is claimed; the other bars are vendor-published values, not our runs.

Three runs of 30 items, 97.7 / 99.9 / 97.7; 95 percent interval 95.4 to 99.95. OpenAI's grader.

3 · Works across models

The same model, with and without Sapience.

The same model with and without Sapience: 86 against 20 on BABILong at 128K, 66.7 against 34.4 at 1M, and a tie on RULER retrieval reading a small fraction of the tokens

The same model scores 86 with Sapience against 20 without it on BABILong at 128K, and 66.7 against 34.4 at 1M; nothing is fine-tuned and no weights change.

The same model on both sides, the same 100 items; judge-scored 86 against 65; the benchmark's first-sentence rule gives 20. The record stays yours, and the model underneath is a part you can swap.

4 · Stays current

What is current in an evolving codebase.

Current-value questions answered correctly, on the same 60 questions: Sapience 88, Sapience with an 8B reader 83, Mem0 48, the same model with git grep 27, a notes file 20, flat RAG 17. The 8B bar reads with Llama-3.1-8B; all other bars read with deepseek-v3.2

Asked what a setting is now in a codebase that keeps changing, Sapience answers 88 percent right on the shipped engine, against 48 for Mem0 and 27 for the same model with git grep, on the same 60 questions.

Real git histories of Django, LiteLLM, huggingface_hub and vllm, graded against the git history, no AI judge; 8B reader 83. The 8B bar reads with Llama-3.1-8B; all other bars read with deepseek-v3.2.

5 · Learns without forgetting · lab result

Keeping the first task while learning a second, 7B to 397B.

Forgetting the first task after learning a second across seven models from 7B to a 397B mixture of experts: sequential fine-tuning 50 to 94 percent, with Sapience 4 to 18 percent

Sapience forgets 4 to 18 percent of a first task instead of 50 to 94, from 7B to a 397B model, and learns the new task as well or better.

Synthetic two-task study on open models. Three seeds a rung; the 397B is a mixture of experts, 17B active; the second task learned as well or better in all 21 seed cells. A laboratory result, not a feature shipped in the beta, and not a guarantee of perfect recall.

6 · When a fact changes · lab result

Answering with the current value after a fact changed.

When a fact changes, current answers after training on two disjoint instruments: two controls at 58 and 59 and at 40 and 33 percent, a model trained from the record at 95 and 82 percent

When a fact changes, a model trained from the record answers the current value on 95 percent of probes on the first instrument and 82 percent on a second one built from five different codebases, against 58 and 40 for a control.

One 8B base, three seeds per instrument, item-clustered intervals; the second instrument is built from five disjoint codebases. A laboratory result.

7 · Reach across fields · backtest

Discoveries whose answer was already published in another field.

Each bar is a real discovery reaching back on a calendar axis to where its answer was already published in another field

Each bar is a real discovery reaching back to where its answer had already been published in another field.

n=48 measured cases; the famous cases (Boole to circuit design, 83 years; Radon to computed tomography) are documented history shown as context, not part of the measured set.

8 · A connection that took eight years

Glaciologists needed a method engineering already had.

engineering distributed fibre-optic sensing glaciologists a Greenland glacier I II III IV V VI VII VIII eight years earlier, with no citation path between the fields credited to both 1 10 100 1,000 10,000 rank at which the method was put forward Sapience: among the first hundred it put forward by similarity: 22,746th
engineering fibre-optic sensing glaciologists a Greenland glacier eight years, no citation path credited to both 1 10 100 1,000 10,000 rank at which the method was put forward Sapience: among the first hundred it put forward by similarity: 22,746th

Engineering had the method, distributed fibre-optic sensing, eight years earlier; by similarity it ranked 22,746th, and Sapience re-found the pairing among the first hundred it put forward.

The glaciologists made the connection by hand in the end; this is a connection recovered, not a discovery claimed.

9 · Cost

Accuracy and cost stay flat as your history grows.

BABILong qa3 accuracy against tokens of stored history, 4K to 100M on one log axis: language models alone fall away; Sapience holds a flat band

Roughly 1,800x fewer tokens per query once your accumulated work passes a million tokens, on our benchmark corpus.

Language models alone fall away as the stored history grows; Sapience holds a flat band from 4K to 10M tokens on BABILong, and on generated corpora out to 100M. On our benchmark corpus; a read scales with the question, not the history.

We trail on single-document question answering, where every fact is already in the prompt and memory has nothing to add.

The theory

How we think about it.

A transformer is one of the memory systems a brain has, scaled on its own. It is sharp up to its training cutoff and then it stops, and every session starts from zero. Sapience builds the missing systems around it, following complementary learning systems theory: a fast store for what just happened, and a slow process that turns it into what you know.

It remembers at scale because the model never reads the history. It reads a few hundred tokens selected from the record, so what it reads stays about the same size whether the record holds a week or a career. That is why accuracy and cost stay flat to a billion tokens (1), and why the same model scores 86 through Sapience against 20 on its own (3).

It stays current because the record tracks what replaced what, so a question about what a setting is now is answered from the record rather than guessed (4).

In the lab, what the record holds can also reach the weights of an open model: a model trained from its record keeps a first task while learning a second (5), and answers with the current value after a fact changes (6). That is a lab result, not what ships.

Still open: broader consolidation into general knowledge is built and not yet established as a measured gain, and we trail on single-document question answering, where the prompt already holds every fact.

Papers

Papers and protocols, on request.

Our papers, the full protocols and the ledger are available to partners on request. oliver@spnc.ai

Your AI, learning with you from the first session.