A transformer is one of the memory systems a brain has, scaled on its own. It is sharp up to its training cutoff and then it stops, and every session starts from zero. Sapience builds the missing systems around it, following complementary learning systems theory: a fast store for what just happened, and a slow process that turns it into what you know.
It remembers at scale because the model never reads the history. It reads a few hundred tokens selected from the record, so what it reads stays about the same size whether the record holds a week or a career. That is why accuracy and cost stay flat to a billion tokens (1), and why the same model scores 86 through Sapience against 20 on its own (3).
It stays current because the record tracks what replaced what, so a question about what a setting is now is answered from the record rather than guessed (4).
In the lab, what the record holds can also reach the weights of an open model: a model trained from its record keeps a first task while learning a second (5), and answers with the current value after a fact changes (6). That is a lab result, not what ships.
Still open: broader consolidation into general knowledge is built and not yet established as a measured gain, and we trail on single-document question answering, where the prompt already holds every fact.