More accurate, with less memory
Selected published results on BEAM 100K. Assertion and Hindsight are scored the same way; the gray rows used other answer models or judges, so read them as context.
Memory given to the model per question
lower is betterMemory is the text given to the model for each question, counted with the same tokenizer for both systems. Assertion gave the model less memory in 19 of the 20 conversations.
Why this result matters
Ahead of a leading memory system, on its own benchmark.
Hindsight is the leading memory system that publishes its BEAM answers, and calls itself #1 on the Agent Memory Benchmark, which its makers built. Run through that same harness, with the same answer model, judge and rubric, Assertion scores 81.1% against 77.9%, and each of its 3 runs scores above each of Hindsight’s 3.
Better memory, not more of it.
Assertion gives the model 12.2k tokens of memory per question, against Hindsight’s 17.7k: about a tenth of the conversation. The accuracy comes from deciding what is true while remembering, not from handing the model more text. Less to send means faster, cheaper calls and more of the context window left for the work.
Summarization
Ask what happened across weeks of sessions and you need one coherent account, not a pile of fragments. It is what an agent needs to pick a project back up after a break.
Information extraction
The version you pinned, the threshold you agreed, the reason an approach was ruled out: the specifics that stop an agent re-asking or re-deciding.
Contradiction resolution
When two statements can’t both be true, Assertion keeps both and flags the conflict instead of quietly picking one, so the agent doesn’t act on a decision you already reversed.
These are the questions developers put to their coding agents every day: what did we decide, why, and what changed since. That is the memory Assertion is built to keep.
What BEAM tests
BEAM is an academic benchmark for long-term memory (arXiv 2510.27246). Each of its 20 conversations runs over many sessions, about 130,000 tokens on average, from coding projects to estate planning. Afterwards the model answers questions only a good memory can: what was decided, what changed, what was said twice and contradicted, what happened in what order. The model never sees the conversation itself, only what the memory kept.
Why BEAM 100K first
BEAM comes in four sizes: 100K, 500K, 1M and 10M tokens per conversation. We started with 100K for three reasons. It is the size we could run completely: all 20 conversations and all 400 questions, several times over, rather than a sample. Hindsight published its memory for every 100K question, so the comparison could be like for like, with the same answer model and judge. And it is where memory has to earn its place: these conversations still fit in a large model’s context window, so memory is only worth having if it is accurate and much smaller. BEAM 1M, with conversations ten times longer and beyond most models’ context windows, is next.
Strengths and gaps
Accuracy by the ten abilities BEAM tests, 40 questions each. Assertion Hindsight
Assertion settles what is true when it remembers, instead of storing every statement and leaving the model to sort it out. When a decision changes, the new answer becomes current and the old one is kept as superseded, so the model reads one consistent picture. That shows most in summarization and information extraction. It is weaker today on knowledge update, the latest value after something changed, which is what we are improving next.
Why some scores are low for everyone
- Some questions can’t be fully earned. When we audited 60 answers Assertion lost, 17 traced to the benchmark’s expected answer or the judge rather than to memory: an expected answer that cites an outdated date, two compatible statements scored as a contradiction, or the assistant’s suggestion counted as the user’s choice.
- The rubric rewards habits, not just memory. A contradiction resolution answer only gets full credit if it asks which statement is correct, whatever the memory supplied.
- The judge varies. Identical answers are sometimes scored differently between runs, which is why scores are averaged over repeated runs.
- Some questions grade by checklist. Multi-session reasoning asks for several items gathered from many sessions, and an answer missing any one loses part of its score; a few of these expected answers are also defective. It scores 56.9% for Assertion and 53.9% for Hindsight.
These affect both systems equally: the same questions, the same rubric and the same judge.
How it compares
Same harness, same answer model, same judge. Only the memory differs.
| Assertion | Hindsight | |
|---|---|---|
| Accuracy, BEAM 100K (400 questions) | 81.1% | 77.9% |
| Memory given to the model per question | 12.2k tokens | 17.7k tokens |
| Answer model | Gemini 3.1 Pro | Gemini 3.1 Pro |
| Judge | Gemini 3.5 Flash | Gemini 3.5 Flash |
We compare against Hindsight because it is the leading memory system that publishes its BEAM answers, so it can be scored exactly as Assertion is. Its figure here is its own published memory for every question, answered by the same model and scored by the same judge as Assertion’s, averaged over 3 runs; Assertion’s is the average of 3 runs. Re-scoring Hindsight’s own published answers instead gives 78.8%. Hindsight originally reported 73.4%, scored by a judge Google has since retired. Memory sizes are counted with the same tokenizer for both, and match Hindsight’s own reported counts.
How we measured
- Harness. The Agent Memory Benchmark’s BEAM loader, answer prompt and judging rubric, unchanged, pinned to commit 5d5e8dbe. AMB is built by Vectorize, the makers of Hindsight.
- Memory. Every exchange of every conversation, in order, through Assertion’s memory pipeline. For each question, the answer model receives that memory, never the conversation itself.
- Scoring. Both systems are answered by the same model and judged by the same model with the same rubric. Assertion’s score is the average of 3 runs; Hindsight’s is the average of 3 runs (77.1% to 78.9%).
- Development. Like other systems on this benchmark, Assertion was developed with BEAM in view.
- What is public. The Assertion plugin is open source (Apache-2.0); the memory engine runs as a service. We publish every answer, the judge’s scores and the exact memory each answer was based on, so anyone can rerun the answering and the scoring and check that the memory holds compact facts, not the conversation. Results for each of the 20 conversations are there too: github.com/Assertion-AI/benchmarks.
- Next. BEAM 1M, with conversations ten times longer, is in progress and will be added to this page.