The first AI platform to lead both major memory benchmarks

93.4% on LongMemEval in an independent evaluation, and 90.0% on LoCoMo, ahead of every published memory system in every category. Both runs are public and reproducible.

93.4%

LongMemEval accuracy

467 of 500 questions correct

Independent evaluation by NewMathData

93.4%

LongMemEval accuracy

467 of 500 questions correct

Independent evaluation by NewMathData

90.0%

LoCoMo accuracy

Highest score in all four question categories

Human baseline ~88%, frontier models ~80%

+14.2 pts

Ahead of the next-best system

Memobase 75.78%, Zep 75.14% on LoCoMo

Mem0, LangMem and OpenAI memory trail further

LongMemEval: 93.4% on the original academic specification

500 questions across six question types, each with roughly 40 to 50 sessions of history. Evaluated independently by NewMathData and scored by four separate LLM judges, which agree within about two points overall.

Accuracy by question type
LongMemEval, 500 questions, scored by the GPT-4o-mini judge.
Overall93.4%
Single-session (assistant)98.2%
Single-session (user)97.1%
Knowledge update93.6%
Multi-session91.7%
Temporal reasoning91.7%
Single-session (preference)90.0%
Category
Questions
Correct
GPT-4o-mini
GPT-4
Gemini 3 Pro
GPT-5.2
Single-session (assistant)
56
55
98.2
96.4
96.4
96.4
Single-session (user)
70
68
97.1
95.7
95.7
95.7
Knowledge update
78
73
93.6
89.7
93.6
93.6
Multi-session
133
122
91.7
94.0
92.5
89.5
Temporal reasoning
133
122
91.7
93.2
92.5
92.5
Single-session (preference)
30
27
90.0
80.0
80.0
66.7
Overall
500
467
93.4
92.8
92.8
91.2
Accuracy (%) on LongMemEval s_cleaned, 500 questions. GPT-4o-mini is the primary judge; GPT-4, Gemini 3 Pro and GPT-5.2 scored the same responses. Backboard ran on GPT-4.1, memory set to Auto for ingestion and Readonly for queries.

LoCoMo: 90.0% overall, first in every category

Ten multi-session conversations ingested turn by turn into isolated assistants, then questioned across single-hop, multi-hop, open-domain and temporal reasoning. Backboard ran on Gemini 2.5 Pro; GPT-4.1 judged every answer.

Accuracy by system
LoCoMo, overall accuracy, GPT-4.1 judge. Hover a row for every category.
Backboard90.0%
Memobase (v0.0.37)75.8%
Zep75.1%
Memobase (v0.0.32)70.9%
Mem0-Graph68.4%
Mem066.9%
LangMem58.1%
OpenAI52.9%
Method
Single-hop
Multi-hop
Open domain
Temporal
Overall
1
Backboard
89.36
75.00
91.20
91.90
90.00
2
Memobase (v0.0.37)
70.92
46.88
77.17
85.05
75.78
-14.2 pts
3
Zep
74.11
66.04
67.71
79.79
75.14
-14.9 pts
4
Memobase (v0.0.32)
63.83
52.08
71.82
80.37
70.91
-19.1 pts
5
Mem0-Graph
65.71
47.19
75.71
58.13
68.44
-21.6 pts
6
Mem0
67.13
51.15
72.93
55.51
66.88
-23.1 pts
7
LangMem
62.23
47.92
71.12
23.43
58.10
-31.9 pts
8
OpenAI
63.79
42.92
62.29
21.71
52.90
-37.1 pts
Accuracy (%) on LoCoMo categories 1–4, adversarial questions excluded. GPT-4.1 judge at temperature 0.1. Points shown under each overall score are the gap to Backboard.
Breaking the 90-percent threshold requires superhuman consistency in recall and reasoning. Most high-performing frontier models currently score around 80 percent on LoCoMo. The system built by Backboard.io is a far better attempt at simulating memory as it manifests in humans.

Adyasha Maharana, creator of the LoCoMo benchmark and research scientist at Databricks, in the Ottawa Business Journal

What this means for agentic AI

Without reliable, shared memory, agents fragment, hallucinate, and reset. Backboard gives agents persistent, shared memory even when they run on different underlying models, so agentic behavior emerges naturally rather than being scripted.

Why leading both benchmarks matters

Most systems optimize for either short-horizon precision or long-horizon persistence, not both. Backboard's results reflect system-level behavior built on Active Temporal Resonance, which preserves meaning and continuity as interactions unfold, rather than benchmark-specific tuning.

Methodology: LongMemEval was evaluated independently by NewMathData on the benchmark's original academic specification (s_cleaned, 500 questions) with Backboard on GPT-4.1. LoCoMo used the standard dataset with adversarial questions excluded, Backboard on Gemini 2.5 Pro and GPT-4.1 as judge; a 2 to 3 point variance between runs is expected. Both runs, including code and per-question outputs, are published on GitHub. Neither involved benchmark-specific tuning.

* The LongMemEval score is a conservative lower bound: several responses that were more precise than the benchmark's expected answer were marked incorrect during evaluation.