The first AI platform to lead both major memory benchmarks
93.4% on LongMemEval in an independent evaluation, and 90.0% on LoCoMo, ahead of every published memory system in every category. Both runs are public and reproducible.
90.0%
LoCoMo accuracy
Highest score in all four question categories
Human baseline ~88%, frontier models ~80%
+14.2 pts
Ahead of the next-best system
Memobase 75.78%, Zep 75.14% on LoCoMo
Mem0, LangMem and OpenAI memory trail further
LongMemEval: 93.4% on the original academic specification
500 questions across six question types, each with roughly 40 to 50 sessions of history. Evaluated independently by NewMathData and scored by four separate LLM judges, which agree within about two points overall.
LoCoMo: 90.0% overall, first in every category
Ten multi-session conversations ingested turn by turn into isolated assistants, then questioned across single-hop, multi-hop, open-domain and temporal reasoning. Backboard ran on Gemini 2.5 Pro; GPT-4.1 judged every answer.
Breaking the 90-percent threshold requires superhuman consistency in recall and reasoning. Most high-performing frontier models currently score around 80 percent on LoCoMo. The system built by Backboard.io is a far better attempt at simulating memory as it manifests in humans.
Adyasha Maharana, creator of the LoCoMo benchmark and research scientist at Databricks, in the Ottawa Business Journal
What this means for agentic AI
Without reliable, shared memory, agents fragment, hallucinate, and reset. Backboard gives agents persistent, shared memory even when they run on different underlying models, so agentic behavior emerges naturally rather than being scripted.
Why leading both benchmarks matters
Most systems optimize for either short-horizon precision or long-horizon persistence, not both. Backboard's results reflect system-level behavior built on Active Temporal Resonance, which preserves meaning and continuity as interactions unfold, rather than benchmark-specific tuning.
Methodology: LongMemEval was evaluated independently by NewMathData on the benchmark's original academic specification (s_cleaned, 500 questions) with Backboard on GPT-4.1. LoCoMo used the standard dataset with adversarial questions excluded, Backboard on Gemini 2.5 Pro and GPT-4.1 as judge; a 2 to 3 point variance between runs is expected. Both runs, including code and per-question outputs, are published on GitHub. Neither involved benchmark-specific tuning.
* The LongMemEval score is a conservative lower bound: several responses that were more precise than the benchmark's expected answer were marked incorrect during evaluation.