Measured performance.
Published results.
Three public benchmark suites and one internal evaluation. Every run is published alongside its result.
84.3%
Terminal-Bench 2.1Terminal agents
75 of 89 tasks · +5.4 pts over the next-best agent on the same modelAgents do real end-to-end work in a live shell. An independent verifier decides pass or fail, with no partial credit.
Backboard-R-CLI-Terminal-Bench-2.1-Results ↗93.4%
LongMemEvalLong-term memory
467 of 500 questions · four separate judges agree within 2 pts500 questions asked against 40 to 50 prior sessions of conversation. Evaluated independently by NewMathData.
Backboard-longmemEval-results ↗90.0%
LoCoMoConversational memory
Leads all four categories · +14.2 pts over the next-best systemEvery system answers the same questions on the same conversations. The human baseline sits around 88%.
Backboard-Locomo-Benchmark ↗3.2–3.8x
BBQ-FP4Quantization
101.4% average recovery vs bf16 · 16 GB of weights against 57 GBBackboard's own FP4 quantization, measured against both a bf16 reference and an external 4-bit build at the same footprint.
Code not public yet