Measured performance.
Published results.

Three public benchmark suites and one internal evaluation. Every run is published alongside its result.

84.3%
Terminal-Bench 2.1Terminal agents
75 of 89 tasks · +5.4 pts over the next-best agent on the same model

Agents do real end-to-end work in a live shell. An independent verifier decides pass or fail, with no partial credit.

Backboard-R-CLI-Terminal-Bench-2.1-Results ↗
93.4%
LongMemEvalLong-term memory
467 of 500 questions · four separate judges agree within 2 pts

500 questions asked against 40 to 50 prior sessions of conversation. Evaluated independently by NewMathData.

Backboard-longmemEval-results ↗
90.0%
LoCoMoConversational memory
Leads all four categories · +14.2 pts over the next-best system

Every system answers the same questions on the same conversations. The human baseline sits around 88%.

Backboard-Locomo-Benchmark ↗
3.2–3.8x
BBQ-FP4Quantization
101.4% average recovery vs bf16 · 16 GB of weights against 57 GB

Backboard's own FP4 quantization, measured against both a bf16 reference and an external 4-bit build at the same footprint.

Code not public yet