We just submitted Backboard CLI to the official Terminal-Bench 2.1 leaderboard with a score of 85.4% ± 0.8%.
If verified as submitted, it would be the highest score on the current leaderboard.
But the score itself is not the most interesting part.
Same Model. Different Harness.
We used Claude Opus 4.8 through Amazon Bedrock.
Claude Code, using the same model, currently scores 78.9%.
Backboard CLI scored 85.4%.
Same model. Different harness. A 6.5-point difference.
That is the story.
For the last few years, most of the AI industry has focused on which model is smartest. Models obviously matter, but increasingly, so does everything around them.
Context management matters. Tool use matters. Planning matters. Recovery matters. State matters. Verification matters. Token efficiency matters.
The model did not suddenly become smarter when we called it.
The system around it got more out of it.
Performance Is Only Half the Story
The economics are equally important.
Our full submitted Terminal-Bench run cost $280.72.
The current verified leader scores 83.8% at a reported cost of $552.67.
So our submitted result achieved a higher score for roughly 49% less reported cost.
That matters because at enterprise scale, small differences in inference efficiency become very large differences in infrastructure cost.
The question will increasingly become less about:
Which model are you using?
And more about:
What can you get the model to accomplish, how reliably can you do it, and what does it cost?
Benchmarks Are Feedback Loops
That is why we benchmark.
Not for the leaderboard itself, but because benchmarks expose where our system fails.
Every failed task gives us data. Every inefficient trajectory exposes something in the harness. We change it, test again, and keep improving.
This submission includes 89 tasks, five attempts per task, and 445 total trials under the same agent configuration. Every result is included, including failures and errors.
The aggregate score is 85.4%.
That is the number we were comfortable submitting.
The Model Is Only One Layer
There is a bigger thesis behind it.
We do not think organizations should have to rebuild their AI infrastructure every time a new model takes the lead.
Today it might be Claude. Tomorrow it might be OpenAI, Gemini, Kimi, DeepSeek, Mistral, or something that has not been released yet.
Models will keep changing.
The intelligence layer around them can persist.
That is what we are building at Backboard.
The goal is to continuously get more from the models underneath us: better execution, better context, better memory, better evaluation, and better economics.
Terminal-Bench gives us one measurable piece of evidence that this approach works.
Take the same frontier model.
Put a different system around it.
Move from 78.9% to 85.4%.
Then outperform the current leaderboard leader at roughly half the reported cost.
That is why we are excited about this result.
Not because one benchmark proves everything.
Because it supports a much bigger idea:
The next major gains in AI will not come only from making models smarter. They will also come from getting much better at using the intelligence we already have.
Our Terminal-Bench 2.1 submission is public as PR #200 and is currently awaiting official verification.
→ https://github.com/harbor-framework/terminal-bench-2-1/pull/200

Rob Imbeault
SHARE