Here’s how we performed—and what else is new.
85.4% on Terminal-Bench 2.1
Backboard R-CLI achieved 85.4% on Terminal-Bench 2.1, our strongest result to date.
The result demonstrates what we've been focused on from the beginning: coding performance isn't determined by the model alone.
The harness matters.
How work is reasoned through, executed, and managed around the model can materially change the result.
91% with Kimi K3
R-CLI also achieved 91% running Kimi K3 on Terminal-Bench 2.1, demonstrating that strong coding performance isn't limited to closed-source frontier models.
Together, Backboard R-CLI sits at the top across both closed- and open-source models.
For developers, that means more flexibility in the models they choose without treating the model itself as the entire coding system.
Top Performance, Lower Cost
Performance isn't the only metric we're optimizing.
This benchmark also produced our lowest-cost run yet, combining our strongest performance with improved cost efficiency.
The goal isn't to spend more inference to get a better result.
It's to get more useful work from the models and tokens you're already paying for.
Automatic Reasoning
R-CLI now handles reasoning automatically.
Instead of requiring developers to manually determine how much reasoning a task needs, the CLI can adapt its reasoning behavior as part of the workflow.
That means less configuration when moving between straightforward edits and more complex coding tasks.
These improvements are focused on making everyday development sessions more stable and predictable as we continue expanding what the CLI can handle.
Open source coming soon.
Ask. Review. Ship.
Get started with R-CLI → backboard.io/products/cli.

Rob Imbeault
SHARE