The AI Industry Has a Benchmark Problem

The AI Industry Has a Benchmark Problem

Benchmarks are for building, not bragging.

The AI industry has a benchmark problem.

Not because we have too many benchmarks.

Because we have too many companies treating them like trophies instead of tools.

That's unfortunate, because benchmarks are one of the most valuable engineering practices we have.

At Backboard, we don't benchmark because we believe benchmarks are perfect. We benchmark because imperfect, transparent measurements are better than subjective claims.

That idea shapes how we think about engineering.

If you've spent any time following AI over the past year, you've probably seen an endless stream of benchmark announcements. Every week another model reaches the top of another leaderboard. Every release claims a new state of the art. Every company seems to have a chart proving they're the best.

It's easy to become cynical.

The problem isn't that benchmarks exist.

The problem is that we've started confusing the measurement with the mission.

Benchmarks Are Feedback Loops

The primary reason we benchmark isn't to publish a score. It's to build a better product.

Every benchmark is a feedback loop. It tells us where we're strong, where we're weak, and whether the changes we made actually improved something meaningful. Sometimes an optimization delivers exactly what we hoped for. Other times, it exposes a regression we never expected. Without objective evaluation, it's remarkably easy to convince yourself that your product is getting better simply because you've spent weeks working on it.

Benchmarks have a way of keeping engineers honest.

In fact, I think every benchmark should pass one simple test:

A benchmark should challenge your engineers before it impresses your marketing team.

If it isn't making your product better, it probably isn't serving its most important purpose.

Benchmarks also create a common language.

Our customers shouldn't have to rely solely on our opinion of our own products. Public evaluations give everyone a shared point of reference. No benchmark captures every aspect of an AI system, but transparent measurements allow different products to be compared using the same criteria. That's healthier than a world where every company simply declares itself to be the best.

When the Measurement Becomes the Mission

Of course, benchmarks have limitations.

Like any measurement, they can be abused.

It's possible to optimize specifically for a benchmark rather than the capability it's supposed to measure. It's possible to leak evaluation data into training. It's possible to cherry-pick configurations. It's possible to publish only the results that make for impressive headlines.

None of those things improve the product.

They improve the marketing.

Build First. Benchmark Second.

There's an important distinction between building for benchmarks and benchmarking what you've built.

Building for benchmarks starts with the test. The goal becomes maximizing a score, even if that improvement never translates into real-world value.

Benchmarking what you've built starts with the customer. You solve real problems first, then use independent evaluations to validate that you're moving in the right direction.

Those approaches may produce similar-looking leaderboard results, but they're fundamentally different engineering philosophies.

Transparency Matters Just as Much as Performance.

Whenever possible, we publish our methodology. We open source our evaluation frameworks. We share logs, configurations, and results so others can reproduce our findings for themselves. If someone discovers we've made a mistake, that's not a failure of the process. It's evidence that the process is working.

Science advances because results can be challenged.

Engineering improves because assumptions are tested.

In our case, we've chosen to open source not only our benchmark results, but also the logs and evaluation artifacts behind them. We want people to inspect our work, reproduce it, challenge it, and improve upon it. That level of transparency creates far more confidence than simply posting a screenshot of a leaderboard ever could.

No Benchmark Will Ever Tell the Entire Story

They don't measure customer trust. They don't measure usability. They don't capture every workflow or every edge case that matters to an enterprise. Public benchmarks should always be complemented by real customer evaluations, production deployments, and continuous feedback.

That's why we see benchmarks as one input, not the only input.

There is another benefit, and it's one we're happy to acknowledge.

Benchmarks create reach.

Strong benchmark performance helps people discover what we're building. It starts conversations with engineers, customers, investors, and partners who otherwise might never have found us.

There's nothing wrong with that.

What's important is the sequence.

Benchmark to build a better product.

Be transparent about the methodology.

Share the results.

Learn from the feedback.

Then repeat.

The Leaderboard is a Byproduct, Not the Objective.

We hope our products perform well on public evaluations. Of course we do. But that's never been the goal.

The goal is to build software that genuinely helps people solve difficult problems.

If we ever stop learning from benchmarks and start treating them as trophies, we'll have missed the point entirely.

Because in the end, benchmarks don't build great products.

Engineers do.

Rob Imbeault

No headings found on page

SHARE