AI and models · 18 Aug 2026

What an evaluation actually measures, and what a vendor means by one

Evergreen background. The word covers three different things, and vendors rely on the confusion.

By Felix Albuerne Jr.

An evaluation, in the sense a research team uses it, is a fixed set of tasks with a known answer key, run against a model to produce a comparable number. The value is comparability, not realism.

An evaluation, in the sense a buyer needs, is a measurement of whether the system does the buyer's job on the buyer's data, including the ugly inputs. It is not comparable to anyone else's and it is the only kind that predicts what happens after purchase.

An evaluation, in the sense a vendor deck uses it, is usually the first kind, presented as evidence for the second. When a chart shows a benchmark score next to a customer logo, those two things are not connected, and the burden of connecting them sits with the seller.

The practical test: ask which inputs the number was produced on, who wrote the answer key, and what the score was on the runs that were not shown. A vendor who cannot answer the third question has not run an evaluation, they have run a demo.

Sources

  • Vendor methodology note