Skip to content
AIWorkBench.fr

Benchmark

How a business benchmark gets built

An executive's question, real documents already annotated by humans, a scoring grid weighted by what a mistake costs, human arbitration, a frozen publication. Five steps, in that order.

A hub benchmark does not start with a model, or with a technology. It starts with a question an executive would ask: how many of my invoices will go through without a human fixing them? If the question cannot be put in one sentence, to someone who has never opened a technical manual, the benchmark is not ready.

Then come the documents. They are real and public, never borrowed from a client nor fabricated by us: the repository is public, and an invented invoice never quite looks like the real thing. The ground truth does not come from us either — it is the annotations already made by humans on those documents, before this test and without knowing which models would take it. This is the hub's hardest constraint: it limits testable tasks to those for which such a dataset exists, and that is why most benchmarks in the catalogue are still on the roadmap rather than measured.

Third step, the scoring grid. Every graded item gets a weight from 1 to 3, set by what a mistake costs the business, not by how technically hard it is. On an invoice, the gross total weighs 3 and the due date 1. Some fields are marked critical: one critical field wrong, and the whole invoice drops into “needs review”. A field is worth its weight or zero, with no half points: an amount is right or it is wrong.

Scoring is automatic, and it is fussy where the business is fussy. Amounts are compared to the cent. Dates follow the convention of the document's country, declared field by field: on an American invoice, 03/04/2020 is 4 March; on a French one, 3 April. It is lenient where the business is lenient too: an identifier may be written with or without spaces. And when the comparator cannot decide, the answer goes to review rather than being counted wrong.

Fourth step, human arbitration. The scoring code is not an authority; it is a sorting tool. Three kinds of case go in front of a human: every hallucination, because they decide the heaviest measure; answers judged correct but written differently from the reference, because that is where the comparator can be too lenient; and a control sample among the answers judged correct. The human decides, and that verdict prevails. On the ranking published so far this stage has not been played yet, and the methodology page says so.

Last step, publication. A run is a timestamped, immutable folder: each model's raw answers are kept exactly as produced, along with the exact version of the model that answered, and the cost and time of every call. Nothing is ever rewritten. Changing the grid and rescoring costs no API call at all, which is what keeps comparisons over time honest.

Two safeguards frame publication. A model that fails on more than a tenth of the set blocks it: better no ranking than one computed only on the documents a model was willing to process. And failed calls stay on the record: they enter no average, but they appear in the ranking.

Today, one benchmark has been through all five steps: invoice reading, on real annotated public invoices. The other twenty-one are at the first — the question is asked, the protocol is written — and are waiting for their test set.