Skip to content
AIWorkBench.fr

Analysis

A one-point gap is a tie

Twenty-five invoices cannot separate two models one point apart. The hub publishes the margin of error of its scores, and says so when a ranking does not settle the matter.

A leaderboard tempts you to read ranks as truth: first beats second, second beats third. On a set of twenty-five invoices, that is often false.

Do the sums at the scale of one document. With twenty-five invoices, a single failed invoice moves the share of invoices handled without review by four points. Two models separated by a gap of that size differ by one document — perhaps a scan a little more skewed than the rest. Run the test again with twenty-five other invoices, and the order may flip.

That is what the margin of error measures: the half-width of the 95% confidence interval around accuracy, in points. A score of 90% with a margin of ± 3 points means that the value you would get on a much larger set probably lies between 87 and 93.

The hub's rule is simple: two models separated by less than the larger of their two margins are not told apart. The rank puts them in order, because a table has to be displayed in some order; it does not distinguish them.

The site flags this in two ways. In tables, the margin sits next to the score whenever it is published: “± 2.1”. And in a benchmark's key takeaways, when the gap between the top two is smaller than the margin, it says so: this test does not separate them.

The Business Index has a margin too, combined from those of the benchmarks that make it up. It tightens as benchmarks are added: several tests say more than one. While a single benchmark is measured, the index's margin is that benchmark's, no more and no less.

The margin does not fix everything. It describes the luck of the draw in the documents, not the limits of the protocol: a single prompt, images rather than PDFs, no per-model tuning. It only tells you how far a gap can be trusted.

In practice: do not choose a model over a one-point gap. When two models tie on accuracy, the other columns decide — cost, time and, above all, hallucinations.