Skip to content
AIWorkBench.fr

Benchmark

First measurement: 27 models read 100 real invoices

The winner costs $0.0009 per invoice and misses none out of a hundred. The two most expensive models in the panel miss one — and three models invent a VAT number that does not exist.

The hub publishes its first real measurement. Twenty-seven models read the same one hundred invoices, and every answer was compared against an annotation written by a human.

The documents are not ours. They are real television advertising invoices that US stations are required to file with the Federal Communications Commission and that law makes public. Journalists entered them by hand, field by field, to track campaign spending. That entry is what serves as ground truth.

What came out:

  • Five models out of twenty-seven read the total amount without a single mistake across one hundred invoices.
  • The winner, GPT-6 Luna, costs $0.0009 per invoice — eighty times less than GPT-6 Astra, from the same provider, which misses one.
  • Three models invent. On a VAT number these American invoices do not carry, Kimi K3, Llama 4 Maverick and Mistral Medium each fabricated a value. The other twenty-four abstain correctly.
  • The last model in the ranking leaves nearly one invoice in three to be redone by hand, while being sold as a premium model.

Sample size is what separates a ranking from an illusion. On nineteen invoices, nineteen models were perfectly tied and none invented anything. On one hundred, only five remain, three inventions appear, and the two most expensive models in the panel drop out of the leading pack. A small sample does not say “these models are equivalent”: it says “this test does not separate them yet”.

How an answer is verified, since that is the question that matters: not by asking an AI to judge an AI. A hand-written comparator reads the model's answer and the annotation, and decides. An amount is compared to the cent. When the comparator cannot decide, the document goes to human review rather than being counted wrong.

The rubric scores only what admits a single right answer: the total amount billed, and the VAT number absent from the document. The other fields on these invoices — contract number, advertiser, period — admit several defensible answers: these documents carry six distinct identifiers and two different periods. So the question is no longer asked at all, rather than asked and then discarded.

All one hundred invoices were read by all twenty-seven models, without exception. It took several retries: a rate limit at a provider, or a model too verbose to fit in the room we give it, is enough to make an invoice unusable — and a ranking compares models on the same documents, or it compares nothing.

What this measurement does not say: the invoices are American and in English. It does not measure line-item extraction, which is the real difficulty of the job — the original annotation does not cover it. And twenty-five models out of twenty-seven exceed 97%: this task remains an entry threshold rather than a league table.

View the benchmark — Advertising invoice extraction