Skip to content
AIWorkBench.fr

Runnable protocolOn documents

Advertising invoice extraction

How many invoices does a model read with nothing for a human to fix?

Updated 5 October 2026 · Run 2026-09-27_facture-fcc+2026-10-02_facture-fcc+2026-10-04_facture-fcc+2026-10-05_facture-fcc · 27 models · 100 invoices

Advertising invoice extractionAccuracy × cost per test

Each square is a model, in its lab's colour and monogram. Click to open its page.

Key takeaways

  • 5 models tie for first place at 100.0% accuracy: GPT-6 Luna, DeepSeek V4.1 Flash, Command A+, Kimi K2.6 and Muse Spark 1.3. This test does not separate them; the table orders them by price, then by speed.
  • Best accuracy for the price: Gemma 4 31B, 0.7 points behind the best score at $0.0003 per invoice — 3.4 times cheaper than GPT-6 Luna.
  • 24 of 27 models invent nothing: 0% hallucinations.
  • The hardest field: “Total amount billed”, answered correctly 97.0% of the time on average across the models.

Which model to choose

If accuracy comes first

GPT-6 Luna

100.0% accuracy

If volume matters

Gemma 4 31B

$0.0003 per invoice

If errors are costly

GPT-6 Luna

0.0% invented items

The leaderboard

Click a header to re-sort. Accuracy orders the table; the other measures are never blended into it.

Advertising invoice extractionAccounting

RankModel (27)
1GPT-6 Luna100.0%±0.00.0%$0.0009
2DeepSeek V4.1 Flash100.0%±0.00.0%$0.0025
3Command A+100.0%±0.00.0%$0.0060
4Kimi K2.6100.0%±0.00.0%$0.013
5Muse Spark 1.3100.0%±0.00.0%$0.013
6Kimi K399.8%±0.71.0%$0.021
7Gemma 4 31B99.3%±1.20.0%$0.0003
8Qwen3.7 Flash99.3%±1.20.0%$0.0005
9Ministral 3 8B 251299.3%±1.20.0%$0.0010
10GLM 5.3 Flash99.3%±1.20.0%$0.0012
11Gemini 3.5 Flash Lite99.3%±1.20.0%$0.0015
12Qwen3.8 27B99.3%±1.20.0%$0.0036
13Gemini 3.8 Flash99.3%±1.20.0%$0.0070
14Claude Sonnet 599.3%±1.20.0%$0.015
15Qwen3.8 Max (0902)99.3%±1.20.0%$0.020

What this benchmark measures

Real television advertising invoices filed with the US telecom regulator and annotated by hand by journalists. The model receives the scanned pages and must return two things: the total amount billed, and the VAT number — the latter absent from these documents, to see who invents rather than abstains. Every answer is compared against the original annotation by a hand-written comparator: no AI grades an AI.

Where the documents come from

Invoices filed by US television stations with the Federal Communications Commission, which law makes public. The annotations come from journalists who entered them by hand to track campaign spending.

Licence
MIT (outillage) · documents publics déposés auprès de la FCC
Sample
100 invoices
What the model receives
Documents (images or PDF)
Models tested
27
Protocol
Runnable protocol
Status of the figures
Real measurement
Run
2026-09-27_facture-fcc+2026-10-02_facture-fcc+2026-10-04_facture-fcc+2026-10-05_facture-fcc

Where models pull apart

The average across all models on each graded field, from best mastered to hardest, and the best model on each.

The documents that split the models

These invoices are real: no difficulty was planted in them. They are the ones the models contradicted each other on most. Here is the page, the original annotation, and what each model answered — enough to judge the scoring without taking our word for it.

Montant total facturé

13 of 27 models got it wrong · field: Montant total facturé
Test invoice 9ed3362a-84b4-1d8e-7aab-1ed898008b16
6 scanned pages · 9ed3362a-84b4-1d8e-7aab-1ed898008b16
Original annotation
6414
Nova Lite 1.0
6414
Nova Pro 1.0
5451.9
Claude Fable 5.1
9347
Claude Opus 5.5
9347
Claude Sonnet 5
6414
Command A+
6414
DeepSeek V4.1 Flash
6414
Gemini 3.5 Flash Lite
6414
Gemini 3.8 Flash
9347
Gemma 4 31B
6414
Llama 4 Maverick
6414
Muse Spark 1.3
6414
Ministral 3 8B 2512
6414
Mistral Large 3 2512
5451.9
Mistral Medium 3.5
6414
Mistral Small 4
6414
Kimi K2.6
6414
Kimi K3
6414
GPT-6 Astra
9347
GPT-6 Luna
6414
GPT-6 Sol
9347
Qwen3.7 Flash
9347
Qwen3.8 27B
9347
Qwen3.8 Max (0902)
9347
Qwen3.8 Max Prime
9347
Grok 4.7
nothing — the model abstains
GLM 5.3 Flash
9347

Montant total facturé

9 of 27 models got it wrong · field: Montant total facturé
Test invoice 93cd08d3-28b1-31c0-8902-4060655fcb91
2 scanned pages · 93cd08d3-28b1-31c0-8902-4060655fcb91
Original annotation
4720
Nova Lite 1.0
4838
Nova Pro 1.0
4012
Claude Fable 5.1
4838
Claude Opus 5.5
4720
Claude Sonnet 5
4838
Command A+
4720
DeepSeek V4.1 Flash
4720
Gemini 3.5 Flash Lite
4838
Gemini 3.8 Flash
4720
Gemma 4 31B
4720
Llama 4 Maverick
4838
Muse Spark 1.3
4720
Ministral 3 8B 2512
4720
Mistral Large 3 2512
4720
Mistral Medium 3.5
4838
Mistral Small 4
4838
Kimi K2.6
4720
Kimi K3
4720
GPT-6 Astra
4720
GPT-6 Luna
4720
GPT-6 Sol
4838
Qwen3.7 Flash
4720
Qwen3.8 27B
4720
Qwen3.8 Max (0902)
4720
Qwen3.8 Max Prime
4720
Grok 4.7
4720
GLM 5.3 Flash
4720

Montant total facturé

6 of 27 models got it wrong · field: Montant total facturé
Test invoice 4b703aaa-2e13-528c-85f8-9bd7879d4cbb
4 scanned pages · 4b703aaa-2e13-528c-85f8-9bd7879d4cbb
Original annotation
7000
Nova Lite 1.0
6500
Nova Pro 1.0
5950
Claude Fable 5.1
7000
Claude Opus 5.5
7000
Claude Sonnet 5
7000
Command A+
7000
DeepSeek V4.1 Flash
7000
Gemini 3.5 Flash Lite
7000
Gemini 3.8 Flash
7000
Gemma 4 31B
7000
Llama 4 Maverick
6500
Muse Spark 1.3
7000
Ministral 3 8B 2512
6500
Mistral Large 3 2512
6500
Mistral Medium 3.5
7000
Mistral Small 4
6500
Kimi K2.6
7000
Kimi K3
7000
GPT-6 Astra
7000
GPT-6 Luna
7000
GPT-6 Sol
7000
Qwen3.7 Flash
7000
Qwen3.8 27B
7000
Qwen3.8 Max (0902)
7000
Qwen3.8 Max Prime
7000
Grok 4.7
7000
GLM 5.3 Flash
7000

The documents analysed

The 100 documents in the test, from the most-missed to the most agreed-on: the annotation used as reference, how many models found it, and a link to the original where it is published. The ranking does not ask to be believed: it can be recounted.

The documents analysed
DocumentExpected Total amount billedModels correctInvented valueOriginal
9ed3362a641414 / 27—open
93cd08d3472018 / 27—open
4b703aaa700021 / 27—open
152629831765023 / 27—open
17bee528209524 / 27—open
3fb2b31f20925024 / 27—open
2388be1199525 / 27—open
4809552b35025 / 27—open
4e1c4012225525 / 27—open
5d56f8db1210025 / 27—open
64780ed04070025 / 27—open
7bbe366d448525 / 27—open
afa6fb202130025 / 27—open
d2bbcfb235525 / 27—open
dd623c624070025 / 27—open
03b940c21635026 / 27—open
08bea5442795026 / 27—open
0be55a7b142826 / 27—open
0e4719a3668026 / 27—open
250e73dd375526 / 27—open
26f3961b3490026 / 27—open
3f335f42654526 / 27—open
4cee59d1368526 / 27—open
59d47fbe958026 / 27—open
5b40f242536026 / 27—open
5d8fb479352626 / 27—open
703ebcab21226 / 27—open
74c779b47250026 / 27—open
97ea6a67760026 / 27—open
98905496188026 / 27—open
991a6b5e172026 / 27—open
9cc36a8344026 / 27—open
a5b12494605026 / 27—open
bc343c3862526 / 27—open
c6ec7473362526 / 27—open
d6bd95e8242526 / 27—open
d748c96419426 / 27—open
e3f3dba66026 / 27—open
f9889d44190026 / 27—open
02ccb7f221035027 / 27—open
0458e5e4212027 / 27—open
05a53fa5602527 / 27—open
0617c152746527 / 27—open
0880bc9a4473027 / 27—open
0a32ce11142527 / 27—open
11a46278826027 / 27—open
1fd3058f1149027 / 27—open
1fd7b5c94321527 / 27—open
21228dd6306027 / 27—open
296043904275027 / 27—open
2aac88c6578027 / 27—open
331ff2f32230027 / 27—open
345c9cc01550027 / 27—open
3499dd36867027 / 27—open
36ff284e225027 / 27—open
3bdbc1a7338427 / 27—open
48e3a9521468527 / 27—open
5c8edc38311027 / 27—open
5e0773ca3450027 / 27—open
627d67b52218027 / 27—open
62e6d540485027 / 27—open
6c10638e220027 / 27—open
707c944c2917527 / 27—open
73e159452870027 / 27—open
76b092d7303527 / 27—open
7b9146bb102527 / 27—open
7e00d87536027 / 27—open
805aafc61604527 / 27—open
8222d0e91009527 / 27—open
8588921b312027 / 27—open
88d79dde86527 / 272open
8b687291885027 / 27—open
8e41b95417998027 / 27—open
9215778a350027 / 27—open
9294d8b31507027 / 27—open
988878c1367527 / 27—open
a21fa4343950027 / 27—open
a24ef1d82196027 / 27—open
a7d60614482027 / 27—open
a8aaf9d51415027 / 27—open
abf2f735169027 / 27—open
b34fc957263527 / 27—open
be9e50f7620027 / 27—open
c0b67c9f5027727 / 27—open
c6974e8e23535027 / 27—open
ca356a4b47927 / 27—open
d06fd8db180027 / 27—open
d1ba92252438027 / 27—open
d297f399692527 / 27—open
d6cd4fa880027 / 27—open
da77aaa05560027 / 27—open
db47993a547027 / 27—open
e0260276445527 / 27—open
e5e2f3433220027 / 27—open
e6846bc9421527 / 27—open
e8d412047320027 / 27—open
ec8415e2121327 / 27—open
ee19ec768027 / 27—open
f56e2fc1269027 / 271open
f5871a461515027 / 27—open

How these figures are produced — the scoring grid, the exact prompt, and what this test does not measure.