Skip to content
AIWorkBench.fr

Methodology

Two parts: the method common to the whole hub, then the protocol specific to each measured benchmark. Everything can be checked in the repository — the prompt sent, the scoring grid applied, the documents submitted and each model's raw answers.

The Business Index

The Business Index sums a model up in one figure. It is computed in two steps:

  • a business function's score is the mean of the model's accuracy on the benchmarks it took in that function, rounded to a tenth of a point;
  • the index is the mean of those per-function scores, rounded the same way.

So every business function weighs the same, however many benchmarks it has: a function with three tests does not count three times. A model scoring 90% and 70% on one function's two benchmarks, and 50% on another's only benchmark, has function scores of 80% and 50%. Its index is 65%, not 70%, the mean of the three tests.

The index covers only the business functions already measured. A function whose benchmarks have not been run does not enter the computation: it does not count as a zero, it does not count at all. How many functions the index covers is written next to it, because an index over one function is not worth an index over ten.

A model missing from a function the others took, however, has neither an index nor a rank: it appears as “Not ranked”, at the bottom of the tables. Its index could not be compared with any other.

A model that cannot read documents — neither images nor PDFs — does not take the benchmarks that submit one. Within each business function, it is scored on the text benchmarks alone; a function with none would leave it without an index. Its index therefore rests on fewer tests than the others': the number of benchmarks taken is shown alongside.

The index aggregates accuracy alone. A model's cost, time and hallucination rate are means over the benchmarks it took, shown separately and never melted into the index. An unknown price stays out of the cost mean.

Rank follows the index, from highest to lowest. On a benchmark, it follows accuracy.

The margin of error

A score measured on a few dozen documents is not known to a tenth of a point. The margin of error says so: it is the half-width of the 95% confidence interval on accuracy, in points. When a ranking publishes it, it is shown next to the score: “± 2.1”.

Two models separated by less than the larger of their two margins are not told apart. The table still puts them in order, because some order is needed; the benchmark page, for its part, states that the test does not separate them. With no published margin, the site declares no tie.

The Business Index has a margin of its own, combined from those of the benchmarks taken: the square root of the sum of their squares, divided by their number. It is published only if every benchmark the model took publishes one.

The margin describes the luck of the draw in the documents. It says nothing about the limits of the protocol, listed further down.

Costs and prices

The prices shown are the models' public prices, in dollars per million tokens, for input and output. They are synced from OpenRouter's public catalogue — the gateway the calls are actually billed through — as are context windows; last synced on 2 October 2026.

They stay in dollars: converting them would date the figure. A model missing from the catalogue is labelled “Undisclosed” rather than given a price copied from memory.

Cost per test is the mean of successful calls, at the price on the day of the test: the tokens each call consumed, multiplied by the catalogue price. It is recorded at run time, never reconstructed afterwards, and recomputed from published prices rather than read from the provider: anyone reading the repository can redo the multiplication.

A failed call enters no average, and is counted separately. A model that fails on more than a tenth of the set blocks publication of the ranking.

A failure may still have been billed: a model that answers and then exceeds the length limit consumes tokens without returning a usable result. Those calls are recorded with their cost, so that the total spent is accurate, but they do not weigh on a model's cost per test.

A failure on our side does not penalise a model's score: account quota, credit reserved by concurrent calls, a host being down. Its count stays in the ranking. When the model itself returns nothing usable, the rule belongs to the benchmark and its protocol states it: on the financial questions, that failure counts against the model, and against it alone.

The time shown is a median, not a mean: a single slow call would skew the mean.

How an answer is verified

Querying a model is the easy part. Everything else comes down to one question: how do we know its answer is right?

No AI grades an AI. The reference is written by humans, before the test and without knowing which models will take it. For the invoices, journalists entered every field by hand. Having one model judge another would measure their agreement, not their correctness.

The comparison is done by a hand-written program, field by field, according to the nature of the information:

  • an amount is reduced to a number, then compared to the cent;
  • a date is reduced to year-month-day from the dozen forms models use — “12/03/2020”, “March 12, 2020”, “2020-03-12” — applying the reading convention of the document's country, not the reader's;
  • an identifier or a name is compared after normalisation: accents, spaces, punctuation, common abbreviations;
  • a list of line items is matched line by line, order not counting.

This program never guesses. When it cannot decide — an unreadable value in the reference, a form it does not recognise — the answer goes to human review rather than being counted wrong. A reference we cannot read fails the preparation of the test: without that rule it would become a silent “nothing”, and a model that read the document correctly would be accused of inventing.

The same document, the same prompt, the same conditions for every model. A document one provider refuses — too many pages, too many pixels — is dropped from the test for everyone, not just for that provider: a ranking must compare models, not subsets of documents.

Ambiguous questions are excluded, not patched afterwards. If a field admits several defensible answers, scoring it measures whether the model guesses what we wanted, not whether it can read. The field is then removed from the rubric and the reason is written into the rubric itself; the models' answers stay published.

Finally, the pipeline's four stages are separate and immutable: call the models, score, arbitrate, publish. Raw answers are written once and for all, and re-scoring costs no call. That is what makes it possible to fix a rubric without ever touching an answer.

The raw answers of all 27 models queried are in the repository, one folder per model: anyone can redo the scoring and get the published figures back.

Nothing the site displays is fabricated. A task is measured and carries its figures, or it is on the roadmap and carries none.

As of today, no published ranking is a demonstration: every figure on the site comes from calls actually made.

Measured or coming

Every benchmark carries one of these two labels:

  • Runnable protocol: the task runs end to end and carries measured figures. Its definition, scoring grid, prompt, documents and annotations are in the repository; anyone can rerun the pipeline and get the same figures.
  • Coming: the protocol is written — the question asked, what the model receives, the subtasks, the target sample size — but the test set has still to be built. No figure is shown.

Measured as of today: 2 of 22 (Financial report analysis, Advertising invoice extraction).

A task that is coming states what it is waiting for, because the causes are not equivalent: a real, annotated public dataset still to be wired in; a task for which no public reference gives the right answer, and which therefore needs a scoring rubric and human arbitration; or documents that never leave the company, to be collected from partners with their consent.

Why publish a protocol before its results? Because it can no longer be adjusted afterwards to suit a ranking, and because it is easier to challenge before than after.

One benchmark's protocol

One protocol per benchmark

Everything above applies to the hub as a whole: the Business Index, the margin of error, where costs come from, how mature a task is. What follows applies to one benchmark only. Its limits, its scoring grid, its verdicts, its prompt and its latest run belong to it, and say nothing about the others.

Every measured benchmark has its own section here, and a task that is coming has none until it is measured. A protocol does not carry over from one task to the next: an amount is checked to the cent, whereas sorting applications or summarising a contract calls for a different grid entirely — and for judgements no automatic comparator can make.

Advertising invoice extraction

The protocol: Advertising invoice extraction

One hundred American advertising invoices, twenty-seven models, two scored fields: the total amount, and the VAT number these documents never carry.

View the benchmark

Advertising invoice extraction

What this test does not measure

  • One prompt per model. A prompt tuned for a given model would improve its results; we measure what a standard integration gives.
  • No fine-tuning, no specialised OCR upstream. A dedicated processing chain would do better.
  • One hundred documents, not ten thousand. Enough to see a twenty-point gap, not to separate two models one point apart. Each score's margin of error says so.
  • American invoices, in English, from a single sector: television advertising buys. Nothing guarantees these results carry over to a French invoice.
  • Images, not PDFs with a text layer. A native PDF is easier for a model to read: published scores are a floor, not a ceiling.
  • Two scored fields. The other fields on these invoices admit several defensible answers — six competing identifiers, two distinct periods — and the question is no longer asked.
  • No line items. That is the real difficulty of the job, and the original annotation does not cover it: we would have to build our own reference.
  • A task easier than expected. Twenty-five models out of twenty-seven exceed 97%: this test separates them poorly, and should be read as an entry threshold rather than a ranking.

Advertising invoice extraction

The scoring grid

Fields are weighted by what a mistake costs in accounting, not by how technically hard they are. One critical field wrong drops the whole invoice into “needs review”.

Scoring grid of this benchmark
FieldComparisonWeightCritical
Montant total facturéto the cent3yes
Numéro de TVAidentical, punctuation ignored1no

Field labels are the original French ones, as written in the task definition.

A field is worth its weight or zero, with no half points. Accuracy is the sum of points earned over the total possible. An invoice counts as “no review needed” when none of its critical fields is wrong, missing or invented; an invoice the model failed to process does not count.

The comparison is fussy where the business is fussy: amounts are compared to the cent. It is lenient where the business is lenient too: an identifier may be written with or without spaces, and the order of line items does not matter.

Dates follow the convention of the document's country, declared field by field in the rubric. On these American invoices, 03/04/2020 is read as 4 March; on a French invoice the same writing would mean 3 April. Reading an American date the French way had in fact produced, on a first attempt, some twenty errors wrongly blamed on the models.

Advertising invoice extraction

The three verdicts

A field absent from the document is part of the test. Answering “nothing” when there is nothing counts as a right answer; producing a value counts as a hallucination, tallied separately from the rest and never melted into an overall score.

  • Correct: the value on the document, or “nothing” when the document contains nothing.
  • Wrong or missing: a value other than the one on the document, or nothing when the information was there.
  • Hallucinated: a value where the document contains none.

Only a correct field earns points. A hallucination also feeds a measure of its own: the share of genuinely absent fields that the model filled in anyway. A model that writes “non trouvé” or “n/a” is abstaining correctly: its answer is reduced to “nothing” before it is scored.

The code does not have the last word. A human arbitration stage selects the answers to review — every hallucination, those judged correct but written differently from the reference, and a sample of the rest — and the human's verdict replaces the code's in the computation.

On the published run, that stage has not been played yet: the figures shown are the comparator's alone. 80 of the 2700 gradings are flagged “needs review”, and the ranking will say otherwise once they have been arbitrated. Saying so, rather than implying a review that did not happen, is part of the protocol.

Advertising invoice extraction

The prompt, in full

Sent as is to every model, along with the scanned pages of the invoice. The same for all twenty-seven, without a word of difference.

You are given the pages of a broadcast advertising invoice, as images.
Extract the requested information and answer with a JSON object only.

## The most important rule

If a piece of information **does not appear** on the document, answer `null` for
that field. Never guess, never infer, never fill in what looks usual. A value you
invent is treated as a serious error — worse than admitting you did not find it.

## Fields

- `contract_num` — the contract or order number identifying this buy, exactly as
  printed. String.
- `advertiser` — the value of the field labelled `Advertiser` or `Advertiser Name`:
  the client who bought the advertising. Not the TV station, not the media agency,
  and not the `Product` or `Brand` field, which often repeats the advertiser's name
  in a different form. String.
- `gross_amount` — the total gross amount invoiced, for the whole document.
  This is usually a grand total, and it is often on a later page than the first.
  Number.
- `flight_from` — the first day of the advertising period covered by this invoice.
  When the document names that period — a field labelled `Flight Dates`,
  `Order Flight` or `Flight` — use it, and prefer it over `Invoice Period`,
  `Bill Period` or `Bill Plan`, which cover the billing cycle and often start on a
  different day. **Many of these documents have no flight field at all**: they
  carry only a billing period, usually labelled `Period`. Use that one then, and
  do not answer `null` — the period is on the document, under another name.
  Format `YYYY-MM-DD`. String.
- `flight_to` — the last day of that flight. Format `YYYY-MM-DD`. String.
- `vat_number` — the VAT registration number of the issuing company, if the
  document carries one. String.

## Value formats

- Dates: `YYYY-MM-DD`. A date printed `02/03/20` on a US document means
  3 February 2020, so it becomes `2020-02-03`.
- Amounts: a plain decimal number, **no currency symbol and no thousands
  separator**. `$1,880.00` becomes `1880.00`.

## Output

A JSON object with exactly these six keys, and nothing around it.

Advertising invoice extraction

This test

Latest published run of this benchmark
Run2026-09-27_facture-fcc+2026-10-02_facture-fcc+2026-10-04_facture-fcc+2026-10-05_facture-fcc
Date5 October 2026
Documents100
Models27
Statusreal measurement
View the benchmark

Financial report analysis

The protocol: Financial report analysis

The first layer of a financial analyst's work, on real reports. The questions were written by analysts for FinanceBench and bear on one or two pages of a US listed company's report: compute an EBITDA margin or days payable outstanding from two financial statements, say whether a margin is improving, name the segment that held growth back, recognise that a metric makes no sense for a bank. One answer per question: a number, a yes or no, or the name of a line item. Sixty-four questions count in the ranking, for twenty-seven models: forty-eight come from a first draw by quotas, sixteen are calculations added afterwards.

View the benchmark

Financial report analysis

What this test does not measure

  • Sixty-four questions. Enough to see a ten-point gap, not to separate two models one or two points apart: one question is worth 1.6 points. Each score's margin of error says so.
  • The top of the ranking is not separated. Seven of twenty-seven models score 100%, and their order only reflects price then speed, which break the ties. They are the same seven on the forty-eight questions of the first draw and on all sixty-four: the sixteen added calculations did not separate them. This test tells which models fail, not which of those seven reasons best on a hard question.
  • Calculations are over-represented: thirty-two questions out of sixty-four. Sixteen were added after the first run, because compound calculations are where models go wrong. All were taken, with no choice made on results, except the dividend payout ratio, dropped before the run. The first draw alone gives the same ranking (rank correlation 0.99).
  • One question trips up more than half of the panel: Netflix's (EBITDA margin), missed by seventeen of twenty-seven models, one of which answered nothing. Seven add the two content-amortization lines ($3.4 billion for streaming, $79 million for DVD) to “D&A” and get 56.8% instead of 5.4%. The reference takes the single line named “Depreciation and amortization of property, equipment and intangibles”, which is what the question asks for (“D&A from cash flow statement”), and we keep it. Without that question, ten models are perfect instead of seven, and the order changes little (rank correlation 0.97).
  • The useful pages are supplied. The model receives the page or pages holding the answer, not the whole report: we measure what an analyst does once the right page is open — read off, compute, judge — not the search for the information in a report of several hundred pages.
  • Famous companies, reports published years ago, a public dataset since 2023: a model may know some of these figures by heart.
  • American reports, in English. The question is asked in English, word for word, on both versions of the site.
  • What an automatic answer can check, and nothing else: 91 questions out of 150. The other 59 — written explanations, opinions, numbers whose unit the question does not fix — are listed line by line in the published sorting, with their reason. These are the most checkable questions, not a representative sample of the job: writing a note, comparing companies, forecasting or defending a view are not graded here.
  • One answer per question, written by an analyst, with whatever errors it may carry. On the 3M question about fixed assets, the reference says 8.70 billion while the page shows 8,738 million: twenty-four models answer 8.738. The 0.5% tolerance absorbs it; at 0.1%, no model is perfect any more.
  • A yes or no has a fifty-fifty chance of being right by luck: the twelve verdicts weigh less than their number suggests.
  • Only three questions have “there is none” as their right answer. The hallucination count is a count out of three, not a reliable percentage.

Financial report analysis

The scoring grid

Each question is worth one point, on a single criterion: the one the question asks. A criterion that a question's ground truth does not declare is not scored for it. Without that rule it would read as “absent from the document” and earn points for a model that answered nothing.

Four criteria, one per form of answer. Reading a figure and computing a metric are compared the same way; they are kept apart so that the site can say where a model goes wrong.

Scoring grid of this benchmark
FieldComparisonWeightCritical
Relever un chiffrewithin 0.5%, or half of the last digit1yes
Calculer un indicateurwithin 0.5%, or half of the last digit1yes
Trancher un faitidentical, punctuation ignored1yes
Nommer un posteidentical, punctuation ignored1yes

Field labels are the original French ones, as written in the task definition.

A number is right within 0.5% of the reference, or within half of its last displayed digit when that is wider: the reference “1.9” accepts 1.85 to 1.95, the reference “1616” accepts 1,607.9 to 1,624.1. References are rounded by their author, and the page often carries a more precise figure: without this tolerance, a model that reads correctly would be marked wrong. It was set before the run.

The decimal point is read as such, and the “%” and “$” signs are ignored: “1.734” is 1.734, not 1,734.

A verdict is “yes”, “no” or “not_applicable”. A label is compared with a list of accepted forms, after stripping accents, case and punctuation. Equality is whole: “Corporate & Investment Bank” does not pass for “Corporate”.

The accepted forms were extended on 7 October, after the run, and that is a correction on our side. For the three questions “which activity brought in the most cash”, we expected “operations” or “operating activities”; twelve answers naming the line printed on the page — “Net cash provided by operating activities” — had been refused, among them Claude Fable 5.1's and Opus 5.5's. We added the printed labels and redid the scoring. A list of accepted answers written before seeing the answers can be too narrow; saying so is part of the protocol.

Financial report analysis

The verdicts

Three verdicts are used here: right, wrong or missing, and hallucinated. A question whose right answer is “there is none” is a hallucination trap: a model that names a security or a line item anyway is making it up.

Only three questions are of that kind: the listed debt securities of American Express and Ulta Beauty, and Ulta Beauty's acquisitions. Three hallucinations in all: Nova Pro makes up two, Nova Lite one. They feed their own measure and are never folded into the overall score.

A failed call does not count against the model when it comes from us: a host's rate-limit refusal, exhausted credit, an answer extractor that was too narrow. It counts against the model when it returned nothing usable: two contradictory answers, a subtraction instead of a number, thinking that never ends. Five cases, each read and explained on the benchmark page: Mistral Large three times, Kimi K3 once, Ministral 8B once. The model is marked “missing” there, and it alone.

This rule was adopted after the first run. The previous one removed the question for the whole panel as soon as one model did not answer it, which made the most disputed questions drop out, Netflix's among them: the test became easier. In two cases Mistral Large's last answer was right; we do not decide for it.

The code does not have the last word, but human arbitration was not played on this run: the figures shown are the comparator's alone. The three hallucinations and a control sample are waiting for it. As written, the step does not present answers judged wrong: a label the comparator does not recognise does not go through it, which is how a too-narrow list of forms could go unnoticed until we looked at the answers question by question.

Financial report analysis

The prompt, in full

The template sent to every model, with the question and the report pages as images. Only the format instruction for the relevant form — a number, a verdict or a label — goes out with the question. The same for all twenty-seven, not a word different.

One sentence of the instruction had an effect we had not foreseen: “a percentage of 12.5% is written 12.5” pushed 21 of 23 models to write Coca-Cola's return on assets as a percentage, when the reference was the rounded ratio, 0.01. That question, its twin on AES and General Mills' retention ratio were dropped after the runs: they measured the scale chosen, not the computation. The prompt was not changed, so that every published figure comes from the same text.

You are given one or two pages of a company's financial report, as images, and one
question. Answer from these pages only, with a JSON object and nothing around it.

Question: {{question}}

<!-- forme: nombre -->
## Answer format

`{"answer": <number>}`

A plain decimal number: no currency symbol, no thousands separator, no unit. Use the
unit the question asks for; a percentage of 12.5% is written `12.5`.

<!-- forme: verdict -->
## Answer format

`{"answer": "yes"}`, `{"answer": "no"}` or `{"answer": "not_applicable"}`

Use `"not_applicable"` only when the question itself invites you to say the metric is
not relevant for this company, and it is not.

<!-- forme: libelle -->
## Answer format

`{"answer": "<name>"}`

The short name of the item, as printed on the page. If the pages show there is none,
answer `{"answer": null}`.

Financial report analysis

This test

Sixty-seven questions prepared, sixty-four ranked. The first run, on 7 October, asked fifty questions drawn by quotas; the second, on 8 October, seventeen added calculations. Three were dropped afterwards for their ambiguous scale: the two returns on assets and the retention ratio. All are listed, with their reason, on the benchmark page, together with the five failures counted against their model.

Some calls were replayed: right answers our extractor had failed to read, rate-limit refusals at a host, credit exhausted mid-run, a raised token ceiling for four models that think at length. The two runs cost about $12.6 charged, against $12.1 tracked by the pipeline: the gap comes from interrupted calls the provider bills without our recording them.

Latest published run of this benchmark
Run2026-10-07_analyse-financiere+2026-10-08_analyse-financiere
Date8 October 2026
Questions ranked64
Models27
Statusreal measurement
View the benchmark