Skip to content
AIWorkBench.fr

Benchmark

Second measurement: 27 models face 64 financial-analyst questions

Seven models answer all sixty-four questions correctly, one of them at $0.001 per question. Eight others miss close to one in three or far more, nearly always on a calculation — and two invent an answer where there is none.

The hub publishes its second measurement, in a second business function: finance. Twenty-seven models answered the same sixty-four questions, asked on fifty-four reports from thirty US listed companies.

This is the first layer of a financial analyst's work: open a report at the right page and draw a metric or a judgement from it. The questions come from FinanceBench and were written by analysts, each with its answer and the page that proves it. Compute an EBITDA margin or a cash conversion cycle, say whether a margin is improving, name the segment that held growth back, recognise that a metric makes no sense for a bank.

What came out:

  • Seven models out of twenty-seven answer all sixty-four questions correctly. The cheapest, DeepSeek V4.1 Flash, costs $0.001 per question — thirty times less than GPT-6 Astra, which scores the same.
  • Eight models fail close to one question in three or far more, and it is the calculation that stops them. Nova Pro and Nova Lite get one calculation right out of thirty-two.
  • One question separates the top seven from the rest: Netflix's EBITDA margin, missed by seventeen models out of twenty-seven. Seven of them add content amortisation to the depreciation and amortisation the question asks for, and find 56.8% instead of 5.4%.
  • Two models invent. On the three questions whose right answer is “there is none”, Nova Pro fabricates two answers and Nova Lite one. The other twenty-five abstain.

What this ranking does not say: which of the top seven is best. They have the same score, and the table orders them only by price, then by speed. We tried to separate them on forty harder questions. The middle of the ranking changed with the way of counting, so we are not publishing that tie-break.

How an answer is verified: a number is right within 0.5% of the analyst's, a yes or no has to be the right one, a line item has to carry the name it has on the page. No AI grades an AI.

We corrected three things after reading the answers, and we say so. Three questions were dropped because their scale was ambiguous: the reference was a ratio, 0.01, while most models answered as a percentage, around 1.42, as our own instruction invited them to. The question measured the scale chosen, not the calculation. A list of accepted answers was too narrow and rejected twelve correct answers. And a rule removed a question from the test as soon as one model failed to answer it, which pushed out the hardest ones.

The limits of this measurement: the right page is supplied, whereas an analyst first has to find it in a report of several hundred pages. The companies are famous and the question set has been public since 2023: a model may know some of these figures by heart. Writing a note, forecasting or defending a view are not graded. And the human review of the three hallucinations has not been carried out yet.

View the benchmark — Financial report analysis