Benchmark8 Oct 2026
Second measurement: 27 models face 64 financial-analyst questions
Seven models answer all sixty-four questions correctly, one of them at $0.001 per question. Eight others miss close to one in three or far more, nearly always on a calculation — and two invent an answer where there is none.
- Seven models out of twenty-seven answer all sixty-four questions correctly. The cheapest, DeepSeek V4.1 Flash, costs $0.001 per question — thirty times less than GPT-6 Astra, which scores the same.
- Eight models fail close to one question in three or far more, and it is the calculation that stops them. Nova Pro and Nova Lite get one calculation right out of thirty-two.
- One question separates the top seven from the rest: Netflix's EBITDA margin, missed by seventeen models out of twenty-seven. Seven of them add content amortisation to the depreciation and amortisation the question asks for, and find 56.8% instead of 5.4%.