Benchmark
Second measurement: 27 models face 64 financial-analyst questions
Seven models answer all sixty-four questions correctly, one of them at $0.001 per question. Eight others miss close to one in three or far more, nearly always on a calculation — and two invent an answer where there is none.