Skip to content
AIWorkBench.fr

Business Index

Can one model hold every desk? The index is the mean of its accuracy scores per business function, every function weighing the same.

Updated 8 October 2026 · Index computed over 2 of 10 business functions and 2 of 22 benchmarks.

Key takeaways

  • DeepSeek V4.1 Flash leads the index with 100.0%, ahead of Kimi K2.6 (100.0%) and Qwen3.8 Max (0902) (99.7%).
  • Best open-weights model: DeepSeek V4.1 Flash, ranked 1 with 100.0%.
  • Best model from a French lab: Mistral Medium 3.5, ranked 20 with 82.4%.

The overall leaderboard

The index, then the score for each function. The best value in each column is highlighted.

Business Index
RankModelBenchmarks
1DeepSeek V4.1 Flash100.0%±0.0100.0%100.0%————————2 / 22
2Kimi K2.6100.0%±0.0100.0%100.0%————————2 / 22
3Qwen3.8 Max (0902)99.7%±0.6100.0%99.3%————————2 / 22
4Claude Opus 5.599.7%±0.6100.0%99.3%————————2 / 22
5Qwen3.8 Max Prime99.7%±0.6100.0%99.3%————————2 / 22
6GPT-6 Astra99.7%±0.6100.0%99.3%————————2 / 22
7Claude Fable 5.199.3%±0.9100.0%98.5%————————2 / 22
8Muse Spark 1.399.2%±1.598.4%100.0%————————2 / 22
9GLM 5.3 Flash98.9%±1.698.4%99.3%————————2 / 22
10Gemini 3.8 Flash98.9%±1.698.4%99.3%————————2 / 22
11Claude Sonnet 598.9%±1.698.4%99.3%————————2 / 22
12GPT-6 Luna98.5%±2.296.9%100.0%————————2 / 22
13GPT-6 Sol98.5%±1.798.4%98.5%————————2 / 22
14Qwen3.8 27B98.1%±2.296.9%99.3%————————2 / 22
15Grok 4.798.1%±2.296.9%99.3%————————2 / 22
16Qwen3.7 Flash97.3%±2.795.3%99.3%————————2 / 22
17Kimi K396.8%±3.093.8%99.8%————————2 / 22
18Command A+96.1%±3.392.2%100.0%————————2 / 22
19Gemini 3.5 Flash Lite89.5%±5.079.7%99.3%————————2 / 22
20Mistral Medium 3.582.4%±5.967.2%97.5%————————2 / 22

How the index is computed

For each business function, we average the model's accuracy over the benchmarks of that function it was able to take. The index is the mean of those scores: a function with three benchmarks weighs no more than one with two.

The index covers only the functions already measured, and the header says which. A function on the roadmap does not count as a zero — it does not count at all. While a single function is measured, the index is worth no more than the benchmark it comes from, and should be read that way.

A model that cannot read documents is scored, within each function, on text benchmarks only. Its index stays comparable but rests on fewer tests: the “Benchmarks” column says so.

The index says nothing about cost, response time or hallucinations. Those measures stay separate, in each benchmark: a model that leads on accuracy can be unusable because it makes things up.

Methodology