Business Index
Can one model hold every desk? The index is the mean of its accuracy scores per business function, every function weighing the same.
Updated 8 October 2026 · Index computed over 2 of 10 business functions and 2 of 22 benchmarks.
- DeepSeek V4.1 Flash100.0%
- Kimi K2.6100.0%
- Qwen3.8 Max (0902)99.7%
- Claude Opus 5.599.7%
- GPT-6 Astra99.7%
- Muse Spark 1.399.2%
- GLM 5.3 Flash98.9%
- Gemini 3.8 Flash98.9%
- Grok 4.798.1%
- Command A+96.1%
- Mistral Medium 3.582.4%
- Nova Lite 1.057.0%
Each lab's best model, ranked by Business Index.
View full resultsKey takeaways
- DeepSeek V4.1 Flash leads the index with 100.0%, ahead of Kimi K2.6 (100.0%) and Qwen3.8 Max (0902) (99.7%).
- Best open-weights model: DeepSeek V4.1 Flash, ranked 1 with 100.0%.
- Best model from a French lab: Mistral Medium 3.5, ranked 20 with 82.4%.
The overall leaderboard
The index, then the score for each function. The best value in each column is highlighted.
| Rank | Model | Benchmarks | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | DeepSeek V4.1 Flash | 100.0%±0.0 | 100.0% | 100.0% | — | — | — | — | — | — | — | — | 2 / 22 |
| 2 | Kimi K2.6 | 100.0%±0.0 | 100.0% | 100.0% | — | — | — | — | — | — | — | — | 2 / 22 |
| 3 | Qwen3.8 Max (0902) | 99.7%±0.6 | 100.0% | 99.3% | — | — | — | — | — | — | — | — | 2 / 22 |
| 4 | Claude Opus 5.5 | 99.7%±0.6 | 100.0% | 99.3% | — | — | — | — | — | — | — | — | 2 / 22 |
| 5 | Qwen3.8 Max Prime | 99.7%±0.6 | 100.0% | 99.3% | — | — | — | — | — | — | — | — | 2 / 22 |
| 6 | GPT-6 Astra | 99.7%±0.6 | 100.0% | 99.3% | — | — | — | — | — | — | — | — | 2 / 22 |
| 7 | Claude Fable 5.1 | 99.3%±0.9 | 100.0% | 98.5% | — | — | — | — | — | — | — | — | 2 / 22 |
| 8 | Muse Spark 1.3 | 99.2%±1.5 | 98.4% | 100.0% | — | — | — | — | — | — | — | — | 2 / 22 |
| 9 | GLM 5.3 Flash | 98.9%±1.6 | 98.4% | 99.3% | — | — | — | — | — | — | — | — | 2 / 22 |
| 10 | Gemini 3.8 Flash | 98.9%±1.6 | 98.4% | 99.3% | — | — | — | — | — | — | — | — | 2 / 22 |
| 11 | Claude Sonnet 5 | 98.9%±1.6 | 98.4% | 99.3% | — | — | — | — | — | — | — | — | 2 / 22 |
| 12 | GPT-6 Luna | 98.5%±2.2 | 96.9% | 100.0% | — | — | — | — | — | — | — | — | 2 / 22 |
| 13 | GPT-6 Sol | 98.5%±1.7 | 98.4% | 98.5% | — | — | — | — | — | — | — | — | 2 / 22 |
| 14 | Qwen3.8 27B | 98.1%±2.2 | 96.9% | 99.3% | — | — | — | — | — | — | — | — | 2 / 22 |
| 15 | Grok 4.7 | 98.1%±2.2 | 96.9% | 99.3% | — | — | — | — | — | — | — | — | 2 / 22 |
| 16 | Qwen3.7 Flash | 97.3%±2.7 | 95.3% | 99.3% | — | — | — | — | — | — | — | — | 2 / 22 |
| 17 | Kimi K3 | 96.8%±3.0 | 93.8% | 99.8% | — | — | — | — | — | — | — | — | 2 / 22 |
| 18 | Command A+ | 96.1%±3.3 | 92.2% | 100.0% | — | — | — | — | — | — | — | — | 2 / 22 |
| 19 | Gemini 3.5 Flash Lite | 89.5%±5.0 | 79.7% | 99.3% | — | — | — | — | — | — | — | — | 2 / 22 |
| 20 | Mistral Medium 3.5 | 82.4%±5.9 | 67.2% | 97.5% | — | — | — | — | — | — | — | — | 2 / 22 |
How the index is computed
For each business function, we average the model's accuracy over the benchmarks of that function it was able to take. The index is the mean of those scores: a function with three benchmarks weighs no more than one with two.
The index covers only the functions already measured, and the header says which. A function on the roadmap does not count as a zero — it does not count at all. While a single function is measured, the index is worth no more than the benchmark it comes from, and should be read that way.
A model that cannot read documents is scored, within each function, on text benchmarks only. Its index stays comparable but rests on fewer tests: the “Benchmarks” column says so.
The index says nothing about cost, response time or hallucinations. Those measures stay separate, in each benchmark: a model that leads on accuracy can be unusable because it makes things up.