ComingOn text
Ticket triage
Does the right ticket reach the right team, first time?
The protocol for this benchmark is written. No model has been queried yet, so this hub shows no figure for this task.
What this test will measure
Classify inbound requests by category and urgency, then route them. High volume, low unit cost: this is where a small model's accuracy-to-price ratio is judged.
What will be graded
- Category
- Urgency
- Language and sentiment
- Routing
What it is waiting for
A real, annotated public dataset exists. It remains to be wired into the pipeline.
- Place on the roadmap
- Wave 1
- Target dataset
- CFPB — la base ne fournit plus le texte des réclamations (constaté le 6 octobre 2026) : une source de tickets réels reste à trouver
- Target sample
- 200 tickets
- What the model receives
- Text
One business function at a time, one task at a time. Every wave ends with a publication.
No figure is shown here because there is none. The day this task is measured, its ranking will appear on this very page.
How the hub measures — how an answer is verified, the rubric and the verdicts, on the task already measured.