ComingOn text
Meeting minutes
Do the minutes assign the right action to the right person?
The protocol for this benchmark is written. No model has been queried yet, so this hub shows no figure for this task.
What this test will measure
From a one-hour transcript, produce the decisions made and the actions with owner and deadline. A decision that was never made in the meeting is an invention, and counts as one.
What will be graded
- Decisions
- Actions and owners
- Faithfulness, nothing added
- Concision
What it is waiting for
No public reference says what the right answer is. A scoring rubric and human arbitration are needed before anything can be ranked.
- Place on the roadmap
- Wave 3
- Target dataset
- QMSum — 232 réunions transcrites et résumées par des humains ; la notation d'un résumé demande une grille
- Target sample
- 40 meetings
- What the model receives
- Text
One business function at a time, one task at a time. Every wave ends with a publication.
No figure is shown here because there is none. The day this task is measured, its ranking will appear on this very page.
How the hub measures — how an answer is verified, the rubric and the verdicts, on the task already measured.