ComingOn text
Campaign analysis
Can the model read a dashboard without mixing up clicks and conversions?
The protocol for this benchmark is written. No model has been queried yet, so this hub shows no figure for this task.
What this test will measure
Analyse a multichannel campaign export and recommend where to move budget. Acquisition cost and return on spend are recomputed by hand; errors by a factor of ten are not rare.
What will be graded
- Reading the metrics
- Calculation accuracy
- Attribution
- Recommendations
What it is waiting for
No public dataset exists: these documents never leave the company. They will be collected from partners, with their consent, and only the scores will be published.
- Place on the roadmap
- Wave 4
- Target sample
- 30 reports
- What the model receives
- Text
One business function at a time, one task at a time. Every wave ends with a publication.
No figure is shown here because there is none. The day this task is measured, its ranking will appear on this very page.
How the hub measures — how an answer is verified, the rubric and the verdicts, on the task already measured.