ComingOn documents
Expense report audit
Is a crumpled restaurant receipt enough to fool the model?
The protocol for this benchmark is written. No model has been queried yet, so this hub shows no figure for this task.
What this test will measure
Read photographed receipts, extract the amount and recoverable VAT, then check them against a given expense policy. The set includes disguised duplicates and out-of-policy spending.
What will be graded
- Receipt reading
- Recoverable VAT
- Policy compliance
- Duplicate detection
What it is waiting for
A real, annotated public dataset exists. It remains to be wired into the pipeline.
- Place on the roadmap
- Wave 2
- Target dataset
- CORD — 1 000 tickets de caisse photographiés, 30 sous-classes annotées, CC BY 4.0
- Target sample
- 60 receipts
- What the model receives
- Documents (images or PDF)
One business function at a time, one task at a time. Every wave ends with a publication.
No figure is shown here because there is none. The day this task is measured, its ranking will appear on this very page.
How the hub measures — how an answer is verified, the rubric and the verdicts, on the task already measured.