Skip to content
AIWorkBench.fr

Analysis

An invented figure costs more than a missing one

An empty box gets noticed and filled in. A plausible but wrong amount gets through review. That is why, for the hub, answering “nothing” when there is nothing is a right answer.

An invoice from a freelance designer under the small-business VAT exemption carries no VAT. No rate, no amount: the line “TVA non applicable, article 293 B du CGI” stands in for both. Ask a model for the VAT amount on that invoice. There are two ways to get it wrong, and they are not equal.

The first: returning nothing when there was something to read. The box stays empty in your software, someone notices, opens the invoice and fills it in. A minute lost.

The second: returning a figure when there was nothing. The model applies 20% to the net amount, because that is what one usually sees, and writes it down. The amount is plausible, neatly formatted, consistent with the rest. Nobody notices. It goes into the books, then into a VAT return.

That is why the hub uses three verdicts rather than two. A field is correct. Or it is wrong or missing. Or it is hallucinated: the model produced a value where the document contains none. Hallucinations are counted separately — the share of genuinely absent fields that the model filled in anyway — and are never melted into an overall score.

The consequence is built into the test set: a field absent from the document is part of the test. An invoice may quite legitimately carry no VAT, no due date, or no intra-EU VAT number. In those cases, answering “nothing” is not an admission of failure: it is the right answer, and it earns the same points as an amount read correctly.

The prompt sent to the models spells it out, ahead of any other instruction: if a piece of information is not on the document, do not guess, do not infer, do not fill in from what seems usual. A model that invents despite that has not been caught out by an ambiguous instruction.

The scoring code takes the same care in the other direction. A model that writes “non trouvé”, “aucun” or “n/a” is abstaining correctly: its answer is reduced to “nothing” before it is scored, rather than counted as an invented string.

For a business, the lesson goes beyond invoices. Before handing a document to a model, the useful question is not only “does it read well?” but “what does it do when there is nothing to read?”