cite - the rows a model read, and whether they carry the value

truecopy/cite

A model cannot produce a fact, only point at one. Number the rows, take back the numbers it cited, and look every value up in THOSE rows and nowhere else.

A model that turns cells into records fails one invisible way. It returns a value the document never printed - a town deduced from a postcode, a figure rounded on the way past - and that reads exactly like a successful extraction.

The deterministic answer is a protocol rather than a cleverer prompt, and this module is it. The model does not produce the value. It points at the row it read the value in, and the lookup happens here.

import { numberedRows, citedText, carriesText, carriesNumber } from 'truecopy/cite';

The four calls, in the order they are used

import { readTable } from 'truecopy/table';
import { decimalMarkOf } from 'truecopy/notation';

const { rows, document } = await readTable(file);
const decimal = decimalMarkOf(document.text);

// 1. What the model is given. The number IS the protocol.
const prompt = numberedRows(rows);
//    0 | 25 avenue Henri | 1 358 522 | 2019
//    1 | 4 rue des Lilas |   412 900 | 2021

// 2. What the model returns: records, each with the rows it read them in.
const records = await yourModel(prompt); // [{ street: '25 avenue Henri', price: '1 358 522', cited: [0] }]

// 3. The only source a value may be checked against.
for (const record of records) {
  const source = citedText(rows, record.cited);

  carriesText(source, record.street); // true
  carriesNumber(source, record.price, decimal); // true
}

Why the cited rows and not the document

Checking a value against the whole document was already possible, and it is the trap this closes. On a two-hundred-page document almost any figure is printed somewhere: a document-wide lookup therefore confirms values the cited rows never carried, and it looks strongest exactly where it protects least.

checkExtraction stays what it is - the arithmetic check against what the document declares about itself - and this check stands beside it, not instead of it. They fail on different things: one catches a reading that contradicts its own total, the other catches a value that came from nowhere in particular.

numberedRows(rows, first?)

One line per row: its number, then its cells, each separated from the next by a space, a pipe and a space. Numbering starts at 0 unless first says otherwise.

first is not decoration. A long table is read in overlapping windows, and a row must keep the same number in every window it appears in - renumbering each window from zero makes the same row cite differently in two prompts, and nothing downstream can tell the two apart.

citedText(rows, cited, first?)

The text of the rows a record cited, joined once. Pass the same first used to number them.

A cited number out of range contributes nothing rather than throwing. An invented row number is precisely the kind of claim this module exists to catch, and it is caught the ordinary way: every lookup against the missing text fails, and the caller counts it like any other refusal.

carriesText(source, value)

All the words of the value, in order - never one contiguous substring.

That is not a loosening for comfort, it is what a page does. A layout throws the tail of a name past the figure columns, so the name is printed there word for word and in order but not in one piece. Measured on a real property schedule, the contiguous form refused 104 of one document’s 111 records for that reason alone.

What it folds, and what it refuses to fold:

Folded Not folded
case accents
runs of whitespace digits
curly apostrophes (U+2018, U+2019) and U+02BC, to the straight one
hyphens, cut on both sides

Hyphens are cut because a hyphenated town arrives in two halves a whole row apart. Accents are not folded because two neighbouring towns differ by one, and digits are not folded because the rounded figure is the catch.

The guard still holds: every word must come from the cited rows, and an invented value has none of its words there. What it can no longer catch is a recombination of words all present and in order - a real risk, and a narrow one, the cited rows covering one record.

carriesNumber(source, value, decimal?)

On the numbers the page writes, never on flattened digits. Flattening the source into one digit soup lets almost any figure occur by accident: measured, a surface of 410 was found three times in a row that carries no 410 at all.

findNumbers is the same reader every other check here uses. It knows a thousands group from two numbers, and it never starts a match inside one.

Values are compared as read, with the document’s decimal mark on both sides:

carriesNumber('surface 1 655 m2', '1655', ','); // true   - one figure written twice
carriesNumber('total 2 415 065,40', '2 415 065', ','); // false  - the rounding this exists to catch

Pass the mark from decimalMarkOf(document.text). Left out, each figure is read on its own, which is a guess and inherits a guess’s failures.

Where to go next

  • Check what a model extracted - the arithmetic check, and why an empty list of problems is not a promise.
  • notation - findNumbers, decimalMarkOf, and how a page writes a figure.
  • table - where the rows come from, and measuredSpaces for a cell whose items must not run together before a model ever sees them.