anchor - the passage a stored value came from, found again

truecopy/anchor

A value read once and served for months cannot re-run its extraction per reader. What proves it is the sentence, found again from a short anchor kept beside the value. The figures the value announces are checked against that sentence.

A document read once, offline, into values a site then serves for months cannot re-run its extraction for every reader. So what proves such a value is not the extraction. It is the sentence - stored nowhere, found again at display time from a short anchor kept beside the value.

Nothing recopies the text, so the value cannot drift from it, and a revised document announces itself by the anchor no longer falling.

import { anchoredPassage, missingFigures, proofSpan } from 'truecopy/anchor';

The three calls, in the order they are used

const passage = anchoredPassage(article, 'hauteur maximale de 9', { depth: 'outline' });
// { heading: ['1.1- Dans une bande de 15 metres :'], line: '- zone UA : 9 metres.', opened: [] }

const quoted = [...passage.heading, passage.line, ...passage.opened].join(' ');

missingFigures(quoted, '9 m', { decimal: ',' }); // [] - the sentence carries what the value announces
proofSpan(quoted, '9 m'); // { start, end }: how far a reader travels to the proof

anchoredPassage returns null when the anchor no longer falls. That is the alarm and not an error: it is how a document revised under a stored reading announces itself, and the surface says so rather than displaying a value with no sentence.

A Passage has three parts, and the join is yours

interface Passage {
  readonly heading: readonly string[]; // the ancestors, outermost first
  readonly line: string; // the paragraph the anchor fell in
  readonly opened: readonly string[]; // the list the line opens, when it opens one
}

They are kept apart because the join is a display decision and the parts are not: a surface too narrow for the whole passage sheds ancestors from the outside in, and it can only do that if it knows which lines they are.

Each part is a defect already paid for. A paragraph quoted without its ancestors is exact and says nothing: -- de moins de 2 ans : 1 mois has lost - pour les ETAM, which is half the rule, and reads as everyone’s notice period. A heading quoted without the list it opens costs as much: l'employeur complete : followed by two items says that something is topped up and never what.

Only what the line opens comes with it, never its neighbour. Quoting the paragraph beside it puts the next sector’s twelve metres inside the citation of a nine, and nothing then says which of the two the displayed value proves.

PassageOptions: neither units nor depth is guessed

Both mistakes are silent. Reading depth off a source that does not mark it returns a passage with no ancestors - the exact failure this module exists to prevent - and returns it looking like a success.

units says where a paragraph ends.

  • lines (the default) trusts the line breaks the source printed. Anything already cut into paragraphs is this: marked-up text, a .docx, a text layer that kept its structure.
  • prose cuts a run-on blob itself, on numbering and on sentence ends. A regulation extracted from a PDF arrives as one paragraph of fourteen hundred characters, and looking for a full stop followed by a capital finds no cut at all in - zone UA: 9 metres. - secteur UAa: 12 metres.

depth says who is under whom.

  • dashes counts the hyphens the source puts in front of a line: one for a first level, two for a second, none for a top-level paragraph. A source that marks its own structure is read, never second-guessed.
  • outline (the default) infers it from typography, for a source that marks nothing: a bullet is a child of what precedes it, a number is as deep as it has components (1.2- is two), a line ending in a colon opens what follows, and a line ending in a full stop heads nothing.

The two are independent: a PDF text layer keeps its line breaks and marks nothing (lines + outline), while a blob may still carry hyphens.

foldAnchor: the shape an anchor is matched in

foldAnchor(text, options?); // the form both sides are compared in

Three tolerances, and each is a document that exists. Curly apostrophes and long dashes, because an extractor hands both forms back from the same document - one planning code writes l'unite fonciere and l'unite fonciere two lines apart - so an anchor demanding the exact sign would fall one time in two, and the page would report a perfectly present rule as unreadable. And runs of horizontal whitespace, because the cut inserts and swallows spaces.

Neither accents nor case, which is not the choice cite makes and is deliberate: an anchor is chosen by whoever compiled the value, out of the document itself, so it can be required to match letter for letter. Folding accents would widen a short anchor until it matches in two places, and the second place is another rule.

Line breaks survive under lines and not under prose, and that is the whole difference: where the source cut its own paragraphs the break carries the structure, and where it did not it is an artefact of the page width that would cut anchors in half.

missingFigures: the check an anchor cannot make

An anchor that still falls proves that the paragraph exists. It proves nothing about what the paragraph says.

Measured on one planning code: a rule displayed as 35 m de l'axe sat beside a sentence about the same road reading un recul minimum de 10 metres - anchor green, value invented, and no mechanical check saw it until this one.

missingFigures(source, value, options?); // the figures the value announces that the source does not carry

Empty means every announced figure is carried. A value announcing no figure at all - A l'alignement, Non reglemente, and that is a third of one measured corpus - announces nothing to miss, and is empty too: its proof is its anchor, not a figure.

FigureOptions is the localisation point

interface FigureOptions {
  readonly decimal?: DecimalMark | null;
  readonly spelled?: ReadonlyMap<string, string>;
}

decimal is the document’s own mark, from decimalMarkOf(document.text). Left out, each figure is read on its own, which is a guess and inherits a guess’s failures.

spelled is what a domain writes instead of digits. A regulation writes its small numbers out. Une place par logement does carry the 1 of 1 place, and ignoring that accuses the most faithful citation of the batch of missing its figure. Which words those are is a property of the language the document is written in, so it arrives from the caller - the way readDate takes its notation.

Masking a domain needs before its figures can be counted stays yours too: a planning code writes 50 m2, which announces fifty and not two. Mask without changing length and the extents stay indices into your own string.

proofSpan: a citation is bounded twice, and the two bounds do not talk

proofSpan(source, value, options?); // Extent | null

A citation is bounded by the string a check reads, and by the box a screen draws. Measured on one served corpus: of 610 values whose citation did carry the announced figure, 331 put it past the cut of a phone screen, and the gate called all 331 correct.

The distance to the proof is what you crop against. How many characters your own surface shows is yours, and a threshold measured on one screen is worth nothing on another.

Both ends are returned because both are used. The end says what the screen has to reach; the start says what a crop must not eat. Keeping only the end moves the start of 2 places par 50 m2 past its 2 - the figure is won on the screen and lost in the text.

Each figure counts at its first reading, the one an eye finds, and the worst of them decides: a citation showing the 1 of 1 place par 40 m2 without its 40 proves half of what it displays.

null twice over, and missingFigures tells the two apart: when the value announces no figure at all, and when it announces one the source does not carry. Neither has a span to point at, and treating either as a distance of zero would report a proof that is on screen because it does not exist.

Extent is not Span

interface Extent {
  readonly start: number;
  readonly end: number; // exclusive, so text.slice(start, end) is what was found
}

labels already exports a Span that carries the text it found. This one carries only the two indices, because what sits between them is a stretch of your own string.

Why this is not cite

cite answers the same question about a reading that happens now: a model points at the rows it read, and the lookup happens against those rows. This module answers it about a reading that happened months ago and cannot be repeated - so the anchor, and not the extraction, is what survives beside the value.

Two applications wrote this separately before it was here, on document families with nothing in common - a planning code printed as a PDF, a collective agreement published as marked-up text - and the second lost a lesson the first had already paid for. That is what says it belongs in a library.