readTable - the two-line path
truecopy/table
It gets rows and cells out of a file with no configuration at all, plus the field that says what the reading could not vouch for.
import { readTable } from 'truecopy/table';
const { rows, warnings, findings, pages, boundaries, document } = await readTable(file, options?);
| Field | |
|---|---|
rows |
string[][]: every row of every page, cut into cells |
warnings |
string[]: what could not be vouched for |
findings |
Finding[]: the same doubts, each with a code to act on |
pages |
string[][][]: the same rows kept in their pages, like boundaries |
boundaries |
number[][]: the cut used, one list per page |
document |
the Document, for a reader that needs more than the cells |
options is a TableOptions: everything OpenOptions takes - caps, deadline, workerSrc, pdfjs - plus the two knobs below, which belong to the cut rather than to the file.
The two knobs of the cut
Neither changes a default, and that is deliberate: a change in how a page is cut never arrives as a new default, it arrives as an option a reader adopts once its own bench is green.
| Option | Default | What it does |
|---|---|---|
rightEdges?: boolean |
on | reads right edges inside a band, cutting flush-right figure columns apart |
measuredSpaces?: boolean |
off | spaces the items of a cell as the page printed them, instead of one space apart |
rightEdges: false restores the cut exactly as it was before right edges were read at all. It exists for compatibility rather than tuning. A finer cut is not free for every reader: one that learned the wider cells - splitting a shared cell itself, on a page printing two tables side by side - loses the row it relied on when a cut lands inside it. Measured on a fund report of that shape, six sections that closed on the wider cells stopped closing on the finer ones, the row carrying one position of each table coming back as neither. It is not a per-document setting to hunt with: a reading that only lands with it off is a reader built on the wider cells, not a better cut.
measuredSpaces: true lets the gap on the page decide. An engine breaks a cell into items wherever glyph spacing changes, so one space between all of them turns 1 207 773 393 into 1 2 07 773 393 - a number no reader parses and no check catches, every digit being there in the right order. Measured on that filing: 0,00 point where glyphs touch, 3,6 to 3,8 where a thousands separator or a word space stands, on characters 6,7 points wide. The cut is a quarter of a character, which clears both by a wide margin and, unlike a fixed number of points, does not depend on the font size.
No width is no measurement: a row assembled without geometry - a paste, a CSV, an OCR line handed over whole - keeps its one space between items, having no gaps to read. document.text is not touched either way; the option moves cell text only.
Whatever was dropped, the same two lines
file is a PDF, a table pasted out of one, a CSV, a TSV, an OCR page. One engine reads all of them, because the cut votes on which left edges come back row after row, and that question does not care what measured them. A document only has to say where its fields start.
| What was dropped | Where a column starts | What the reading prints |
|---|---|---|
| a PDF | the item’s x, in points | cut at 100, 290 |
| a table pasted with spaces | the character the field starts at | cut at characters 7, 22 |
| a file written with a delimiter | the field’s index: a CSV lines nothing up | cut on the delimiter |
| prose, or a file that quotes its fields | nowhere | every row came back whole |
Tab, semicolon, pipe and comma, in that order, and only when the count recurs line after line. What makes a delimiter real is what makes a column real.
Two refusals come with that, and both are the point rather than a gap:
- A file that quotes its fields is not split. A quoted field may hold the delimiter itself, and splitting anyway shifts every column after it. Half-parsing a CSV is exactly the plausible-but-wrong reading this library argues against, so the rows come back whole with a warning.
- An aligned paste is not read as a CSV because its amounts hold commas. Every French amount does, so counting commas finds exactly one on every line of a pasted statement, and cutting on it would turn
12,40into two columns. Alignment wins; only a tab overrides it, a tab never being punctuation.
warnings is the whole point
Whoever types “extract tables from pdf” wants rows, and making them write a kindOf and choose thresholds before the first result is twenty minutes most people do not spend. So this is two lines.
What it does not do is return a bare array. A plausible-looking table out of a document that was misread is worse than nothing, because nobody checks a table that looks right. Hand one over with no signal and this library becomes the thing it was written to prevent, and a worse one than the extractors that have spent years tuning their heuristics.
An empty
warningsis not a promise. It says nothing looked wrong from the shape of the page. That is a much smaller claim than “this reading is right”, and the distance between the two is what the rest of the library is for.
Every warning is computed without knowing anything about your document:
page N carries no text at all: a blank page, a scan, an image;page N shows no column at all: the page is prose, or the cut failed and every row came back whole;column C of page N is filled on only X%: the cut may have invented that column out of a letterhead;the pages disagree on how many columns there are: usually a different table, and joining them makes a third that is neither;column C of page N holds two values on X% of its rows: the cut missed a boundary and two neighbouring columns landed together.
That last one is the opposite failure to a thin column, and the dangerous one. A column the cut never separated is filled on every row, exactly like a good one, so no fill rate can see it: the reading is wrong and silent about being wrong. Read as one figure, one such column produced 97 wrong values out of 162 on a real property schedule.
It only fires when a cell certainly holds two values: nothing but numbers and separators, and every number carrying its own decimal mark. Without that second condition two integers separated by a space are indistinguishable from one number, a space being exactly what French notation puts between thousands - measured, the loose rule fires on 48% of a column that is perfectly well cut, and the strict one on none of it.
findings, for whatever has to act on one
A sentence is for whoever reads it; a code is for whatever acts on it. warnings is derived from findings, so a message and a code can never say different things.
const { findings } = await readTable(file);
// { code, message, page?, column?, shareFilled?, shareDoubled? }
if (findings.some((f) => f.code === 'blank-page')) offerTheScanRoute();
The Doubt codes are the five warnings above: blank-page, no-column, thin-column, merged-column, pages-disagree. Same reason UnreadableDocument names its reason - a program, or an agent, that has to branch on a doubt should not be matching English prose to do it, and a message rewritten for clarity should not break it.
pages, when the flat list is not enough
rows is exactly pages.flat(), and it is the field most readings want. But flattening loses which page a row came from, and some documents cannot be read without it: a page that prints two tables side by side carries two runs of headings, and walking the rows in order alternates between them - a position silently inherits the heading of the other column. pages[i] and boundaries[i] are one page.
boundaries, and why they are not the page’s own
readTable does not cut on page.columnBoundaries. That one is the spread of x over everything printed (letterhead, address block, footer), and on a real page it proposes columns the table never had: measured, twelve where the table has five.
It uses boundariesFromRecurrence instead: keep only the x that come back row after row, because a real column’s left edge recurs and a word in the middle of a description does not.
boundaries is handed back so you can explain or re-cut the same page against the same lines:
import { explainDocument } from 'truecopy/explain';
const { document, boundaries } = await readTable(file);
console.log(
explainDocument(document, {
boundariesOf: (page) => boundaries[page.pageNumber - 1]
})
);
Without that, the heading and the cells would tell two different stories.
When two lines are not enough
The moment something downstream acts on the rows (a budget, a report, a decision), you want a reading that checks itself rather than one that merely looks fine:
findRowAnomaliesdrops the rows that break the table’s own shape;validatesays whether enough well-formed records came back;readDocumentrefuses to return a reading that contradicts its document.