Article
The shape agreed in advance
A schema is written for a reader who cannot ask a follow-up question. That constraint, not tidiness, is why schemas exist -- and it explains why a false sentence in a document is not merely hard to find but structurally invisible.
what a schema is, why shape buys detectability, and what it costs · 2026-08-06 · 1 source
A will is the clearest case. It is a document whose author is, by the time anyone needs to read it, definitionally unavailable for comment. Every question a reader might have wanted to ask -- what did you mean by "personal effects", which grandchildren, in what proportion -- has to have been answered in advance, in the shape of the document itself. Wills are among the most heavily schematised texts people produce, and the reason is not ceremony. It is that the one person who could clarify is gone.
The same constraint, in weaker forms, produces nearly every schema in existence. A telegram is read by someone who cannot cheaply ask a follow-up. A sequence record deposited this year may be opened in 2060 by a researcher with no idea who deposited it. A laboratory notebook is read by a patent examiner a decade later, and a clinical protocol by a regulator who was not in the room. In each case the writer and the reader are separated -- by distance, by cost, by time, or by death -- and the shape has to carry what the conversation would have carried.
So, a definition worth being precise about:
A schema is the set of shapes a message is allowed to have, agreed before any message is sent.
It is not a template, which is a shape you are offered rather than held to. It is not a file format, which concerns how bytes are laid out rather than what they may say. It is not a validation rule, which is a check performed after the fact on something already written. A schema is a constraint on the space of possible messages, held in common by both ends, and it does its work before anyone writes anything at all.
That last property is the one that pays, and the payment is larger than tidiness.
Why shape buys detection
In April 1950 Richard Hamming published a paper in the Bell System Technical Journal that is filed under coding theory and belongs, on the evidence of its first page, under epistemology.1Hamming, R. W. (1950), "Error Detecting and Error Correcting Codes," Bell System Technical Journal 29(2), pp. 147-160. DOI 10.1002/j.1538-7305.1950.tb00463.x; full text at the Internet Archive. All three passages quoted above are from the paper's opening pages; the second and third are given verbatim. (T1 -- primary source, publicly retrievable) He states the problem plainly: he was led to it "from a consideration of large scale computing machines in which a large number of operations must be performed without a single error in the end result."
Then the observation that matters most here. Comparing a computer with a telephone system, where a moment of noise is merely an annoyance, he writes that in a digital computer
"a single failure usually means the complete failure, in the sense that if it is detected no more computing can be done until the failure is located and corrected, while if it escapes detection then it invalidates all subsequent operations of the machine."
Note the asymmetry, because everything below depends on it. A detected error is merely expensive: work stops until someone fixes it. An undetected error is categorically worse, because every result computed downstream inherits it silently and looks exactly as sound as it did before. Detection is therefore a different problem from correctness, and a more urgent one.
Hamming's mechanism is geometric, and once seen it is difficult to unsee. Picture every possible message as a point in a space. The messages you have agreed to call legal -- his code points -- are some subset of those points. His condition for detecting a single error reads straight off the picture:
"If all the code points are at a distance of at least 2 from each other, then it follows that any single error will carry a code point over to a point that is not a code point ... This in turn means that any single error is detectable."
Read that again with the weight on not a code point. The error is detectable because of where it lands. Nothing about a corrupted message announces itself; what announces it is that the agreed shape has no room for it.
Which yields a first principle, stronger than it first appears:
detection is possible only if some messages are illegal
If every string is a legal message, no string is a detectable error. This is not a statement about how hard you are looking. It is a statement about what there is to find. Detection is bought with redundancy -- by spending part of the message space on being unusable, so that corruption has somewhere illegal to land.
Figure 1. Hamming's argument with one axis of abstraction removed. The
displacement is identical in both panels -- the same error, the same magnitude. Only
the legal set differs. On the left it lands on another legal point and passes; on
the right it lands in a gap, and the gap is the signal. A schema does not prevent
the error. It arranges for the error to have nowhere innocuous to go. Drawn by
fig-1.py.
A useful sanity check on how cheap this can be. A single parity bit -- one extra bit appended so the number of ones is always even -- makes exactly half of all possible strings illegal, and thereby makes every single-bit error detectable. One bit of redundancy, half the space forfeited, a whole class of error rendered visible.
The date that cannot be wrong
Consider 03/04/26.
In the United States that is 4 March 2026. In much of Europe it is 3 April 2026. In a system using the Japanese convention it may be 26 April 2003. The string is not wrong. It is worse than wrong: there is no observation that could contradict it, because it never committed to anything specific enough to be contradicted. It is unfalsifiable, in the sense that makes the word an insult rather than a compliment.
Write the same instant as 2026-03-04 and it becomes capable of being wrong. That
is not a small change. A claim that cannot be wrong cannot be checked, and a claim
that cannot be checked is not doing any work.
This is the general upgrade a schema performs, and it is worth stating as a rule: a good schema converts ambiguity into falsifiability. It does not make claims true. It makes them the kind of thing that could be found false, which is the precondition for anyone ever finding out.
Prose is the space where everything is legal
Now turn Hamming's condition on natural language.
Every grammatical sentence is a legal message. There is no string a competent reader would reject on the grounds that the language has no room for it. The legal set is, for practical purposes, the whole space -- the left panel of the figure.
It follows, with unpleasant directness, that a false claim in a document is not hidden. It is structurally indistinguishable from a true one. It has the same shape, the same grammar, the same confident cadence. There is no anomaly to notice, because an anomaly is a departure from an expected form and prose expects every form.
This is why "read it more carefully" does not scale, and why more careful reading by more senior people does not fix it either. Searching a haystack for a needle is hard but tractable, because the needle differs from the hay. Searching a stack of needles is a different problem, and diligence is not the missing ingredient.
And then Hamming's asymmetry does its work: whatever rests on the undetected claim inherits it, silently, looking exactly as sound as everything around it.
What a claim looks like once it has a shape
So give a claim a shape. The minimum that does real work turns out to be four fields, and each one exists to make a specific family of message illegal.
| field | what it holds | what it makes impossible to say |
|---|---|---|
proposition | one claim, with one truth-maker | a compound assertion that can be half true |
locator | where it came from, content-addressed | a source that cannot be resolved or was quietly edited |
quote | the source's own words | a paraphrase that drifted from what was actually said |
at | when it held in the world | a fact that was true once and is asserted in a tense |
Admission is then a conjunction, and each clause carves away another region of the space:
admit(f, c) iff prov(locator(f)) >= floor(c)
and every number, date and negation in prop occurs in quote
and no judge has rejected f
and f has not expired
The second clause is worth pausing on because of what it is not. It is not a judgement about meaning; it is a literal occurrence check. If the claim says 42 per cent, the quote must contain 42 per cent. If the claim says the endpoint was not met, the negation must be there in the source's own words. A machine can run this, it never guesses, and it rejects rather than interprets.
The locator carries the most weight. It is content-addressed: the address of a source is derived from the bytes of the source. Alter the source and the address stops resolving. The failure mode is the right way round -- a tampered source does not quietly admit a false claim, it fails to resolve at all. The system's response to corruption is silence rather than a confident wrong answer.
And there is no field anywhere in which an author writes down how good the evidence is. The grade is computed from what the locator is. There is nothing there to forge, which is a stronger property than a rule against forging it.
What it costs
This is not free, and a piece that pretended otherwise would be selling something.
A closed vocabulary means a new kind of claim is a governed change rather than a keystroke. Every source needs an address before it can be cited, which front-loads work that prose lets you defer indefinitely -- and "indefinitely" is doing real labour in that sentence, since the deferred work is usually never done at all. Writing a fact this way is slower than asserting it. It is slower in the specific way that a lab notebook is slower than remembering.
There is a reason schemas of this kind have historically been the output of committees, and have therefore existed only where a committee could be justified: standards bodies, large institutions, funded consortia. The construction cost, not the idea, was the barrier.
And an honest residue. The whole structure rests on the assumption that you cannot find two different sources with the same address. That is a cryptographic assumption, not a logical impossibility, and it is the sort of thing that should be said out loud -- a schema that oversells its guarantees is one you will keep trusting past the point where it has earned it.
What this has to do with a diligence report
Which returns to the will.
A due-diligence report is read by someone who cannot ask the analyst. They are reading it months later, or after the analyst has moved on, or in a meeting where the analyst is not present and the decision is being made anyway. The reader is separated from the writer in precisely the way that produces schemas in the first place.
So the report has to answer, in advance and in its shape, every question a careful reader would have asked. Where did this number come from. What were the source's own words. When did it hold. What would falsify it. That is what the citation apparatus on these reports is for: not decoration, and not a display of rigour, but the pre-recorded half of a conversation that cannot happen.
The practical difference this makes is smaller and more useful than it sounds. "Which claims here rest only on the company's own statements?" becomes a question I can run, rather than an audit I have to perform and might perform inconsistently at the end of a long week. Unsourced is a query. That is the whole return on the structure.
It is worth being equally clear about what none of this buys. It says nothing about whether a claim is true. Hamming's codes never made a computation correct; they made corruption detectable, and left correctness to the people. A schema does the same and no more. It makes one class of error visible, cheaply and mechanically, so that attention is free for the part that was always going to require judgement.
That part does not get automated. It gets defended.
Footnotes
-
Hamming, R. W. (1950), "Error Detecting and Error Correcting Codes," Bell System Technical Journal 29(2), pp. 147-160. DOI 10.1002/j.1538-7305.1950.tb00463.x; full text at the Internet Archive. All three passages quoted above are from the paper's opening pages; the second and third are given verbatim. (T1 -- primary source, publicly retrievable) ↩