The data that is not a number

A corpus that was not generated

Every collection in five strands has been copies of a generated text with a fraction of its characters replaced — three numbers, one dial. Here is one that was not — twelve essays, eight source modules, ten revisions of one file — measured beside the model of it.

Three essays in a row ended with the same sentence, and it was the oldest thing on this collection’s list of what was missing.

Every collection here is copies or near-copies of a generated text, and the properties the document strand is about — how many generations of copying, how many documents, what alphabet — are exactly the ones a synthetic corpus does not settle.

This is a corpus that was not generated, and the first four measurements taken on it.

Two measures of repetition, on real text and on the model of itRuns of the transform and phrases of the parse, both divided by the length of the text so that three parts of different sizes can sit on one plate. The version history is the only part that is repetitive in the sense this field means: 0.126 runs a character against 0.491 for twelve unrelated essays. The generated collection at the setting used throughout this collection lands near the version history on runs and well below it on phrases, and nowhere near either of the other two — which is the whole question this strand was written to ask, drawn once before it is answered properly.twelve essays — runs0.491generated, matched0.109twelve essays — phrases0.213generated, matched0.046eight source modules — runs0.461generated, matched0.122eight source modules — phrases0.210generated, matched0.050ten revisions of one file — runs0.126generated, matched0.113ten revisions of one file — phrases0.067generated, matched0.047per character24,576 characters a partdivergence 0.02
Fig. 1 Two measures of repetition on three real collections, each with a generated collection matched to it in documents and in length.

What the generator does

One line and one dial. Take a text; make k copies; replace a fraction of each copy’s characters with random ones. A collection is fixed by three numbers — the base length, the number of copies, and that fraction — and every measurement in five strands has been taken on one.

It is exact, it is reproducible, and it made every figure in those strands possible. Nothing below is a criticism of it. What follows is a question about what its numbers mean.

The three parts, and why there are three

The corpus is frozen: written down once and never re-read from anywhere. A measurement about a moving collection is true on the day it is taken and unreproducible after, and this collection’s rule is that a number in an essay is a number a reader could get back.

Its three parts repeat for three different reasons, which is the whole reason there are three.

Prose — twelve essays. Repeats a vocabulary: the same few thousand words in different arrangements, with almost no long repeated stretch. Twelve documents, 144,606 characters, 95 distinct symbols.

Source — eight modules of the program that draws this collection’s figures. Repeats a form: long identifiers, boilerplate lines, the same indentation on every line of every file. Eight documents, 106,276 characters, 111 symbols.

Revisions — ten successive versions of one file, oldest first, as they were written. Repeats almost everything, with localised changes. Ten documents, 75,649 characters, 79 symbols. This is the case every compressed index in this collection is built for, and the only one of the three whose documents stand in a known relation to each other.

Real documents are not the same lengthEach bar is one document. The version history's ten revisions run from 1,086 to 4,079 characters and grow monotonically, because a file that is edited across ten versions mostly gets longer. The eight source modules vary by 3.12x. The twelve essays vary by only 1.26x, and that is the one uniformity here that is genuinely an artefact: they come from a collection that requires every essay to sit inside a length band. A generated collection matched to any of them has 12 documents of exactly 2,048 characters each, every time.twelve essaysrepeats a vocabulary1.26xeight source modulesrepeats a form3.12xten revisions of one filerepeats almost everything3.76xgenerated, matchedone length by construction1.00xdocument length · the ratio of longest to shortest at the right24,576 characters a part3.76x at the widest
Fig. 2 The documents of each part, drawn as lengths, with a matched generated collection below them. The generated one has documents of exactly one length, every time.

What made it measurable

The measure this strand turns on is the phrase count of the parse, and computing it by comparing characters is quadratic. That is why every measurement in this collection before now was taken on a few thousand characters, and it is why a corpus of a hundred and forty thousand was not a thing anybody could weigh.

The parse in one pass of the text is the construction that removed it, and it landed one strand before this one. The whole of this corpus is parsed in the time one of those small texts used to take.

That is worth stating because it is the reason this strand exists now rather than long ago. The instrument came first and the question was already there.

Why three parts rather than one

The temptation with a real corpus is to take one and call it real. Three were taken because the word “real” hides a distinction this collection needs.

A compressed index’s whole argument is that a repetitive collection is small. Repetitive is a property, not a synonym for real, and most real text is not repetitive: twelve unrelated essays share their vocabulary and almost nothing else. If the corpus had been prose alone, every structure in these strands would have looked bad on it and the conclusion would have been that the structures do not work on real data — which is false, and would have been an artefact of choosing the one real case they are not for.

Three parts spanning the range make the range visible instead. Prose is the case the structures are not for; a version history is the case they are for; source is in between and is the one real collections most often actually look like.

Two measures of repetition, on real text and on the model of itRuns of the transform and phrases of the parse, both divided by the length of the text so that three parts of different sizes can sit on one plate. The version history is the only part that is repetitive in the sense this field means: 0.143 runs a character against 0.527 for twelve unrelated essays. The generated collection at the setting used throughout this collection lands near the version history on runs and well below it on phrases, and nowhere near either of the other two — which is the whole question this strand was written to ask, drawn once before it is answered properly.twelve essays — runs0.527generated, matched0.106twelve essays — phrases0.244generated, matched0.051eight source modules — runs0.499generated, matched0.118eight source modules — phrases0.243generated, matched0.054ten revisions of one file — runs0.143generated, matched0.111ten revisions of one file — phrases0.084generated, matched0.052per character12,288 characters a partdivergence 0.02
Fig. 3 The same three parts at half the size. The ordering between them is what this plate is for, and it does not move.

The first measurement

Runs of the transform and phrases of the parse, both divided by the length of the text so that three parts of different sizes sit on one plate.

part symbols runs a character phrases a character
twelve essays 87 0.491 0.213
eight source modules 98 0.461 0.210
ten revisions 75 0.126 0.0668
generated, matched 22 0.109–0.127 0.046–0.050

Only the version history is repetitive in the sense this field means. Twelve unrelated essays have four times the runs per character that ten revisions of one file do, and prose and source are indistinguishable from each other on both measures.

That is not surprising and it is worth having measured, because the strands built on compressed indexes have been measuring collections whose repetitiveness was set by a dial, and this says which of the three real cases the dial has been imitating.

The alphabet, which is the largest difference

A generated text here draws from 22 symbols. Real prose has 87, real source 98, a version history 75.

Punctuation, digits, capitals, brackets, quotation marks, the characters nobody thinks of as part of a text. Between three and four and a half times the alphabet, on every part.

The alphabet is where the two are least alikeA generated text here draws from an alphabet of 22 symbols. Real prose has 87, real source 98, and a version history 75 — punctuation, digits, capitals, brackets, the characters nobody thinks of as part of a text. It matters more than it sounds, because the alphabet sits inside a logarithm in every size this field measures, and because the document strand's headline about separators turns on whether adding one crosses a power of two. At 22 symbols it does; at 87 nothing near it does.twelve essays877 bits/chargenerated, matched225 bits/chareight source modules987 bits/chargenerated, matched225 bits/charten revisions of one file757 bits/chargenerated, matched225 bits/chardistinct symbols24,576 characters a part4.45x at the widest
Fig. 4 The symbol counts, real against generated. The right-hand column is what each costs to pack, and the crossing of a power of two is where a whole strand’s headline lives.

It matters more than it sounds, for two reasons.

The alphabet sits inside a logarithm in every size this field measures — a packed text is n⌈log₂ σ⌉ bits, a wavelet tree is ⌈log₂ σ⌉ levels deep, a backward search step costs a symbol rank whose price is the code length. Four times the alphabet is about two extra bits a character everywhere.

And a headline in the document strand turns on whether adding a symbol crosses a power of two. At 22 symbols it does; at 87 nothing near it does. Documents that are not the same length is where that comes apart.

What is frozen, and how

Twelve slugs, eight module names and ten dated revisions, listed explicitly and written out in full. Not a rule for selecting documents — a list, so that adding an essay to this collection changes nothing here.

Truncation, where it is needed, scales every document by the same factor rather than cutting each to a common length. That is a deliberate choice and it is what documents that are not the same length is about: a corpus whose documents are all cut to one length is a corpus whose length distribution has been thrown away, and the distribution is one of the things this strand exists to measure.

Every document is kept whatever the factor, for the same reason.

What a real collection looks like in the small

The revision history is the part most worth looking at, because it is the case the compressed-index strands are built for and the one the generator claims to model.

Ten revisions, from 1,086 characters to 4,079, growing monotonically — a file that was edited across ten occasions and mostly got longer. The generated collection matched to it has ten documents of exactly 2,458 characters each.

That difference is not cosmetic. Several of the real pairs differ in nothing at all over their common prefix and simply have a few hundred characters appended; none of the generated pairs does, because the generator’s model of an edit is a substitution and a substitution does not change a length.

Where the changes are, not how many of them there areEach row is one consecutive pair of documents; the bar is the fraction of positions that differ over their common length, and the number beside it is how many characters a run of differing positions holds. Real revisions average 7.3 characters a block and generated ones 1.04 — a scattered substitution is a block of one, by definition. Several real pairs differ in nothing at all over their common prefix and simply have 333 characters appended, which the generator never does: its documents are all one length. And the large rates are an insertion seen positionally — everything after an inserted paragraph is at the wrong offset, so a comparison that only knows about substitutions reports most of the file as changed, which is exactly the model the dial implements.119.3 a block217.3 a block3nothing changed415.1 a block5nothing changed6nothing changed7nothing changed8nothing changed914.1 a blockfraction of positions differing · the thin bar is the generated pairreal 7.3 characters a block · generated 1.04ten revisions of one file · 10 documents7.05x apart
Fig. 5 Where the changes are, pair by pair. Real revisions change nothing, or a great deal; generated ones change three per cent of everything, every time.

What the numbers mean, in units a reader can hold

Runs per character and phrases per character are compact and abstract, so it is worth converting them once.

A version history of 24,580 characters has 3,103 runs in its transform and 1,642 phrases in its parse. That means an index proportional to the run count is holding about three thousand things where the text has twenty-four thousand characters — a factor of eight — and a phrase index is holding about sixteen hundred.

Twelve essays of the same length have 12,065 runs and 5,234 phrases: a factor of two, not eight. An index proportional to either measure is barely smaller than the text it indexes, which is an index larger than what it indexes in its natural habitat.

The structures work, on the collection they are for, and the collection they are for is a version history. That is the plainest thing this corpus says.

The measures do not see the order

One thing that might have been expected and is not there.

A version history has an order, and the intuition is that a compressed index is small on it because consecutive versions sit next to each other. Shuffle the documents into an order they were never in — keeping every document, every character, every symbol — and the measures move by 0.06% in runs and 0.4% in phrases.

They do not see it. Both measures find a previous occurrence wherever it is, and every version of this file resembles every other, so which one a phrase copies from moves and how many phrases there are does not.

That matters because it is the assumption under every collection in five strands: the generated ones are exchangeable copies, and this says the real one is exchangeable too, for the two quantities being measured. It is the first thing this corpus has confirmed rather than contradicted.

Where the changes are, not how many of them there areEach row is one consecutive pair of documents; the bar is the fraction of positions that differ over their common length, and the number beside it is how many characters a run of differing positions holds. Real revisions average 16.5 characters a block and generated ones 1.04 — a scattered substitution is a block of one, by definition. Several real pairs differ in nothing at all over their common prefix and simply have 879 characters appended, which the generator never does: its documents are all one length. And the large rates are an insertion seen positionally — everything after an inserted paragraph is at the wrong offset, so a comparison that only knows about substitutions reports most of the file as changed, which is exactly the model the dial implements.115.3 a block214.4 a block316.3 a block415.6 a block515.8 a block616.6 a block721.7 a blockfraction of positions differing · the thin bar is the generated pairreal 16.5 characters a block · generated 1.04eight source modules · 8 documents16x apart
Fig. 6 The same comparison on the source part, whose documents are unrelated modules rather than versions of one file. Nearly everything differs, and the blocks are still long.

The control that must fail

A measure of repetition has to be shown to be measuring repetition, and the sharpest available control is a shuffle of the characters.

Shuffling a text’s characters keeps its length, its alphabet and every symbol’s frequency exactly — so its zero-order entropy is identical to the last bit — and destroys every structure above that.

Measured on the revision history: the run count multiplies by 7.42 and the phrase count by 6.09, with the zero-order entropy unchanged.

That is the check that licenses everything else here. The entropy that cannot see a copy is the essay that first made the point in this collection — a text and its shuffle have the same entropy at every order the text is long enough to estimate — and this is the same demonstration on data nobody constructed.

Why a frozen corpus rather than a live one

There is an obvious alternative — read the collection as it stands whenever a figure is drawn — and it was rejected for a reason worth recording.

Every number in this strand would then be a number about whatever the collection held on the day, and no reader could reproduce one. Worse, the numbers would drift: this collection grows by twenty essays at a time, so the prose part’s alphabet, length distribution and repetition would all move, and an essay’s claim about a factor of four would become a claim about whatever factor happened to hold later.

A frozen corpus is a fixed object. Its provenance is stated, its documents are listed, and its measurements are reproducible in the only sense that matters — the same input produces the same number.

The cost is that it ages. In a year the source part will describe a program that has changed and the revision history will stop at revision ten. That is the right cost to pay: an aged measurement that is reproducible is worth more than a current one that is not.

The count transfers exactly, and the share does notA concatenation of d documents has exactly (m − 1)(d − 1) windows of length m spanning a join, whatever the documents are — arithmetic about a concatenation, and it holds here to the last window on all three parts. How many of those windows are strings that no document contains is a property of the documents: twelve unrelated essays invent every one of them, and a version history invents 66.7% at the shortest window length, because near-copies of one file share their substrings and some spanning window turns out to occur inside a document anyway. The generated collection is near-copies by construction, so it measured the second case and reported it as the general one.twelve essaysm = 433 of 33m = 655 of 55m = 877 of 77m = 12121 of 121eight source modulesm = 420 of 21m = 634 of 35m = 848 of 49m = 1277 of 77ten revisions of one filem = 418 of 27m = 642 of 45m = 863 of 63m = 1299 of 99spanning windows · solid is the ones no document holds24,576 characters a partcount exact, share not
Fig. 7 One of the measurements the freezing makes reproducible: the strings a concatenation invents, counted exactly on documents that will not change.

The ratio between the two measures is what transfers

The table reports two measures per collection and reads them down the columns. Reading it across — dividing one by the other — gives a quantity that is far more stable than either, and it is where the generator’s mismatch shows up as a single number.

Twelve essays: 0.491 runs a character against 0.213 phrases, a ratio of 2.31. Eight source modules: 0.461 against 0.210, 2.20. Ten revisions: 0.126 against 0.0668, 1.89. The matched generated collection: about 0.118 against 0.048, 2.46.

So across four collections whose absolute repetitiveness spans a factor of four — prose at half a run a character, a version history at an eighth — the ratio between the two measures moves by 30%. The measures are nearly proportional, and the constant of proportionality is a much smaller thing than either measure.

That is the useful form of the relationship this strand keeps asserting. A collection with more runs per character has more phrases per character is true and weak; the two are within 30% of a fixed multiple of each other across every collection measured is the same observation with a number, and it is a number a reader can check against their own text in one pass.

It also locates the generator’s error precisely. The generated collection has the highest ratio of the four, 2.46, and the version history it is meant to model has the lowest, 1.89 — so the model produces too many runs for the phrases it produces, by about 30%.

Which is the same fact the dial that has no setting reports in a different shape. There the generator is fitted to the real collection twice, once on each measure, and the two settings come out a factor of 1.40 apart; here the mismatch is one ratio against another, 2.46 against 1.89, which is a factor of 1.30. Same direction, comparable size, and one number instead of two fitted parameters.

That suggests the calibration a future generator should be checked against. Fitting the divergence to reproduce rr leaves zz wrong; fitting it to reproduce zz leaves rr wrong; and fitting it to reproduce r/zr/z would catch the shape error in one comparison rather than requiring two fits and a discrepancy between them. It would not fix the generator — the model’s edits are scattered where real edits are localised, which is a structural difference no dial reaches — but it would report the failure in a single number computable from one pass over each collection.

And it explains why the relationship survives where the values do not. A run and a phrase are both bounded below by the same thing: the number of distinct contexts the text contains. So a collection’s two measures are two views of one property, differing in how they count it, and the ratio between them is a property of how the repetition is arranged rather than of how much there is. The collection decides which index is small is where the two measures decide between two structures; this is the quantity that says when they will disagree.

What this corpus cannot settle

Three limitations, stated because a real corpus invites more confidence than it earns.

It is one corpus. Three parts, thirty documents, three hundred and twenty thousand characters. Every number here is about it, and a different version history would give different numbers.

Its prose is unusually uniform. The twelve essays vary in length by a factor of 1.26, which is the narrowest spread of the three parts and is an artefact: they come from a collection that requires every essay to sit inside a length band. Real prose corpora are not like that.

And it is small. A hundred and forty-five thousand characters is a large collection by this site’s standards and a small one by any real measure. Nothing here says what happens at a million characters or at a hundred million, which is where the structures this collection weighs are actually deployed.

Two measures of repetition, on real text and on the model of itRuns of the transform and phrases of the parse, both divided by the length of the text so that three parts of different sizes can sit on one plate. The version history is the only part that is repetitive in the sense this field means: 0.089 runs a character against 0.449 for twelve unrelated essays. The generated collection at the setting used throughout this collection lands near the version history on runs and well below it on phrases, and nowhere near either of the other two — which is the whole question this strand was written to ask, drawn once before it is answered properly.twelve essays — runs0.449generated, matched0.118twelve essays — phrases0.175generated, matched0.043eight source modules — runs0.391generated, matched0.130eight source modules — phrases0.163generated, matched0.047ten revisions of one file — runs0.089generated, matched0.123ten revisions of one file — phrases0.042generated, matched0.045per character65,536 characters a partdivergence 0.02
Fig. 8 The same table at nearly three times the size. The ordering of the three parts holds and the numbers move a little, which is what a second size is for.

The one thing that is not a limitation

It is worth defending one choice that looks like a weakness.

The prose part is twelve of this collection’s own essays. That is self-referential, and a reader might reasonably wonder whether a collection measuring its own text is measuring anything.

It is, and the self-reference is incidental. What the prose part is for is to be natural language written by people about a technical subject, with the punctuation, the numbers, the capitals and the varied vocabulary that comes with it. Any twelve documents of that kind would do; these were to hand, their provenance is exactly known, and they can be listed by name so that a reader can see what was measured.

The one place the self-reference does bite is the length uniformity, and that is stated above rather than hidden: these essays are unusually similar in length because they were written to a standard, and a corpus of twelve arbitrary essays would have a wider spread.

A separator is free on an alphabet nobody choseThe document strand measured that giving every document its own separator costs a whole bit per character, and it does — on a generated collection whose alphabet is 21 symbols, where adding a few crosses 32 and the packed width goes up. Real text of this kind has 86 symbols and sits 42 short of the next power of two, so 3 more make no difference at all: the whole premium is 0.0%. The finding was correct and it was a finding about the generator.run together867 bits a characterone separator877 bits a charactera separator each897 bits a charactergenerated, run together215 bits a characterdistinct symbolstwelve essays0.0% premium
Fig. 9 A measurement that depends only on the alphabet, taken on the prose part. Where it lands is a property of natural language rather than of whose language it is.

What is checked

The corpus is what it says: three parts, thirty documents, none empty, none unlabelled, with their character counts and symbol counts asserted rather than described.

The documents are uneven, in at least one part, by a factor of more than 1.5 — and the matched generated collection is required to have a length spread of exactly one, which is the assertion that makes the comparison mean something.

The alphabets differ by more than a factor of two on every part, asserted in that direction so that a change to the generator’s alphabet would fail here rather than quietly removing this essay’s second finding.

And two things must fail. A corpus cut to one length loses 55.8% of itself and moves both measures. A corpus with its characters shuffled multiplies the run count by seven and leaves the entropy alone.

What follows

Three questions, in the order they matter.

The dial that has no setting asks whether the generator can be made to be this corpus — whether some fraction reproduces the real collection’s measures — and finds two answers a factor of 1.4 apart.

Documents that are not the same length takes the document strand’s own findings to real documents, and one of its headlines does not survive.

And what the generated collection was right about is the accounting: which of five strands’ conclusions hold on data nobody made, and which were properties of the maker.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 16 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AlphabetControlCorpusDocument collectionEntropyGenerated collectionIndex sizeMeasurementPhrase countRepetitionRun countVersion history