The data that is not a number

The shape of a real history's depth

Four essays here are about capping how far an extraction follows a chain of copies, and every number in them came from a generated collection. Here is the depth histogram of a real version history, which is a bell, and of twelve unrelated essays, which is nearly the same bell.

A phrase in a self-referential parse copies from an earlier position. That position may itself sit inside a copied phrase, whose source sits inside another, and producing one character means following the chain to the bottom.

The length of that chain is the position’s depth, and the character that costs a chain is the essay about why it is the quantity an extraction pays. Four essays here cap it, sweep it, and price what the cap costs in phrases.

Every number in all four came from a generated collection: a base text, sixteen copies, two per cent of the characters replaced.

The depth profile of a real version historyEvery position of ten revisions of one file — 75,658 characters — placed at the number of copies an extraction has to follow to produce it. The shape is a bell with a peak at 6, holding 18.4% of the collection, and a thin tail out to 14. Drawn behind it is twelve unrelated essays, which is not a version history and has no generations of copying at all: its mean depth is 5.17 against this one's 5.71. Whatever produces depth in real text is mostly not what the cap ladder was about, because it is here in full when generations are absent.02468101214mean 5.71depthshare of the collectionten revisionstwelve essays75,658 characters · 2,966 phrasesworst 14, mean 5.71
Fig. 1 The depth of every position of ten real revisions of one file, with twelve unrelated essays drawn behind it. Both are real; only one of them has generations of copying in it.

Nobody had computed the depth histogram of a real version history, which is a thing that takes four seconds once a real version history exists. Here it is, and it says three things — one confirming, one refining and one retracting.

The shape is a bell

Ten successive revisions of one file, 75,658 characters. Depths run from 0 to 14. The most common depth is 6, holding 18.4% of the collection; the mean is 5.71; two thirds of the characters sit between 4 and 7.

That is a bell with a short left tail and a longer right one, and it is not what one would guess. A geometric decay would be the natural guess — most positions shallow, a few deep, halving at each level — and it is what a parse of a text that copies mostly from nearby produces. This is the opposite: almost nothing at depth 0 or 1, a rise to a peak in the middle, a fall to a thin tail.

The reason is the shape of the collection rather than of the parse. A version history is a chain, not a star: revision 7 copies from revision 6, which copies from 5. A character that survives from the first revision to the last is at depth 9 by arithmetic and nothing shallower is available to it. So the distribution of depths is roughly the distribution of how long a passage has survived, and in a file being edited most passages have survived a while — which is a bell.

That matters for the cap immediately. A cap truncates the right-hand side, and if the mass were at the left a cap could be tight and cheap. Here the mass is under the middle of the range, so a cap set at the peak cuts half the collection.

The depth profile of a real version historyEvery position of fourteen revisions of another — 174,282 characters — placed at the number of copies an extraction has to follow to produce it. The shape is a bell with a peak at 8, holding 11.6% of the collection, and a thin tail out to 19. Drawn behind it is a generated collection, which is not a version history and has no generations of copying at all: its mean depth is 4.63 against this one's 7.92. Whatever produces depth in real text is mostly not what the cap ladder was about, because it is here in full when generations are absent.024681012141618mean 7.92depthshare of the collectionfourteen revisionsthe model174,282 characters · 3,462 phrasesworst 19, mean 7.92
Fig. 2 Fourteen revisions of a second file, with the generated collection behind it. The generated one is not a chain: every copy is one generation from the base, so its depths come from the base text’s own repetition rather than from the history.

The generated collection has a different shape for a reason

Put the model behind the measurement and the difference is structural rather than a matter of parameters.

The generator makes k copies of one base, each with a fraction of its characters replaced. Copy 7 is not derived from copy 6; it is derived from the base. So the parse finds its sources mostly in the base text, one hop away, and the depth it produces comes from the base text’s own internal repetition rather than from the copying.

A real history is a chain and a generated one is a star. That is not a divergence parameter set wrongly — no setting of the dial turns a star into a chain — and it is why the depth strand’s numbers are the numbers of a different object.

The dial that has no setting reached the same conclusion from the repetition measures: the generator’s r asks for one divergence and its z for another, a factor of 1.40 apart, because real edits are localised and generated ones are scattered. This is that finding one measure further on. The edits differ in where; the derivation differs in from what; and depth is the quantity that sees the second.

And a collection with no generations has nearly the same depth

Here is the retraction, and it is the thing that makes this page worth writing rather than filing.

Twelve unrelated essays — not a version history, no copying, no generations at all — have a mean depth of 5.17 and a worst depth of 15. The ten-revision history has a mean of 5.71 and a worst of 14.

The depth profile of a real version historyEvery position of twelve unrelated essays — 144,617 characters — placed at the number of copies an extraction has to follow to produce it. The shape is a bell with a peak at 5, holding 21.1% of the collection, and a thin tail out to 15. Drawn behind it is ten revisions of one file, which is not a version history and has no generations of copying at all: its mean depth is 5.71 against this one's 5.17. Whatever produces depth in real text is mostly not what the cap ladder was about, because it is here in full when generations are absent.02468101214mean 5.17depthshare of the collectiontwelve essaysten revisions144,617 characters · 21,166 phrasesworst 15, mean 5.17
Fig. 3 The same two collections with the roles exchanged: twelve unrelated essays in front, the version history behind. The shapes are nearly the same, and one of the two collections has nothing the depth strand is about.

Whatever produces depth in real text is mostly not generations of copying, because it is there in full when generations are absent.

What produces it instead is ordinary repetition at every scale: a word that occurred earlier is copied from a phrase that was itself copied from an earlier occurrence, and the chain of derivations is long even though no document was derived from another. Prose repeats a vocabulary — that is what its part of the corpus is for — and a vocabulary produces chains.

So the cap ladder’s central sentence needs weakening. Depth is generations of copying is true of the collection it was measured on and is not true of real text: it is one contributor among several, and on this evidence not the largest.

The part that holds, sharply

The confirming result is the sharpest of the three, and it needs the longer history to see.

One more revision is one more level of depthThe worst and mean depth of the first k revisions of one real file. A single revision, with no generations of copying at all, already reaches depth 8: that is the file's own internal repetition, and it is the larger term for the first several generations. From there every added revision adds exactly one to the worst depth — 9 times in a row — until the twelfth, after which nothing gets deeper and the last revisions only add width. The mean rises steadily and far more slowly, at about 0.36 a generation, because most characters are copied from the revision before rather than through the whole chain.05101520510revisions in the collectiondepthone revision alone: 8worst depthmean depthfourteen revisions of another9 generations added exactly one
Fig. 4 The worst and mean depth of the first k revisions of one real file. The horizontal rule is the depth of a single revision on its own.

A single revision of the file — one document, no generations — already reaches depth 8. Add a second and the worst depth is 10. From there, every added revision adds exactly one to the worst depth: 11, 12, 13, 14, 15, 16, 17, 18, 19, for nine consecutive generations, and then it stops at 19 and the last three revisions add nothing.

That is the ladder’s premise as an equality rather than a trend, and it is worth having in that form. The depth of a k-generation history is k plus what one document had on its own, until it saturates.

The offset is the part the ladder’s wording misses. Eight is a big number when k is two. It says that for the first several generations, most of the depth in a version history is not the history at all — it is the file’s own repetition, the same thing the essays’ 5.17 is made of.

And the saturation is worth a sentence too. It stops at 19 because the deepest chain in the collection stops being extended: the passages that survive from revision 1 to revision 12 are the ones nobody edits, and after that they are being copied from a recent revision rather than deepening a chain. A history long enough saturates, which means a cap set at the saturation depth is free forever after.

How a depth is computed, and what it is not

Worth being exact about the quantity, because three plausible things could be meant by “depth” and only one of them is this.

The parse produces phrases. A phrase at position i copying from position s is a claim that the characters at s…s+L are the characters at i…i+L, and the parse’s output is the list of those claims — the phrases a text copies from itself is the construction. To produce the character at i, a decoder goes to s; if s is itself inside a copied phrase it goes to that phrase’s source, and so on until it reaches a literal.

Depth is the number of hops in that walk, computed position by position: a literal is 0, and a copied position is one more than its source. It is computed from the phrase list, by a separate function from the one that produces the list, and that separation is deliberate — a construction that both enforces a bound and reports whether the bound holds is a check of nothing.

Three things it is not. It is not the number of times a passage has been copied in the world; a passage duplicated ten times in one document has depth 1 if all ten copies point at the first. It is not the phrase’s length or its distance from its source. And it is not a property of the text alone: a different parse of the same text gives different depths, which is exactly why the capped parse can exist at all.

That last point is what makes a histogram of depths a measurement of a decision as much as of a text. The parse here is the greedy one — the longest match, always — and the greedy choice is what puts the mass in the middle.

What the mean does, and what it means for a cap

The mean grows differently from the worst, and the difference is the whole content of the cap trade.

Across fourteen generations the mean depth goes from 3.19 to 7.92 — about 0.36 a generation, against the worst’s 1.00. Most characters are copied from the revision immediately before, not through the whole chain, so the mass moves slowly while the extreme moves at one level per generation.

A cap is a promise about the extreme. It says no character costs more than D copy-follows to produce, and it is paid for in phrases, because a source too deep to use forces the parse to take a shorter match or a literal.

What a cap costs, on five collectionsThe phrase count under a cap, as a multiple of the phrase count without one, against the cap. Every number the cap ladder published came from the generated collection, which is the line marked as such; the real version histories are far steeper. A cap of one costs 35x on fourteen revisions of another against 21x on the model, and the two collections of unrelated documents are flatter than either. The knee — the smallest cap within a twentieth of the free parse — is at 10, 10, 10, 10, 10 respectively, against the four to eight the ladder reported.110110cap on the depthphrases ÷ free parseten revisionsfourteen revisionstwelve essayseight modulesthe model5 collections · free parse = 1worst 35x at a cap of one
Fig. 5 The phrase count under a cap, as a multiple of the free parse’s, on five collections. The two version histories are the steep lines; the collections of unrelated documents are the flat ones.

Since the mass is under the middle of the distribution and the tail is thin, a cap near the mean is expensive and a cap near the worst is nearly free. On the ten-revision history a cap of 4 costs 5.7 times the free parse’s phrase count and a cap of 12 costs 0.4%. That is a very sharp knee, and it sits at a depth the generated collection put in a different place.

Where the tail is, which is somewhere else entirely

There is one collection here whose histogram does not fit any of this, and it turned out to be the most interesting.

How many of the deep positions are a run of one charactereight source modules, 106,283 characters. The bars are the share of the collection at each depth and the darker part of each is the share that sits inside a run — a position whose character equals the one before it. A parse may copy from a source that overlaps it, which is what makes a thousand identical characters two phrases rather than five hundred, and it makes the depth of the k-th character of a run exactly k. So the right-hand end of this histogram is not the most-copied text in the collection: it is its longest run, one level per character.051015202530354045505560657075depthshare of the collectioninside a runeverything elseeight source modules · worst depth 78mean 6.03
Fig. 6 Eight modules of source code. The bars are the share of the collection at each depth, and the darker part of each is the share inside a run of one repeated character.

Eight source modules reach depth 78. Not 15, not 19 — seventy-eight, on a collection with no generations at all.

And the tail is flat: the same number of characters at every depth from about 20 to 72. A flat tail is not something a distribution of copying produces; it is what a single object produces when its positions are spread one per level.

The deepest text is punctuation is the page about what that object is. The short version is that a phrase may copy from a source overlapping itself, so a run of identical characters parses as two phrases and its depths increase by one per character — and the deepest positions in eight files of source code are the dashes in a comment separator.

When the history overtakes the file

The retraction above and the confirmation above sit uneasily together, and the arithmetic that reconciles them is worth doing, because it turns “depth is not mostly generations” into a statement with a boundary rather than a flat contradiction.

Both quantities are affine in the generation count. The worst depth is 8+k8 + k — an offset the file has on its own, plus one per revision, exactly. The mean is close to 3.19+0.36k3.19 + 0.36k across the fourteen-revision curve. So each has a baseline that owes nothing to the history and a term that is entirely the history, and the question of which dominates is a question about kk.

Setting the two terms equal answers it. For the worst depth, k=8k = 8. For the mean, k=3.19/0.369k = 3.19/0.36 \approx 9. A version history’s own copying overtakes the file’s internal repetition at somewhere between eight and nine revisions, and the two statistics agree on the number despite being computed from entirely different parts of the distribution.

That is the boundary the retraction needed. Below it, the earlier sentence is right and depth is mostly not generations: at two revisions the history contributes a fifth of the worst depth and a tenth of the mean, and the rest is what one document already had. Above it the original ladder’s premise recovers, and by fourteen revisions the history owns 64% of the worst depth and 60% of the mean. Neither reading is wrong; they describe opposite sides of a crossing nobody had located because nobody had separated the two terms.

Two cautions go with it, and the second is the one that stops this becoming a rule of thumb.

The offset is not a constant of the world. Twelve unrelated essays have a mean depth of 5.17 — higher than one revision of the file at 3.19 — because a larger and more varied collection has more internal repetition to copy from. So the baseline grows with the collection, which means the crossing moves with it: a bigger single-generation corpus needs more generations before its history dominates. Eight to nine is this file’s number, not a universal one, and it is exactly the sort of quantity a corpus that was not generated exists to keep honest.

And the two statistics cross at the same place for different reasons, which makes the agreement a coincidence worth not over-reading. The worst depth’s slope is 1.00 because a surviving passage gains exactly one level per revision; the mean’s is 0.36 because most characters are copied from the immediately preceding revision and never join the long chain. That the ratio of baseline to slope lands in the same place for both is arithmetic rather than mechanism, and treating the mean and the extreme as one quantity is precisely the error expected is not average is about. One revision, one level takes the slope apart, and the cap that would ship is where the crossing turns into a setting.

What a reader should take from the histogram

Three things, in the order they change a decision.

Depth is not a proxy for how many times something was copied. It is a chain length in a parse, and the chain can come from generations, from ordinary repetition, or from a run of one character. On real text all three are present and the third owns the extreme.

A real history’s depths are a bell, so a cap is a cliff. There is no cheap tight cap on a version history, because the mass is in the middle rather than at the left. The useful setting is near the worst depth, which is near the generation count, which is a number a system usually knows about itself.

And the generated collection was the wrong shape rather than the wrong size. A star is not a chain. That is worth remembering the next time a model is chosen for a measurement: the parameter that was swept was divergence, and the parameter that mattered was one nobody had written down.

One more revision is one more level of depthThe worst and mean depth of the first k revisions of one real file. A single revision, with no generations of copying at all, already reaches depth 9: that is the file's own internal repetition, and it is the larger term for the first several generations. From there every added revision adds exactly one to the worst depth — 4 times in a row — until the twelfth, after which nothing gets deeper and the last revisions only add width. The mean rises steadily and far more slowly, at about 0.36 a generation, because most characters are copied from the revision before rather than through the whole chain.051015246810revisions in the collectiondepthone revision alone: 9worst depthmean depthten revisions of one file4 generations added exactly one
Fig. 7 The generation curve on the shorter history, where the offset is nine and the steps are less regular — a second file, behaving the same way with different constants.

What this leaves for the rest of the strand

The histogram answers what the depth of a real collection is. It leaves three questions that the next three pages take in order.

Where the extreme actually comes from, which is the flat tail on the source code, and which turns out to be a single kind of object rather than a distribution.

What the cap costs on real text, since the ladder’s sweep was drawn on the star-shaped collection and the real ones are steeper — a cap of one costs 24 times the free parse on the ten-revision history and 48 on the fourteen, against the model’s 21.

And what a system should actually set, which needs both: the knee on a real history is at 9 to 12 rather than the 4 to 8 the ladder reported, and the reason it is there is the saturation this page measured.

One more thing is worth recording because it is a measurement about the instrument rather than about the text. Every histogram here comes from a parse that is linear in the length of the text — the parse in one pass of the text — and the quadratic construction it replaced could not have drawn a single one of these plates. A hundred and seventy thousand characters would have been a hundred and seventy thousand times an eight-thousand-character measurement, which is why the depth profile of a real history had never been computed and not because anybody decided it did not matter.

Where this sits

Four pages on real depth. This one is the histogram. One revision, one level is the same quantity against the generation count. The deepest text is punctuation is what the tail is made of. The cap that would ship is the parameter, set from all three.

The ladder they are all arguing with is a parse that will not follow a long chain and the three essays after it, every number of which came from a generated collection — and the collection this page measures instead is the one a corpus that was not generated froze.

Where each collection stops paying for its capThe knee: the smallest cap whose phrase count is within a twentieth of the free parse's. The cap ladder reported four to eight, on a generated collection at 8,192 characters. On real text it is nine to twelve, and the two version histories are the slowest — which is what one would expect once the depth profile is drawn, because the knee is near the worst depth and the worst depth is the number of generations plus what one document had on its own. The row for the generated collection is the ladder's own number, re-measured here at the same size as the others.ten revisions of one filecap 10worst 14 · mean 5.7fourteen revisions of anothercap 10worst 17 · mean 7.2twelve unrelated essayscap 10worst 14 · mean 4.9eight source modulescap 10worst 78 · mean 5.9a generated collectioncap 10worst 13 · mean 4.6within a twentieth of the free parsefour to eight was the model
Fig. 8 Where the four pages end up: the smallest cap each collection stops paying for, against the four-to-eight the earlier sweep reported.

What this makes readable

Essays that name this one as a prerequisite.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

Every essay whose body links to this one.

The objects this essay names

Each one links to every other essay that touches it.

ControlCopyCorpusDepthExtractionGenerated collectionHistogramParsePhrase countRunVersion history