The shape of a real history's depth
A phrase in a self-referential parse copies from an earlier position. That position may itself sit inside a copied phrase, whose source sits inside another, and producing one character means following the chain to the bottom.
The length of that chain is the position’s depth, and the character that costs a chain is the essay about why it is the quantity an extraction pays. Four essays here cap it, sweep it, and price what the cap costs in phrases.
Every number in all four came from a generated collection: a base text, sixteen copies, two per cent of the characters replaced.
Nobody had computed the depth histogram of a real version history, which is a thing that takes four seconds once a real version history exists. Here it is, and it says three things — one confirming, one refining and one retracting.
The shape is a bell
Ten successive revisions of one file, 75,658 characters. Depths run from 0 to 14. The most common depth is 6, holding 18.4% of the collection; the mean is 5.71; two thirds of the characters sit between 4 and 7.
That is a bell with a short left tail and a longer right one, and it is not what one would guess. A geometric decay would be the natural guess — most positions shallow, a few deep, halving at each level — and it is what a parse of a text that copies mostly from nearby produces. This is the opposite: almost nothing at depth 0 or 1, a rise to a peak in the middle, a fall to a thin tail.
The reason is the shape of the collection rather than of the parse. A version history is a chain, not a star: revision 7 copies from revision 6, which copies from 5. A character that survives from the first revision to the last is at depth 9 by arithmetic and nothing shallower is available to it. So the distribution of depths is roughly the distribution of how long a passage has survived, and in a file being edited most passages have survived a while — which is a bell.
That matters for the cap immediately. A cap truncates the right-hand side, and if the mass were at the left a cap could be tight and cheap. Here the mass is under the middle of the range, so a cap set at the peak cuts half the collection.
The generated collection has a different shape for a reason
Put the model behind the measurement and the difference is structural rather than a matter of parameters.
The generator makes k copies of one base, each with a fraction of its characters replaced. Copy 7 is not derived from copy 6; it is derived from the base. So the parse finds its sources mostly in the base text, one hop away, and the depth it produces comes from the base text’s own internal repetition rather than from the copying.
A real history is a chain and a generated one is a star. That is not a divergence parameter set wrongly — no setting of the dial turns a star into a chain — and it is why the depth strand’s numbers are the numbers of a different object.
The dial that has no setting reached the same conclusion from the repetition measures: the generator’s r asks for one divergence and its z for another, a factor of 1.40 apart, because real edits are localised and generated ones are scattered. This is that finding one measure further on. The edits differ in where; the derivation differs in from what; and depth is the quantity that sees the second.
And a collection with no generations has nearly the same depth
Here is the retraction, and it is the thing that makes this page worth writing rather than filing.
Twelve unrelated essays — not a version history, no copying, no generations at all — have a mean depth of 5.17 and a worst depth of 15. The ten-revision history has a mean of 5.71 and a worst of 14.
Whatever produces depth in real text is mostly not generations of copying, because it is there in full when generations are absent.
What produces it instead is ordinary repetition at every scale: a word that occurred earlier is copied from a phrase that was itself copied from an earlier occurrence, and the chain of derivations is long even though no document was derived from another. Prose repeats a vocabulary — that is what its part of the corpus is for — and a vocabulary produces chains.
So the cap ladder’s central sentence needs weakening. Depth is generations of copying is true of the collection it was measured on and is not true of real text: it is one contributor among several, and on this evidence not the largest.
The part that holds, sharply
The confirming result is the sharpest of the three, and it needs the longer history to see.
A single revision of the file — one document, no generations — already reaches depth 8. Add a second and the worst depth is 10. From there, every added revision adds exactly one to the worst depth: 11, 12, 13, 14, 15, 16, 17, 18, 19, for nine consecutive generations, and then it stops at 19 and the last three revisions add nothing.
That is the ladder’s premise as an equality rather than a trend, and it is worth having in that form. The depth of a k-generation history is k plus what one document had on its own, until it saturates.
The offset is the part the ladder’s wording misses. Eight is a big number when k is two. It says that for the first several generations, most of the depth in a version history is not the history at all — it is the file’s own repetition, the same thing the essays’ 5.17 is made of.
And the saturation is worth a sentence too. It stops at 19 because the deepest chain in the collection stops being extended: the passages that survive from revision 1 to revision 12 are the ones nobody edits, and after that they are being copied from a recent revision rather than deepening a chain. A history long enough saturates, which means a cap set at the saturation depth is free forever after.
How a depth is computed, and what it is not
Worth being exact about the quantity, because three plausible things could be meant by “depth” and only one of them is this.
The parse produces phrases. A phrase at position i copying from position s is a claim that the characters at s…s+L are the characters at i…i+L, and the parse’s output is the list of those claims — the phrases a text copies from itself is the construction. To produce the character at i, a decoder goes to s; if s is itself inside a copied phrase it goes to that phrase’s source, and so on until it reaches a literal.
Depth is the number of hops in that walk, computed position by position: a literal is 0, and a copied position is one more than its source. It is computed from the phrase list, by a separate function from the one that produces the list, and that separation is deliberate — a construction that both enforces a bound and reports whether the bound holds is a check of nothing.
Three things it is not. It is not the number of times a passage has been copied in the world; a passage duplicated ten times in one document has depth 1 if all ten copies point at the first. It is not the phrase’s length or its distance from its source. And it is not a property of the text alone: a different parse of the same text gives different depths, which is exactly why the capped parse can exist at all.
That last point is what makes a histogram of depths a measurement of a decision as much as of a text. The parse here is the greedy one — the longest match, always — and the greedy choice is what puts the mass in the middle.
What the mean does, and what it means for a cap
The mean grows differently from the worst, and the difference is the whole content of the cap trade.
Across fourteen generations the mean depth goes from 3.19 to 7.92 — about 0.36 a generation, against the worst’s 1.00. Most characters are copied from the revision immediately before, not through the whole chain, so the mass moves slowly while the extreme moves at one level per generation.
A cap is a promise about the extreme. It says no character costs more than D copy-follows to produce, and it is paid for in phrases, because a source too deep to use forces the parse to take a shorter match or a literal.
Since the mass is under the middle of the distribution and the tail is thin, a cap near the mean is expensive and a cap near the worst is nearly free. On the ten-revision history a cap of 4 costs 5.7 times the free parse’s phrase count and a cap of 12 costs 0.4%. That is a very sharp knee, and it sits at a depth the generated collection put in a different place.
Where the tail is, which is somewhere else entirely
There is one collection here whose histogram does not fit any of this, and it turned out to be the most interesting.
Eight source modules reach depth 78. Not 15, not 19 — seventy-eight, on a collection with no generations at all.
And the tail is flat: the same number of characters at every depth from about 20 to 72. A flat tail is not something a distribution of copying produces; it is what a single object produces when its positions are spread one per level.
The deepest text is punctuation is the page about what that object is. The short version is that a phrase may copy from a source overlapping itself, so a run of identical characters parses as two phrases and its depths increase by one per character — and the deepest positions in eight files of source code are the dashes in a comment separator.
When the history overtakes the file
The retraction above and the confirmation above sit uneasily together, and the arithmetic that reconciles them is worth doing, because it turns “depth is not mostly generations” into a statement with a boundary rather than a flat contradiction.
Both quantities are affine in the generation count. The worst depth is — an offset the file has on its own, plus one per revision, exactly. The mean is close to across the fourteen-revision curve. So each has a baseline that owes nothing to the history and a term that is entirely the history, and the question of which dominates is a question about .
Setting the two terms equal answers it. For the worst depth, . For the mean, . A version history’s own copying overtakes the file’s internal repetition at somewhere between eight and nine revisions, and the two statistics agree on the number despite being computed from entirely different parts of the distribution.
That is the boundary the retraction needed. Below it, the earlier sentence is right and depth is mostly not generations: at two revisions the history contributes a fifth of the worst depth and a tenth of the mean, and the rest is what one document already had. Above it the original ladder’s premise recovers, and by fourteen revisions the history owns 64% of the worst depth and 60% of the mean. Neither reading is wrong; they describe opposite sides of a crossing nobody had located because nobody had separated the two terms.
Two cautions go with it, and the second is the one that stops this becoming a rule of thumb.
The offset is not a constant of the world. Twelve unrelated essays have a mean depth of 5.17 — higher than one revision of the file at 3.19 — because a larger and more varied collection has more internal repetition to copy from. So the baseline grows with the collection, which means the crossing moves with it: a bigger single-generation corpus needs more generations before its history dominates. Eight to nine is this file’s number, not a universal one, and it is exactly the sort of quantity a corpus that was not generated exists to keep honest.
And the two statistics cross at the same place for different reasons, which makes the agreement a coincidence worth not over-reading. The worst depth’s slope is 1.00 because a surviving passage gains exactly one level per revision; the mean’s is 0.36 because most characters are copied from the immediately preceding revision and never join the long chain. That the ratio of baseline to slope lands in the same place for both is arithmetic rather than mechanism, and treating the mean and the extreme as one quantity is precisely the error expected is not average is about. One revision, one level takes the slope apart, and the cap that would ship is where the crossing turns into a setting.
What a reader should take from the histogram
Three things, in the order they change a decision.
Depth is not a proxy for how many times something was copied. It is a chain length in a parse, and the chain can come from generations, from ordinary repetition, or from a run of one character. On real text all three are present and the third owns the extreme.
A real history’s depths are a bell, so a cap is a cliff. There is no cheap tight cap on a version history, because the mass is in the middle rather than at the left. The useful setting is near the worst depth, which is near the generation count, which is a number a system usually knows about itself.
And the generated collection was the wrong shape rather than the wrong size. A star is not a chain. That is worth remembering the next time a model is chosen for a measurement: the parameter that was swept was divergence, and the parameter that mattered was one nobody had written down.
What this leaves for the rest of the strand
The histogram answers what the depth of a real collection is. It leaves three questions that the next three pages take in order.
Where the extreme actually comes from, which is the flat tail on the source code, and which turns out to be a single kind of object rather than a distribution.
What the cap costs on real text, since the ladder’s sweep was drawn on the star-shaped collection and the real ones are steeper — a cap of one costs 24 times the free parse on the ten-revision history and 48 on the fourteen, against the model’s 21.
And what a system should actually set, which needs both: the knee on a real history is at 9 to 12 rather than the 4 to 8 the ladder reported, and the reason it is there is the saturation this page measured.
One more thing is worth recording because it is a measurement about the instrument rather than about the text. Every histogram here comes from a parse that is linear in the length of the text — the parse in one pass of the text — and the quadratic construction it replaced could not have drawn a single one of these plates. A hundred and seventy thousand characters would have been a hundred and seventy thousand times an eight-thousand-character measurement, which is why the depth profile of a real history had never been computed and not because anybody decided it did not matter.
Where this sits
Four pages on real depth. This one is the histogram. One revision, one level is the same quantity against the generation count. The deepest text is punctuation is what the tail is made of. The cap that would ship is the parameter, set from all three.
The ladder they are all arguing with is a parse that will not follow a long chain and the three essays after it, every number of which came from a generated collection — and the collection this page measures instead is the one a corpus that was not generated froze.
What this makes readable
Essays that name this one as a prerequisite.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- A boundary that costs nothing control · corpus · phrase count
- A collection is a construction control · corpus · phrase count
- Documents that are not the same length control · corpus · generated collection
- The half of a fall that is the logarithm control · corpus · phrase count
- What repetition is worth once the logarithm is gone control · corpus · phrase count
- What the generated collection was right about control · corpus · generated collection
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
ControlCopyCorpusDepthExtractionGenerated collectionHistogramParsePhrase countRunVersion history