Logarithm — where it appears
Named by 2 essays across 2 fields — each of them below, with the objects they name alongside it.
A million characters of the same thing
Every measurement this collection has published about real text was taken on twenty-four thousand characters, because the phrase count was quadratic. It is linear now, so here is the same corpus at forty times the size — and what forty times does to its own numbers.
The half of a fall that is the logarithm
Phrases per character on a real collection of essays fall by a factor of 2.34 as it grows. A shuffle of the same characters falls by 1.70. Nearly three quarters of the movement is arithmetic, and no definition of the measure says so.
Named alongside it
The objects these essays reach for when they reach for this one.
AlphabetControlCorpusEntropyPhrase countRepetitionRun countScaleDenominatorDocument collectionIndex sizeShuffle