The floor was the marks
A bidirectional index is two self-indexes and one interval. The reverse half exists to be counted in — an extension asks it for ranks and never for a position, because its rows are positions in a text written backwards and the occurrences a search reports come out of the forward interval.
So the reverse half’s locating apparatus can go. An index that cannot locate is the structure and the half that is never asked where is the measurement, which put the saving at 16.7% of both halves at one sampled position in thirty-two — the sixth the deferral naming it had predicted.
That essay also drew the saving across the sampling dial and reported it flattening: 31.2% at one in four, 16.9% at one in thirty-two, 13.5% at one in two hundred and fifty-six, and levelling. Its own prose named the reason correctly: the marks are n bits however rarely a row is kept, and that is the floor the curve flattens onto.
The floor was the marks, the marks did not have to be n bits, and with them represented properly the saving is a fiftieth.
The two curves
With plain marks: 31.2% at one in four, 25.1%, 20.2%, 17.0%, 15.1%, 14.0%, 13.5% at one in two hundred and fifty-six. A fall of 2.31 across the dial, and the last three points are within one and a half percentage points of each other.
With Elias–Fano marks: 31.4%, 23.8%, 16.3%, 10.2%, 5.9%, 3.3%, 1.8%. A fall of 17.6, and still falling at the sparse end.
At the dense end the two are within two tenths of a per cent, because at one in four the sparse representation costs the same as the plain vector. At the sparse end they differ by a factor of 7.5.
The two curves crossing at the dense end and diverging at the sparse one is the signature of a constant term being removed, and it is worth recognising because it is what such a removal always looks like. Two quantities that agree where the constant is a small share and diverge where it is most of the total.
If the two curves had been parallel, the sparse representation would be saving a fixed fraction and the floor would still be there at a lower level. If they had crossed and swapped, something would be wrong. Diverging monotonically from a common point is the shape of “one of these has a term the other does not”.
What was actually being dropped
The reverse half’s locating apparatus is two arrays: the sampled positions, (n/s)⌈log₂ n⌉ bits, and the marks, n bits.
At one in thirty-two those are 7,695 and 16,385. Sixty-eight per cent of what the earlier essay was dropping was the marks.
At one in two hundred and fifty-six: 963 and 16,385. Ninety-four per cent.
So the flat part of the published curve was almost entirely one array, and the array was a dense encoding of a set with n/s elements. The saving being celebrated was the saving from not storing a badly stored thing.
Represent it well and there is much less to drop, which is what the lower curve is.
Putting those percentages into bits makes the earlier result read differently. At one in two hundred and fifty-six, the reverse half’s locating apparatus is 17,348 bits, of which 963 are the sampled positions the apparatus exists for and 16,385 are the array saying which rows they are.
So the essay reporting “the reverse half’s locating apparatus is worth 13.5% of both halves” was reporting, in substance, that a sixteen-thousand-bit index over nine hundred and sixty-three bits of data can be deleted. Which is true and is a strange thing to celebrate.
Is the earlier result wrong
No, and the distinction is worth being careful about because “wrong” and “superseded” are different.
Every number in the earlier essay is a correct measurement of the structure it was measuring: a bidirectional index whose marks are plain bit vectors, which is what every implementation of one has. Dropping the reverse half’s locating apparatus from that structure saves 16.7% at one in thirty-two, and it does.
What has changed is that the structure has a better version. On the better version the same operation saves 10.2%, and at sparser rates far less.
So the earlier essay’s measurement stands and its conclusion does not travel. “Dropping the reverse half’s locating apparatus is worth about a sixth” was a claim about a structure; it is now a claim about a structure nobody should build.
That is the ordinary way a result ages and it is worth saying explicitly because the alternative — quietly updating a number — loses the information that the two structures are different.
What the deferral had asked for
The strand that produced the earlier result ended with a list of what it had not done, and one item was: a sparse representation of the sample marks, which is the floor under the half-index saving and is a bit vector with one bit in s set.
That sentence is exactly right. It names the component, names why it matters, and names what would fix it.
What it does not say is that fixing it would reduce the saving the strand had just reported — and it would not have been reasonable to expect that. A deferral is written to say what is left, and this one describes a floor as a limit on how much better the result could get. It is a limit on how much of the result was real.
That reversal is worth noticing because it says something about how deferrals should be read. A deferral naming a component as a floor is a deferral naming a component that is large, and a large component being replaced changes every number it was part of — including the ones in the essay that named it.
What the corrected number is
The saving from dropping a counting-only half’s locating apparatus, with marks properly represented:
At one in four: 31.4%. At one in eight: 23.8%. At one in sixteen: 16.3%. At one in thirty-two: 10.2%. At one in sixty-four: 5.9%. At one in a hundred and twenty-eight: 3.3%. At one in two hundred and fifty-six: 1.8%.
The number a reader should carry is a range and a direction: between a tenth and a fiftieth, falling steeply with the sampling rate, and negligible on any index sampling sparsely.
The old summary — about a sixth, flattening at an eighth — is replaced by a summary with no flat part in it, which is a worse summary in the sense that it cannot be reduced to one number, and a better one in that the number it replaces was about an artefact.
What the reverse half is still worth dropping
The retraction should not be read as saying the half-index idea was empty, and the numbers say why.
At one in thirty-two, 10.2% of a bidirectional index is still a tenth of a structure, removed by not building something. That is a better return than almost anything in this field costs to obtain: no new machinery, no query slowdown, no compromise — the reverse half genuinely is never asked for a position, and the arrays genuinely are not needed.
What has changed is the ordering of the improvements available to a bidirectional index. Before this strand, dropping the reverse half’s locating apparatus was the largest single thing anybody could do to one. Now representing the marks is larger at sparse rates and comparable at dense ones, and the two together are larger than either.
That ordering matters to somebody deciding what to implement. A tenth for a deletion and a fifteenth for an encoding, or a sixth for the deletion alone as the earlier essay would have suggested — the second reads as one improvement being clearly dominant and the first as two of comparable size.
A sixth of what, exactly is the essay that established this collection’s rule about quoting a share with its denominator, and it was written about this same saving one strand earlier. The rule it produced was right and insufficient: the share was quoted with its sampling rate, correctly, and the thing it was a share of had a badly represented component in it that nobody had questioned.
So the rule wants a second clause. A share carries its denominator, and a denominator carries what it is made of.
What a reader of the earlier essay should do
The practical question a retraction raises is what to do with the thing being retracted, and there are three defensible answers.
Leave it and link. The measurement stands, the structure it measured is one people build, and a reader who arrives at it from a search engine should find the number and the correction together. That is what this collection does.
Rewrite the number. Tempting, and it destroys the information that two structures exist. A reader who finds 10.2% in an essay about dropping the reverse half’s apparatus cannot tell whether that is a measurement of a plain-marked index or a sparse-marked one, and the difference is the entire point.
Withdraw it. Appropriate for a measurement that was wrong. This one was not: it correctly measured a structure that correctly implemented what every published account describes.
The first is right and it has a cost worth naming, which is that the collection now contains two numbers for one saving and a reader has to notice which structure each is about. That is a real burden and the alternative burdens are worse.
The composition, which is the useful half
There is a way of reading the two savings together that makes both worth having, and it is the one this strand ends on.
The half-index saving removes an array from one of two halves. The sparse representation shrinks an array in both. They are not alternatives; they compose, because one is a deletion and the other is an encoding.
Applied cumulatively at one in thirty-two on a bidirectional index over sixteen thousand characters: the plain structure is 250,584 bits; dropping the reverse half’s locating apparatus takes it to 226,504 (90.4%); representing the surviving marks sparsely takes it to 210,934 (84.2%).
The second step is worth 15,570 bits against the first step’s 24,080 — which is to say the encoding is worth nearly two thirds of what the deletion was, on a structure where the deletion was the headline.
The ladder, and the rung that spends is where that cumulative accounting is drawn, and it has a third rung: the whole saving spent on sampling the surviving half four times as densely, which lands at 96.4% of the original size with a locate several times faster.
Other flattening curves in this collection
If the habit is worth having then it should have somewhere to be applied, so it is worth naming the curves in this collection that flatten and saying whether their asymptotes have been priced.
A sampling that costs more than the array draws a run-length index’s parts against the copy count and finds one of them not shrinking. That asymptote was priced — it is the regular suffix-array sampling — and the answer was to change the sampling policy entirely, which is what the sampling that follows the runs is.
What is still proportional to n is the same question asked systematically of one structure, and it is the closest thing this collection has to the habit already being a habit.
The threshold that reaches zero draws a filter’s selectivity flattening, and the asymptote there is a property of the data rather than of a representation — so there is nothing to price.
Two of the four have been chased and two have not, which is a better rate than this strand’s own history suggested. What distinguishes the ones that were chased is that their asymptote was a structure rather than a term: a sampling policy is a thing to replace and “the marks are n bits” reads as a fact.
What a flattening curve means
The reusable part of this is a reading habit and it is short.
A curve that flattens has a constant term in it. The constant is where the next result is.
Both of this neighbourhood’s earlier essays drew a flattening curve, identified the constant correctly, and stopped — because a curve going to a limit reads as a quantity that has been understood, and naming the reason for the limit completes the explanation.
An explanation and a target look identical on the page. What separates them is whether anybody asks what the constant costs to remove, and in this case the answer was a fifty-year-old encoding and about two hundred lines.
The general instruction: when a sweep flattens, price the asymptote. Not “identify” — that had been done twice — but price, which means building the alternative.
The check that would have caught it earlier
There is a check that would have found this at the time and it is not a check anybody would have thought to write, which is the interesting part.
The earlier essay’s rejection test was that a saving needs its sampling rate: it required the saving at two ends of the dial to differ by a factor of three, and found them differing by 2.19 — so the check was wrong and the measurement was the finding, which the essay recorded.
That 2.19 was the flattening. The check was asking “does this saving depend on the rate?” and the answer was “less than expected”, and the reason it was less than expected is the whole of this strand.
So the information was there, in a failed check, correctly recorded, one strand early. What was missing was a step from “this saving depends on the rate less than expected” to “so a rate-independent term is most of it, and what is it?”.
That step is one question and it has a name: when a quantity depends on a dial less than predicted, find the part that does not depend on it. The prediction is the thing that makes it noticeable — a check expecting a factor of three and finding 2.19 is a check that has located a constant without saying so.
This collection’s habit of writing rejection tests that state an expected magnitude, rather than merely a direction, is what made that possible. A check requiring only that the saving fall would have passed and said nothing.
The three things this strand changed
Collecting them, because they are of decreasing size and increasing generality.
A number. The half-index saving is a tenth rather than a sixth at the usual rate, and a fiftieth at sparse ones.
A structure. An index’s marks are an Elias–Fano array, which is 15% of the whole index at one in thirty-two and 18% at one in a hundred and twenty-eight.
A habit. Two essays named a floor and neither priced it. The instruction is to price it, and the reason it needs to be an instruction is that naming feels like finishing.
The third is the one that would have found this two strands earlier, and it is the one worth carrying to the next flattening curve — of which this collection has several, in fields with nothing to do with text.
There is a fourth thing, and it is about how a collection of essays behaves rather than about indexes. Three essays in this neighbourhood now say different things about the same saving: a sixth, a sixth-with-a-dial, and a tenth-falling-to-a-fiftieth. Each was right about the structure in front of it.
What makes that a collection rather than a contradiction is that the structures differ and each essay says which one it measured. A number without its structure is the failure this whole neighbourhood keeps rediscovering, and it has now been rediscovered at three levels: a share without its denominator, a saving without its dial, and a denominator without its components.
Named alongside this one
Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.
- A position split in two elias fano · index size · sample marks
- A price with no structure under it elias fano · index size · sample marks
- The flat bottom of a shallow curve elias fano · index size · sample marks
- Twenty bits apart elias fano · index size · sample marks
- Where the sparse representation loses elias fano · index size · sample marks
- A constant factor, not a term index size · space accounting
What links here
Every essay whose body links to this one.
The objects this essay names
Each one links to every other essay that touches it.
Bidirectional indexElias fanoIndex sizeLocatingRetractionSample marksSpace accounting