Error propagation — where it appears
Named by 2 essays across one field — each of them below, with the objects they name alongside it.
Also named here as inclusion exclusion, jaccard index — the same set of essays touches all of them, so they are one junction rather than several.
The error of a difference
Three sketches, each within a per cent or two of its own answer, subtracted into an intersection. At a Jaccard index of 0.82 the answer is 1.3% out. At 0.005 it is 146% out — the same three sketches, the same accuracy, a different question. The error never grew: it stayed a fixed fraction of the union, and the union stopped being the thing being asked about.
A budget split before the question arrives
Forty thousand bits a set, divided between a HyperLogLog and a sample of the smallest hashes before anyone knows which intersections will be asked for. Given all to either structure, the worst query is 2.7 or 12.8 times as far out as the best possible. Split in half, it is 1.8 — but only if the choice between the two estimates is made from a crossing measured in advance. The textbook error formulas make the choice wrongly on every one of 640 draws.
Named alongside it
The objects these essays reach for when they reach for this one.
Bottom-kCardinalityHyperLogLogInclusion exclusionJaccard indexRelative errorSketchState bitsEstimatorHonest limitMergeable summaryRegret