What the libraries do

Sized for a rate that does not hold still

A four-second window on a stream at a hundred arrivals a second holds four hundred items on average and between 105 and 2,169 when the rate moves. An allocation set at that average overflows at 47 per cent of instants while 35 per cent of it stands empty, which is the same decision failing in both directions at once.

A structure that holds a time window has to be allocated for something, and the number it is allocated for is nearly always the mean occupancy — the rate times the duration, which is the first thing anybody computes and the only thing the window’s specification supplies.

The window that is not full established that the occupancy is a distribution rather than a number, and drew its range. What it did not do is convert the range into the two things a sizing decision is actually judged on.

There is no allocation at which both costs are smallA time window of 4.0 s over the drifting stream at a mean 100 Hz. The allocation is the mean occupancy times the headroom on the horizontal axis. The falling curve is the share of sampled instants at which the window holds more than the allocation — where a real structure drops something. The rising curve is the mean share of the allocation standing empty. At a headroom of one, 47% of instants overflow; at the headroom that stops the overflow, 67% of the allocation is idle. A capacity chosen from the mean is a point on this curve rather than an answer to it.0%25%50%75%100%1.5×shareheadroom over the mean occupancyoverflowingstanding idledrifting · 4.0 s window · 100 Hz47% overflow at the mean
Fig. 1 The allocation on the horizontal axis, as a multiple of the mean occupancy. One curve is the share of sampled instants at which the window holds more than the allocation; the other is the mean share of the allocation standing empty. There is no point at which both are small.

Two failures, usually quoted as one

A capacity decision is wrong in two ways and they are not the same way.

Overflow is the share of instants at which the window holds more than the allocation. At those instants a real structure does something — drops an item, falls back to a slower path, grows and pays a reallocation. It is the failure that shows up as an incident.

Idle is the mean share of the allocation that is empty. It is the failure that shows up as a bill, and it never shows up as anything else, which is why it is the one that goes unmeasured.

They move in opposite directions on one dial, and the mean is one point on that dial rather than an answer to it.

The two ways a capacity taken from the mean is wrongAn exact time window of 4.0 s, and an allocation set at 1× its mean occupancy on each stream. The upper bar is the share of sampled instants at which the window holds more than the allocation; the lower is the mean share of the allocation standing empty. The evenly spaced stream pays neither, which is what makes the others readable. The drifting stream overflows at 47% of instants while 35% of the allocation is idle on average — the same decision failing in both directions at once.even0% over · 0% idlePoisson48% over · 2% idlebursty51% over · 2% idledrifting47% over · 35% idleallocation 1× the mean · upper: instants that overflow · lower: allocation idle4.0 s window · 100 Hz · 20,000 arrivals47% over and 35% idle at once
Fig. 2 Both failures across four arrival patterns, at an allocation set exactly at the mean. The evenly spaced stream pays neither, which is what makes the other three readable.

What the mean delivers

At an allocation equal to the mean occupancy, on a four-second window at a mean of a hundred arrivals a second:

The evenly spaced stream overflows at 0 per cent of instants and wastes 0 per cent. It holds exactly four hundred items, always, and the mean is the whole distribution.

The Poisson stream overflows at 47.5 per cent and wastes 2.0 per cent. Half the instants are above the mean, which is what a mean is; the excursions are small, so the waste is small.

The bursty stream overflows at 50.8 per cent and wastes 2.0 per cent — barely different from Poisson, because its bursts are short compared with a four-second window and the window averages them.

The drifting stream overflows at 46.7 per cent and wastes 35.1 per cent. It holds between 105 and 2,169 items against a mean of 951.

What a 4.0 s window actually heldThe occupancy of an exact time window over the a rate that rises and falls process at a mean 100 Hz, sampled at 120 instants. The duration is exact — that is what the model guarantees — and the count runs from 105 to 2,169, a factor of 20.7. The horizontal line is the mean, 951, which is what a structure sized from a single figure would be sized for and is a value the window holds almost never.mean 951items in the windowsampled instantdrifting · 4.0 s · 100 Hz105 – 2,169 against a mean of 951
Fig. 3 What the drifting stream’s window actually held, sampled. The horizontal line is the mean, and the trace crosses it twice a cycle and spends its time far away from it in both directions.

That last row is the one worth stopping on. Nearly half the instants are over the allocation and a third of the allocation is empty on average, at the same time, from the same decision. Those are not competing risks to be balanced — they are both being paid, simultaneously, because the distribution is wide enough to be on both sides of any single number.

The frontier

Give the allocation headroom and the overflow falls. It falls slowly.

On the drifting stream: at 1.5 times the mean, overflow 32.5 per cent and idle 43.5 per cent. At twice, 15.0 and 51.4. At three times, 0 and 66.7.

So stopping the overflow entirely costs an allocation three times the mean, of which two thirds is empty on average. That is the trade, and it is a real one — three times the memory for a structure that never drops anything.

There is no allocation at which both costs are smallA time window of 4.0 s over the poisson stream at a mean 100 Hz. The allocation is the mean occupancy times the headroom on the horizontal axis. The falling curve is the share of sampled instants at which the window holds more than the allocation — where a real structure drops something. The rising curve is the mean share of the allocation standing empty. At a headroom of one, 48% of instants overflow; at the headroom that stops the overflow, 33% of the allocation is idle. A capacity chosen from the mean is a point on this curve rather than an answer to it.0%25%50%75%100%1.5×shareheadroom over the mean occupancyoverflowingstanding idlepoisson · 4.0 s window · 100 Hz48% overflow at the mean
Fig. 4 The same frontier on a Poisson stream, where the whole problem is a different size. Fifty per cent more than the mean stops the overflow completely, at a third of the allocation idle.

The Poisson comparison is the point of drawing both. There the overflow goes to zero at 1.5 times the mean and the idle share is 33 per cent, so the decision is cheap and obvious. The drifting stream needs twice that headroom for the same result.

The cost of a capacity decision is set by how much the rate moves, not by the rate. Both streams have the same mean and the same window; one of them costs three times the memory to serve without dropping anything.

Why the overflow share sits near a half

Three of the four streams overflow at close to fifty per cent of instants at an allocation equal to the mean, and the coincidence is not one.

An allocation at the mean is exceeded whenever the occupancy is above its own mean, and for a roughly symmetric distribution that is half the time. The Poisson stream’s occupancy is very nearly symmetric, so 47.5 per cent. The drifting stream’s is not symmetric at all — it spends long stretches low and shorter stretches very high — and still lands at 46.7 per cent, because the time spent above the mean happens to be about half a cycle.

What differs between them is not how often they overflow but how far. The Poisson stream’s worst instant is 440 against an allocation of 400, ten per cent over. The drifting stream’s worst is 2,169 against 952, 128 per cent over.

So how often is a poor summary of the failure and is the one a percentile rule controls. That is the same distinction expected is not average draws for running times: a frequency and a magnitude are different quantities, and a rule that fixes one is silent about the other.

The two ways a capacity taken from the mean is wrongAn exact time window of 4.0 s, and an allocation set at 2× its mean occupancy on each stream. The upper bar is the share of sampled instants at which the window holds more than the allocation; the lower is the mean share of the allocation standing empty. The evenly spaced stream pays neither, which is what makes the others readable. The drifting stream overflows at 15% of instants while 51% of the allocation is idle on average — the same decision failing in both directions at once.even0% over · 50% idlePoisson0% over · 50% idlebursty0% over · 50% idledrifting15% over · 51% idleallocation 2× the mean · upper: instants that overflow · lower: allocation idle4.0 s window · 100 Hz · 20,000 arrivals15% over and 51% idle at once
Fig. 5 The same four streams at twice the mean. Three of them stop overflowing entirely and one does not, and the idle share of all four has gone past half.

A percentile is the usual answer and it is worth pricing

The standard response is to size at a high percentile rather than the mean, and the measurements say what that buys here.

The 95th percentile of the drifting stream’s occupancy is 2,134 items — 2.24 times the mean. So a p95 rule lands between the two-times and three-times rows of the frontier, at an overflow of a few per cent and an idle share around 55 per cent.

That is a defensible choice and it is not free of the same problem: it is still one number standing in for a distribution, and the quantity it fixes is the frequency of overflow rather than its size. A structure that overflows by ten items at three per cent of instants and one that overflows by a thousand items at three per cent are indistinguishable to a percentile rule, and they are not indistinguishable to anything downstream.

The even stream is the check, not the case

Every plate here draws the evenly spaced stream and it pays nothing at all — no overflow at any allocation from the mean upwards, no idle space at the mean. That row exists to make the others readable and it is worth saying why it is not evidence about anything else.

A rule that overflowed on the even stream would be a rule about the rule rather than about the rate, and the assertion that ships with these measurements checks exactly that: the mean-allocation rule must overflow on the drifting stream at a fifth of instants or more, and must not overflow on the even stream at more than one instant in twenty. Either half failing would mean the finding is not about a moving rate.

That two-sided shape is the habit these measurements are written to. A run is a property of the input makes the same move for sorting: a claim that an algorithm exploits structure has to be paired with an input that has none, or the claim is about the algorithm’s constant factor.

The other window model does not escape it

An arrival-counted window has an exact occupancy by construction, so the sizing problem disappears. It reappears as the problem the window that is even in the wrong currency measures: the duration the structure covers is then the thing that varies, by up to twenty-six times on this stream.

So the choice is between a structure whose memory is predictable and whose meaning is not, and one whose meaning is predictable and whose memory is not. Both are exactly even in one currency and pay for it in the other, and a capacity plan and an alert threshold want different ones.

Neither window is uneven. Each is even in its own currency.Two windows over the same stream, matched so that the time window's duration is the mean duration one block of the arrival window turned out to cover. The upper bar is how much the arrival window's duration varies; the lower is how much the time window's occupancy varies. On the evenly spaced stream both are exactly one and the two models are one object. On the drifting stream they are 26.1× and 17.7×. Choosing a window model is choosing which axis to take the unevenness on, not whether to have it.even1.00× / 1.00×Poisson1.16× / 1.21×bursty1.00× / 1.14×drifting26.14× / 17.69×upper: W arrivals, duration varies · lower: D milliseconds, count varies4,096 arrivals in 8 blocks · 100 Hzmatched at the block, not at the window
Fig. 6 The two spreads over the same streams. A deployment gets to pick which of these two columns is a column of ones.

A system that wants both is buying two structures, or is buying one and accepting that a stated figure somewhere in its documentation is a mean of something that varies by twenty times.

What the summary structures do about it

None of this is specific to an exact window, and it is worth saying where the approximate structures sit.

A summary with a fixed table — kk counters, a sketch of fixed width — does not have an occupancy problem at all, because its allocation is a parameter rather than a consequence. What it has instead is an accuracy that moves with the load: the floor a histogram already knows computes the smallest counter such a table settles at, and that number is proportional to the mass the table is not holding, which rises with the arrival rate.

So the same variation shows up in a third place. An exact time window converts a moving rate into moving memory; an arrival-counted window converts it into moving duration; a fixed-width summary converts it into moving error. The variation is conserved across all three and only its currency changes, which is a statement about what a structure can do rather than about any particular one.

What a timestamp costs, against the length of the windowThe one parameter that moves the clock. A structure in the sliding-window model compares a stored stamp against t − W to decide whether an arrival has left, so its stamps must stay ordered across a wrap — which takes ⌈log₂ 2W⌉ bits, one more than the window itself. From W = 64 to W = 65,536 that is 7 bits rising to 17: a thousandfold longer window for 10 more bits per stamp. It is the accuracy dial that does not touch this quantity, and the window length that does, and the growth is slow enough that the clock is a fixed overhead in every practical setting rather than something to tune.7649256111,024134,0961516,3841765,536window length W, in arrivalsbits per stamp⌈log₂ 2W⌉one bit morethan the windowsliding-window model · a stamp orders arrivals across one wrapcomputed, not measured
Fig. 7 The state four windowed structures hold, split into what each part is for. The two with fixed tables have a flat allocation and the two exact ones do not, which is the same trade at the level of the structure rather than of the window.

The window length is a dial too

One dial has been held fixed throughout and it is the one with the most authority over the answer.

A four-second window on a rate that cycles slowly contains a small slice of the cycle, so its occupancy tracks the instantaneous rate and inherits its whole range. A forty-second window on the same stream contains a whole cycle and holds a nearly constant count, wherever in the cycle it sits.

That is the same averaging the boundary that hides the burst describes at the block boundary, in its benign form: a long window is a low-pass filter, and here the filtering is what a capacity plan wants. Lengthening the window converts an occupancy problem into a staleness problem, and staleness is a quantity somebody has to be willing to name.

So the honest statement of the sizing decision has three terms in it — the window length, the allocation, and the acceptable staleness — where the usual statement has one. Two of them are traded against each other and the third is what the trade is measured in, and none of the three is the mean occupancy.

There is no allocation at which both costs are smallA time window of 4.0 s over the bursty stream at a mean 100 Hz. The allocation is the mean occupancy times the headroom on the horizontal axis. The falling curve is the share of sampled instants at which the window holds more than the allocation — where a real structure drops something. The rising curve is the mean share of the allocation standing empty. At a headroom of one, 51% of instants overflow; at the headroom that stops the overflow, 33% of the allocation is idle. A capacity chosen from the mean is a point on this curve rather than an answer to it.0%25%50%75%100%1.5×shareheadroom over the mean occupancyoverflowingstanding idlebursty · 4.0 s window · 100 Hz51% overflow at the mean
Fig. 8 The bursty stream’s frontier, which looks like the Poisson one because a four-second window contains many bursts and averages them. The shape of the curve is set by how much of the variation the window has already removed.

An exact window is not the only thing being sized

Everything above is about a structure that keeps the arrivals themselves, and that is the case where occupancy and allocation are the same quantity. Two other cases are worth separating, because the same variation reaches them differently.

A structure that keeps a summary of the window — a table of counters, a sketch of fixed width — has an allocation that does not move at all. What moves is the accuracy, because the summary’s error is proportional to the mass it is not holding and that mass rises and falls with the rate. So the failure is not an overflow but a quiet widening of the error bars, at exactly the moments when the load is highest and the answer is most likely to be looked at.

A structure that keeps keys rather than arrivals — a distinct-count over a window — sits between them. Its occupancy is the number of distinct keys in the window rather than the number of arrivals, and on a stream whose popular keys drift those two vary differently: the arrival count follows the rate and the distinct count follows how much the population is turning over.

None of the three is exempt and each converts the variation into a different currency. That is the same pattern the two window models show, one level up, and it is the reason a sizing rule expressed as rate times duration is not portable between them even when it is right about one.

The frontier has a corner, and one ratio puts it where it is

The frontier is drawn as two curves crossing, which invites the reading that the trade is smooth and a judgement has to be made about where on it to sit. It is not smooth. It has a corner, the corner is the only point on the curve worth considering, and its position is fixed by one number.

The idle share at an allocation AA is E[max(0,AX)]/AE[\max(0, A - X)]/A, and that is at least 1E[X]/A1 - E[X]/A with equality exactly when the window never exceeds AA. So the instant the overflow reaches zero, the idle share is pinned:

idle  =  1meanpeak\text{idle} \;=\; 1 - \frac{\text{mean}}{\text{peak}}

and no allocation below the peak has zero overflow while no allocation above it does better on idle. Everything past the corner is pure waste and everything before it is pure risk.

The measurements sit on that identity. At three times the mean the drifting stream never overflows and wastes 66.7 per cent, which is exactly 11/31 - 1/3. At twice the mean it wastes 51.4 against a floor of 50.0, the 1.4 being the instants where the window is over the allocation and the empty space is clipped at zero. The corner itself is at the peak, 2,169 against a mean of 951 — a ratio of 2.28 — so the true price of never dropping anything on this stream is 56 per cent idle, not the 66.7 the sweep’s three-times row reports.

That gap matters more on the well-behaved stream than on the badly behaved one. The Poisson window’s peak is 440 against a mean of 400, so its corner is at 1.10 times the mean and its price is 9.1 per cent idle — against the 33 per cent the frontier’s 1.5-times row shows, because 1.5 is simply the next point on a grid somebody chose. A sweep at halves and doubles cannot see a corner at 1.10, and reading the price off the grid overstates it by a factor of three and a half.

So the practical instrument is not a sweep at all. It is two numbers off one trace — the mean occupancy and the peak occupancy — and their ratio is the whole decision: it is the headroom that eliminates overflow, and its complement is what that costs. The range the essay asks a reader to record is that ratio, and the reason to record it is that it is the price rather than a caveat on the mean.

The peak, of course, is a peak over whatever was observed, which is the one term in this that a longer trace can move. That is the same caution the tuples a summary does not report resolves in the other direction — there the peak is reached once per compression period by construction and is not a tail event at all, so it is a property of the schedule rather than of the sample. Here it is genuinely a sample maximum, and a stream whose rate cycles will produce a larger one given more cycles. Which is expected is not average applied to the statistic that replaces the average: the peak is a better summary and it is still a summary.

What to record with a capacity number

Three things, and none of them is the mean.

The range, because it is the thing the two failure modes are computed from and it is usually a factor rather than a percentage. A twenty-fold range is not an error bar on the mean; it is a statement that the mean is not a description.

Both failure shares at the chosen allocation, because quoting one of them makes the decision look better than it is. On the drifting stream at the mean, the overflow number alone reads as a system that is half-failing and the idle number alone reads as a system with slack; together they read as a decision that is wrong in two directions.

The timescale the rate varies on, relative to the window — the quantity a window that is a duration found governing whether the two window models agree at all. It decides everything else: a window long compared with the variation averages it away and the whole problem disappears, which is why the bursty stream costs no more than the Poisson one here and would cost far more with a window a tenth the length.

What a 4.0 s window actually heldThe occupancy of an exact time window over the Poisson, constant rate process at a mean 100 Hz, sampled at 120 instants. The duration is exact — that is what the model guarantees — and the count runs from 356 to 440, a factor of 1.2. The horizontal line is the mean, 399, which is what a structure sized from a single figure would be sized for and is a value the window holds almost never.mean 399items in the windowsampled instantPoisson · 4.0 s · 100 Hz356 – 440 against a mean of 399
Fig. 9 The Poisson stream’s occupancy over the same window, for scale. This is the case every sizing rule is designed against and it is the one where the mean nearly is the distribution.

The last of those is the practical version of everything in this group of essays. Measuring what an algorithm keeps argued that space has to be instrumented rather than asserted, and a window’s occupancy is the case where the assertion is so natural that nobody notices making it: rate times duration is arithmetic, it is correct about the mean, and it describes a quantity that spends almost none of its time there.

There is one more reason the mean survives as a rule of thumb, and it is worth naming so that the rule can be retired properly rather than argued with. On a stream whose rate is steady, the mean is the distribution to within a few per cent, and the great majority of load a system sees over its life is steady by that standard. The rule works, most of the time, on most streams — and the cases where it does not are the ones where the system is under unusual load, which is when a capacity decision is being tested. A guarantee names its model is the general form; here the model is the rate holds still, and it is nowhere written down.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

What links here

The 8 essays that link to this one and share the most of its objects, of 10 that link here.

The objects this essay names

Each one links to every other essay that touches it.

AllocatorArrival processBurstinessExpiryMeasurementOccupancyParameter choicePeak spaceSliding windowTailTemporal resolutionTime windowTrade off