One pass, and no room

A detector that learns its own quiet

A burst detector built from a fast decayed counter needs a threshold, and the one it was measured with came from sixty independent runs of the background — an oracle no deployed system has. Two counters and a running mean square of their difference can supply it from the stream itself, and on a stream that does not move they hit a 5% false-alarm rate at every setting tried. What they lose is the burst: a scale that learns every reading learns the burst as noise and forgets it in 20 to 40 seconds whatever its size. Holding the scale while alarmed brings the oracle's horizon back for small bursts, turns a slow drift into an alarm that will not stop, and still cannot see past the lag at which the slow counter holds more of the burst than the fast one.

The counter with no window in it kept a rate as one number that fades by half every half-life, and found it settles on a steady stream at exactly the count of a window 1.44 half-lives long. What a heavier tail actually buys asked how long such a counter keeps a single burst visible. It matched each decay scheme to the same 5% spread on a steady stream of ten arrivals a second — for the exponential, a half-life of 13.3 seconds — added a burst, and measured the lag at which the estimate stopped standing out. “Standing out” needed a threshold. The page took it from an oracle: sixty independent runs of the same background with no burst, and the 95th percentile of the estimate across them at each lag.

Its closing section named that oracle as the thing a deployed detector does not have. A real system watches one stream, and the spread it must compare against belongs to an arrival process that may itself be drifting. The section proposed a structure that supplies its own threshold. It keeps the fast counter, adds a second at a much longer half-life to estimate the background, and uses the slow counter’s own variation as the spread — “two values and a stamp”. It asked two things: how much horizon is lost against the oracle, and what happens when the background drifts slowly, since a burst and a drift differ only in their timescale.

The detector can be built and it calibrates itself. The horizon it loses is most of the horizon, for a reason that the obvious repair turns into a different failure. The question the section asked last — whether one structure can tell a burst from a drift — has a plain answer, and it is no.

The detector, and the stream it calibrates on

The detector keeps three numbers. The fast rate ff is the earlier page’s matched exponential counter, half-life H=13.3H = 13.3 seconds. The slow rate ss is the same counter with half-life RHRH. The scale vv is an exponentially weighted mean square of the difference d=f−sd = f - s, with the slow half-life. It is read once a second and alarms when d>zvd > z\sqrt{v}, with vv as it stood before the current reading, so the reading being judged is not yet part of the scale it is judged against. z=1.645z = 1.645 is the one-sided 5% point of a normal distribution, and the false-alarm rate that results is measured, not assumed. The detector calibrates for its first minute, starting both counters at the rate of that minute, which removes a start-up transient without an oracle. Streams are long enough for the slow counter to settle, twelve slow half-lives and more.

On a stream that does not move, the detector calibrates itself: with its scale learning every reading it alarms on 4.8%, 4.8%, 5.3%, 5.6%, 5.2% of readings at R = 4, 8, 16, 32, 64, against the 5% its threshold is set for; held while alarmed, the scale reads low and z has to rise from 1.645 to 1.85 to bring the rate backThe share of once-a-second readings at which the detector alarms, on twenty-four streams of a steady background of ten arrivals a second, after the counters and scale have settled, against R. Scale learns, z = 1.645: R 4 4.8%, R 8 4.8%, R 16 5.3%, R 32 5.6%, R 64 5.2%. Scale held, z = 1.645: R 4 8.2%, R 8 7.5%, R 16 8.3%, R 32 8.7%, R 64 7.7%. Scale held, z = 1.85: R 4 5.0%, R 8 4.8%, R 16 5.3%, R 32 5.7%, R 64 5.1%. The dashed line is 5%. The horizontal axis is logarithmic.48163264slow half-life, in fast half-lives (R)readings that alarm0.0%2.5%5.0%7.5%10.0%scale learns, z = 1.645scale held, z = 1.645scale held, z = 1.85a steady stream, 24 runsdashed: the 5% it is set for
Fig. 1 The share of once-a-second readings that alarm on a steady background of ten arrivals a second, twenty-four streams, against R, the slow half-life in fast half-lives. Scale learning every reading, z = 1.645: 4.8%, 4.8%, 5.3%, 5.6% and 5.2% at R = 4, 8, 16, 32 and 64. Scale held while alarmed, z = 1.645: 7.5% to 8.7%. Scale held, z = 1.85: 4.8% to 5.7%.

On a stream that does not move, the detector calibrates itself: it alarms on 4.8% to 5.6% of readings at every R, against the 5% its threshold was set for. No independent runs were used. The difference between two counters on one stream has a spread the stream itself reveals, and a running mean square of that difference is a sufficient estimate of it at every ratio of half-lives tried. The difference is not quite normal, since it is a difference of Poisson-driven sums, but it is close enough that the normal’s 5% point gives 5%.

The scale it learns can be checked against a formula. A decayed counter with rate constant k=ln⁡2/Hk = \ln 2 / H on a Poisson stream of rate λ\lambda has variance λk/2\lambda k/2, and two counters on the same stream share arrivals, so their difference has variance λ (k1/2+k2/2−2k1k2/(k1+k2))\lambda\,(k_1/2 + k_2/2 - 2k_1k_2/(k_1+k_2)). At ten arrivals a second that predicts a spread for dd of 0.342, 0.464 and 0.499 arrivals a second at RR = 4, 16 and 64. The detector’s own scale, averaged over eight settled streams, is 0.336, 0.456 and 0.502. The running mean square finds the same number the formula does, without being told the rate or that the stream is Poisson. The spread grows with RR because a slower background counter shares less of the fast one’s noise, so less of it cancels in the difference. The error of a difference found the same thing for subtracted sketches: what a difference’s error depends on is how much the two estimates share.

That half of the proposal works as stated, and it is the half that would have been hardest to believe without measuring. A detector can set its own false-alarm rate from its own readings. The other half is the horizon.

The burst, and the scale that learns it

One stream through a burst of 900 at R = 16: the difference leaps to 42.6 arrivals a second and decays with the fast counter; the learning scale takes the leap into itself and its threshold rises past the difference 24 s after the burst, while the held scale stays put and the alarm lasts 51 sOn one stream of ten arrivals a second with 900 extra arrivals in the second at time zero: the fast counter less the slow one (R = 16), and the threshold each detector compares it with, once a second from 30 s before the burst to 120 s after. The learning detector's threshold goes from 0.9 before the burst to a peak of 11.6; the held detector's stays at 0.9 until its alarm ends. The learning detector stops alarming 24 s after the burst, the held one 51 s. Values above 12 are drawn at the top edge.0510050100seconds after the burstarrivals a secondfast less slowthreshold, scale learnsthreshold, scale heldone stream, R = 16dashed: the burst
Fig. 2 One stream with 900 extra arrivals in the second at time zero, R = 16: the fast counter less the slow one, and each detector’s threshold, once a second. The difference leaps to 42.6 arrivals a second and decays with the fast counter. The learning scale’s threshold rises from 0.9 to a peak of 11.6 and crosses the difference 24 s after the burst. The held scale’s threshold stays at 0.9, and the alarm lasts 51 s.

The trace shows what goes wrong. At the burst the difference leaps from its usual fraction of an arrival a second to 42.6. The detector alarms at once. At the next reading, the scale has taken in the square of that leap. A mean square with a slow half-life of 213 seconds moves only a little at each reading, but the leap is nearly fifty times the threshold’s usual size, and its square over two thousand times the usual mean square. The threshold climbs more than tenfold, to a peak of 11.6. The difference is falling at the fast counter’s rate, halving every 13 seconds, and 24 seconds after the burst it passes under a threshold that is now measuring the burst itself. The scale has learned the burst as part of the stream’s ordinary variation. The held scale, which took none of the leap in, keeps its threshold at 0.9, and its alarm ends 51 seconds after the burst. By then the difference itself has fallen under that unchanged threshold, a few seconds before the lag at which, as the next section shows, it would have turned negative anyway. The two detectors see the same difference; they disagree only about what to compare it with.

How long a burst stays visible without an oracle: a detector whose scale learns every reading holds a burst for 21.4 s at R = 16 and 37.6 s at R = 64 whatever the burst's size; holding the scale while alarmed reaches 52.7 s against the oracle's 54.8 at a burst of 300, and 78.4 against 119.1 at 9,000The lag after a burst, in seconds, at which half of sixty paired streams still alarm, against the burst's size in arrivals within one second, on a background of ten arrivals a second. The oracle's threshold: 300 54.8 s, 900 79.0 s, 3,000 105.0 s, 9,000 119.1 s. Scale held, R = 64: 300 52.7 s, 900 73.0 s, 3,000 78.4 s, 9,000 78.4 s. Scale held, R = 16: 300 41.2 s, 900 54.7 s, 3,000 55.4 s, 9,000 55.4 s. Scale learns, R = 64: 300 37.6 s, 900 39.2 s, 3,000 39.2 s, 9,000 39.2 s. Scale learns, R = 16: 300 21.4 s, 900 20.3 s, 3,000 19.6 s, 9,000 19.6 s. R is the slow counter's half-life in fast half-lives. Both axes are logarithmic; dashed lines are R = 16.3009003,0009,00010100arrivals in the bursthalf-detection horizon, sthe oracle's thresholdscale held, R = 64scale held, R = 16scale learns, R = 64scale learns, R = 16ten arrivals a second, sixty streamsdashed: R = 16
Fig. 3 The lag at which half of sixty paired streams still alarm, against the burst’s size. The oracle’s threshold: 54.8 s at a burst of 300, 79.0 at 900, 105.0 at 3,000, 119.1 at 9,000. Scale learns, R = 16: 21.4, 20.3, 19.6, 19.6. Scale learns, R = 64: 37.6, 39.2, 39.2, 39.2. Scale held, R = 16: 41.2, 54.7, 55.4, 55.4. Scale held, R = 64: 52.7, 73.0, 78.4, 78.4.

A learning scale’s horizon does not grow with the burst: 20 seconds at R = 16 and 39 at R = 64, whether the burst is 300 arrivals or 9,000. A larger burst leaps higher and inflates the scale in proportion, so the ratio between them — which is all the alarm tests — is the same. The oracle’s threshold does not move with the burst, and its horizon grows as the earlier page found for an exponential: by a fixed number of seconds for every factor of ee in the burst, 55 seconds at 300 and 119 at 9,000. Against it, the learning detector gives up between a third of the horizon, at the smallest burst and R = 64, and five sixths of it, at the largest and R = 16.

This is not a defect of the proposal’s arithmetic. It is the proposal doing what it said. The scale is the spread of dd over the recent past, and a burst is part of the recent past. A detector that knows only its own stream has no way to say that one excursion was not variation, except by deciding so — which is what an alarm is.

Holding the scale, and its ceiling

The obvious repair acts on that decision. While the detector is alarmed, the scale is not updated: a reading judged to be a burst is kept out of the scale the next reading is judged against. That biases the scale low, because the largest ordinary excursions are now excluded as well. The false-alarm rate on a still stream rises to 7.5–8.7%. Raising zz once, to 1.85, for every RR at once, brings it back to 5%. That is a constant of the design, fixed from a still stream in advance, not a threshold drawn per stream — the kind of constant the threshold somebody chose went looking for in library code, except that this one has a measurement behind it.

With the scale held, the horizon comes back. At a burst of 300 and R=64R = 64 it is 52.7 seconds against the oracle’s 54.8, which is essentially the oracle. At larger bursts the held detector’s horizon grows at first and then stops: 73.0 seconds at 900, then 78.4 at 3,000 and at 9,000. At R=16R = 16 it stops at 55.4.

The ceiling has a closed form, and it does not depend on the scale at all. A burst of BB arrivals adds Bln⁡2/H⋅2−L/HB \ln 2 / H \cdot 2^{-L/H} to the fast rate LL seconds later, and Bln⁡2/(RH)⋅2−L/(RH)B \ln 2 / (RH) \cdot 2^{-L/(RH)} to the slow one. The difference is positive only while the first exceeds the second, which holds until

L∗=Hlog⁡2R⋅RR−1.L^* = H \log_2 R \cdot \frac{R}{R-1}.

At H=13.3H = 13.3 seconds that is 56.8 seconds for R=16R = 16 and 81.1 for R=64R = 64, against 55.4 and 78.4 measured. Past L∗L^* the slow counter holds more of the burst than the fast one, and the difference the detector watches has turned negative whatever the burst’s size. A larger RR raises the ceiling, but only by the logarithm of RR. The formula also says what reaching the oracle’s horizon would take. To keep a burst of 9,000 visible for the oracle’s 119 seconds, the ceiling must be at least that, which needs log⁡2R\log_2 R near 9 — a slow half-life of about five hundred fast ones, almost two hours. A background estimate that slow cannot follow any change in the stream that happens within a working day, which is the drift problem below in its strongest form. A decay measured from where it started found that a counter’s response to a step is set by the shape of its weights. Here the response that matters is the difference of two counters, and its shape is fixed by their ratio of half-lives alone.

A burst of 900 at R = 16, detection against lag: every design sees it for the first 16 s; the learning scale loses it by 24 s, the held scale by about 64, and the oracle's threshold keeps half the streams alarmed until 79.9 sThe share of sixty paired streams alarming at each lag after a burst of 900 arrivals, R = 16. The oracle's threshold: 2 s 100%, 4 s 100%, 8 s 100%, 16 s 100%, 24 s 100%, 32 s 100%, 48 s 100%, 64 s 98%, 96 s 10%, 128 s 10%, 192 s 5%. Scale held: 2 s 100%, 4 s 100%, 8 s 100%, 16 s 100%, 24 s 100%, 32 s 100%, 48 s 92%, 64 s 0%, 96 s 0%, 128 s 0%, 192 s 0%. Scale learns: 2 s 100%, 4 s 100%, 8 s 100%, 16 s 100%, 24 s 15%, 32 s 0%, 48 s 0%, 64 s 0%, 96 s 0%, 128 s 0%, 192 s 0%. The horizontal axis is logarithmic.248163264128seconds after the burststreams alarming0%25%50%75%100%the oracle's thresholdscale heldscale learnsR = 16, sixty streamspaired with no-burst runs
Fig. 4 The share of sixty paired streams alarming at each lag after a burst of 900 arrivals, R = 16. The oracle’s threshold: 100% up to 48 s, 98% at 64 s, 10% at 96 s. Scale held: 100% to 32 s, 92% at 48 s, none at 64 s. Scale learns: 100% to 16 s, 15% at 24 s, none after.

At a single burst size the three designs agree completely for the first 16 seconds, and part in order: the learning scale by 24 seconds, the held scale between 48 and 64, and the oracle between 64 and 96. The held detector’s loss against the oracle is the ceiling, and the learning detector’s loss is the scale. Detection falls from all to nothing within one step of the lag grid in both detectors, because each fails for a deterministic reason — the difference passing under the scale, or under zero — rather than by noise.

A drift is a burst that does not end

The closing section’s second question was drift. A background whose rate rises slowly makes the fast counter run ahead of the slow one by a steady amount: the slow counter lags the rise by about its own half-life. To the detector, that looks like a small burst that never decays. A window that is a duration found that on a drifting stream even the meaning of recent shifts — a window of arrivals and a window of time disagree — and a detector comparing two timescales of the same stream meets the same shift as a signal.

A drift is a burst that does not end: at R = 16, with the background rising 10% every 1,000 s, the learning detector alarms on 9.5% of readings, the held one on 18.3%; at a rise of 40% the held one alarms 63.3% of the time, and a limit of 60 s on the hold brings it to 14.8%The share of readings that alarm on streams whose background rises linearly by the stated share of its starting rate every 1,000 seconds, R = 16, twenty-four streams each. Scale learns: 5% 7.9%, 10% 9.5%, 20% 10.0%, 40% 6.6%. Scale held: 5% 11.2%, 10% 18.3%, 20% 32.4%, 40% 63.3%. Held for at most 60 s: 5% 11.1%, 10% 17.5%, 20% 25.8%, 40% 14.8%. On a still stream all three alarm on about 5%. The horizontal axis is logarithmic.5%10%20%40%rise in the background rate every 1,000 sreadings that alarm0.0%5.0%20.0%40.0%60.0%scale learnsscale heldheld for at most 60 sR = 16, 24 streamsdashed: 5%
Fig. 5 The share of readings that alarm on streams whose background rises linearly by 5%, 10%, 20% or 40% of its starting rate every 1,000 seconds, R = 16. Scale learns: 7.9%, 9.5%, 10.0%, 6.6%. Scale held: 11.2%, 18.3%, 32.4%, 63.3%. Held for at most 60 s: 11.1%, 17.5%, 25.8%, 14.8%.

The learning detector survives a drift; the held one does not. With the background rising 10% every 1,000 seconds, the learning detector alarms on 9.5% of readings, twice its target. The steady offset between the counters enters its scale and raises its threshold with it. At steeper rises the offset dominates the scale so completely that the alarm rate falls again, to 6.6%: a drift the scale has learned entirely is invisible. The held detector alarms on 18.3% at a 10% rise and 63.3% at 40%. The first time the offset pushes it over its threshold it stops learning, and the offset keeps it over, so it alarms for good. Limiting the hold to 60 seconds lets it relearn and brings the steepest case down to 14.8%. At the gentler rises, where the offset is comparable to the ordinary spread and the detector crosses in and out, it helps little.

So the two versions of the detector fail in opposite directions, and each fails where the other succeeds. The learning scale treats every sustained excursion as the new normal: it adapts to drift and forgets bursts. The held scale treats every excursion as an event until told otherwise: it remembers bursts and cannot adapt to drift. A limit on the hold is a timescale that decides which of the two an excursion is, and a burst and a drift differ in exactly that — the timescale. The proposal predicted the problem in its last sentence, and the measurement puts numbers on it. There is no setting of this structure that holds a burst of 900 for the oracle’s 80 seconds and alarms on a 10% rise per 1,000 seconds at no more than 5%.

What was measured and what was not

Poisson backgrounds at ten arrivals a second. Every stream is Poisson with a steady or linearly rising rate. A bursty background — arrivals in clumps — has a larger spread of dd than Poisson, which the learning scale would learn correctly, and larger ordinary excursions, which the held scale would mistake for events. The formula that matched the learned scale so closely is a Poisson formula, and on a clumped stream it would no longer be the thing to check against.

One burst shape. Every burst is its arrivals spread over one second, as on the earlier page. A burst spread over a minute is closer to a drift and would be forgotten sooner by the learning scale and held longer by the held one.

Readings once a second. The detector is read once a second. It could be read at every arrival; the scale would then be a mean over arrivals rather than over time, and on a rising background it would weigh the rise more.

The oracle is the earlier page’s, and it is generous. Its threshold is the 95th percentile of the fast counter across sixty independent streams at exactly the lag being tested, which assumes the detector knows both the background’s distribution and how long ago the burst happened. The self-calibrating detectors know neither, and the gap between them and the oracle is partly the price of not knowing when to look.

Three numbers of state. The proposal said two values and a stamp; the detector here keeps three values — the two rates and the scale — and a stamp. The fading nobody computes showed that decayed values need no work between arrivals, and that holds for all three.

Still open: a scale that forgets bursts on a schedule

The summary that has to forget found that a structure asked about only the recent past must forget on a schedule, and pays bits for the schedule that nothing else pays. The scale here is such a structure, and the question is what its schedule should be.

The two failures suggest a structure between them. The held scale fails on drift because it never learns what it has once alarmed on. The learning scale fails on bursts because it learns them immediately. A scale that learns what it has alarmed on, but only after the alarm has lasted longer than a burst could — longer than the ceiling L∗L^* — would forget nothing a burst could explain and learn everything a burst could not.

The measurement that follows sets the hold’s limit at L∗L^*, 57 seconds at R=16R = 16 and 81 at R=64R = 64, and repeats both measurements. The prediction is that the burst horizon stays at the held scale’s ceiling, since no burst’s alarm outlasts it, and that on drift the alarm rate falls towards the learning scale’s, since the relearning starts as soon as the excursion has outlasted anything a burst could produce. It could fail on the gentle rises, where the offset is small enough that the detector crosses in and out of alarm and never holds long enough to relearn. If it does, the right quantity to wait for is not how long the alarm has lasted but how long the difference has stayed positive, which a burst’s difference, decaying and then turning negative past L∗L^*, never does for long.

Named alongside this one

Essays reaching for the same objects. Nobody chose these; they are what the concept index makes visible.

The objects this essay names

Each one links to every other essay that touches it.

Design parameterEstimatorExponential decayFalse alarmHalf lifeHonest limitMeasurement designState bitsStreaming modelVariance