workshop private

← all creations

In Phase

viz · created 2026-10-04

A retry storm, simulated rather than diagrammed, so the claim can be counted instead of asserted. Sixty clients, a three-second outage, and a dependency that could clear the whole backlog in 900 ms of work — replayed under five retry policies over identical demand. Exponential backoff with no jitter takes 25.6 s and spends 95.8% of that recovery with nothing in flight at all; full jitter takes 6.0 s and spends 61 calls, one per client. Jitter is not a refinement of backoff. Backoff without it is the thing making the outage long.

architecturesimulationinterview-prepcanvas

Everyone knows to back off exponentially. The received wisdom draws one client and one server, a wait that doubles, and the moral that a polite client stops hammering a sick dependency. Nothing in that picture is wrong and nothing in it is the problem, because the problem only exists in the plural. A retry storm is a property of a cohort, and a cohort is exactly what a shared failure manufactures: everybody failed at the same instant, so everybody is now holding the same clock.

Doubling a wait does not change that. It changes the period and leaves the phase alone — and the phase is the whole hazard.

So this is a simulator rather than an animation. Sixty clients each need one success from a dependency that is hard-down for three seconds and then healthy with a concurrency limit of 8. The same demand is replayed under five retry policies. That is the entire experimental design, and it buys the one claim worth making.

The control

Every policy completes exactly the same work: 60 admissions, 7.2 seconds of service time, not one request more or less. The dependency is identical, the outage is identical, the cohort is identical. Nothing varies except when clients choose to ask. Across 40 seeds:

backoffcallsin recoverypeak /100 msdependency idleall done
flat retry225245260.036.8%4.5 s
exponential55625660.095.8%25.6 s
equal jitter421618.341.1%7.1 s
full jitter457617.715.7%6.0 s
decorrelated442639.319.4%6.0 s

The dependency can serve 8 concurrent requests at 120 ms each, so the entire cohort is 900 ms of work. Exponential backoff turns 900 ms of work into 25.6 seconds, and spends 21.7 of the 22.6 recovery seconds with zero requests in flight. Full jitter spends 61 calls on the recovery — sixty clients, one call each, give or take one — and finishes in 6.0.

Read those two rows together, because the pairing is the point: exponential backoff manages to overload and starve the same machine. Sixty calls land inside one millisecond, which is 7.5× what the dependency could absorb in a hundred. Then nothing for 3.2 seconds. Then sixty again.

Backoff does not de-correlate. It preserves.

The easiest way to see this is to stop assuming the cohort fails on a single tick. Spread the initial failures over a window and watch what survives (exponential, 40 seeds):

initial spreadwidth of the worst burstpeak /100 msall done
0 ms1 ms60.025.6 s
50 ms49 ms60.025.7 s
200 ms194 ms35.017.9 s
400 ms387 ms20.911.1 s
1600 ms1551 ms8.96.9 s

The burst is always as wide as the spread it started with. Not approximately — to within 3%, at every setting, after six doublings and a cap. A deterministic wait is a translation: it moves every client in the cohort by the same amount and therefore moves none of them relative to each other. The period grows to 3.2 seconds and the burst stays 49 ms wide forever.

Which is why the fix is not a longer wait. The fix is the only operation that changes relative position, and that is noise.

The dumbest policy is the fastest one

Flat retry — try again every 100 ms, forever, no backoff at all — finishes in 4.5 seconds, sooner than anything else on the board. It costs 2252 calls to do it, five times full jitter’s, and 1800 of those are thrown at a dependency that is hard-down and refusing everything.

This is worth stating plainly because the usual framing has backoff as strictly better behaviour. It is not: backoff buys the dependency protection by spending the client’s latency, and that is a real trade with a real price. What jitter does is make the price reasonable. At a 400 ms cap, full jitter finishes in 4.2 s — faster than flat retry — for 1168 calls; at an 800 ms cap it finishes in 4.5 s, matching flat retry to within 30 ms, for 712 calls. Same recovery, a third of the load.

backoff capexponentialfull jitter
400 ms796 calls · 68% idle · 6.0 s1168 calls · 1% idle · 4.2 s
800 ms616 calls · 84% idle · 8.8 s712 calls · 4% idle · 4.5 s
1600 ms556 calls · 92% idle · 14.4 s518 calls · 11% idle · 5.0 s
3200 ms556 calls · 96% idle · 25.6 s457 calls · 16% idle · 6.0 s
6400 msnever finishes (40/40)452 calls · 38% idle · 8.9 s

A shorter cap mitigates the phase-locked case and never fixes it — the idle fraction is still 68% at the setting where exponential is at its best, because a cohort that sleeps together leaves the machine empty together however briefly it sleeps. And the direction of the cap is a genuine trade for jittered clients and a free lunch for nobody: a tighter cap costs full jitter calls and buys it time.

The outage is not required

Delete it. Start sixty clients at once against a healthy dependency with a limit of 8 and nothing else wrong:

backoffcallspeak /100 msall done
flat retry45260.01.5 s
exponential30860.012.8 s
full jitter236125.32.1 s
decorrelated16460.01.9 s

Exponential still takes 12.8 seconds to clear 900 ms of work. The outage was never the cause; it was one way of creating a cohort, and a deploy, a cron minute boundary, a cache expiry or a client library whose clients all started together will do it just as well. Phase comes from a shared event, and failure is only the most memorable kind.

There is an honest caveat sitting in that table, too. Full jitter’s peak here is 125 per 100 ms — worse than exponential’s 60. The first attempt is simultaneous for everybody no matter what policy they hold, because no backoff has happened yet; jitter cannot smooth a burst that precedes it. All it can do is make sure there is never a second one. (Decorrelated jitter looks best on this table because its first wait is already drawn from a range rather than fixed.)

Jitter does not create capacity

Push it past what the dependency can do — 240 clients against a limit of 4, a floor of 7.2 seconds of unavoidable work (20 seeds, 60-second horizon):

backoffcallspeak /100 msall done
flat retry21600240.014.9 s
exponential4908240.0never, in 20 of 20
equal jitter199722.814.9 s
full jitter226224.613.6 s
decorrelated232931.113.3 s

Every jittered policy lands at roughly twice the work floor, and so does flat retry. Jitter buys nothing in completion time here — the dependency is the constraint and no scheduling of the asking changes that. What it buys is the same 14 seconds for 2262 calls instead of 21600, which is the difference between a saturated dependency and a dependency that is saturated and being asked ten times as often as it can answer.

Exponential, meanwhile, does not recover at all. 168 of the 240 clients are still waiting after a full minute — the same 168 in all 20 seeds — while the machine they are waiting for sits idle 96% of the time.

What’s on screen

Time runs left to right on one ruler, shared by every policy and never rescaled — it spans the slowest policy’s finish, so a run that ends in the first fifth of the panel is making its argument with the empty four fifths. The grey block at the left is the outage. The playhead is a frontier; nothing to its right is drawn.

Notes