The usual drawing of a distributed lock is a mutual-exclusion diagram: the lock service grants a lease, the holder works, the holder releases. Nothing in it is wrong and nothing in it protects any data, because the lock service and the resource are different machines and the resource was never asked. What actually reaches the resource is a write, arriving at some wall-clock instant, from a process whose belief about its own lease was formed at an earlier one. Everything interesting lives in that gap.
So this is a simulator rather than an animation. A schedule — think time, work time, when a thread stops, when a clock jumps, how long a write spends on the wire — is drawn once per seed, before anything runs. Then the same schedule is replayed under three postures for the resource. That is the whole experimental design, and it buys the one claim worth making.
The control
both held it counts the moments when two or more clients simultaneously
believed they held the lock. Across 150 trials:
| posture | races | overlap (ms) |
|---|---|---|
| resource accepts anything | 1652 | 3,527,400 |
| client re-checks its lease | 1652 | 3,527,400 |
| resource compares a token | 1652 | 3,527,400 |
Identical. Not close — identical, to the millisecond, and checked seed by seed across 240 seeds with zero mismatches. Fencing does not make the race rarer. Every statement below is about what happens to one fixed set of races.
What changes
Same 150 trials, all three failure sources on, three clients, a 3-second lease:
| posture | writes sent | arrived unlocked | applied anyway | work destroyed | refused | given up |
|---|---|---|---|---|---|---|
| accepts anything | 3749 | 1015 | 1015 | 610 | 0 | 0 |
| client re-checks | 2803 | 74 | 74 | 11 | 0 | 946 |
| fencing token | 3749 | 1015 | 380 | 0 | 635 | 0 |
Work destroyed is the number that matters: a write carrying an older token landing on top of a value a newer token already committed. It is the only column that describes something a human loses.
Three readings fall out.
Fencing is exactly zero, and it is zero for a structural reason. A write is applied only when its token is at least the highest the resource has served, so after it is applied the stored token is the highest ever seen. An older token can never be on top of a newer one, whatever the schedule does. It is not a rate that got small. It is a quantity the arithmetic will not permit.
Fencing does not stop the unlocked writes. 1015 of them still arrive, and 380 are still applied — the ones that arrive when no higher token has written yet. Fencing never claimed to stop them. Its guarantee is an ordering, not an exclusion, and 140 of the 150 trials still contain a write applied without a live lock. If you need “no write without a lock” you need something other than this.
The safety is free and the check is not. Fencing aborts nothing and sends as many writes as the unprotected resource. The expiry check buys its 55× reduction by throwing away 946 writes, 25.2% of every attempt — work done inside a critical section and then dropped.
The three things that open the gap
Each is independently sufficient, and the simulator’s baseline proves it is not manufacturing failures: with all three off, unlocked writes, overlaps and aborts are all zero.
The pause does the damage
A stop-the-world GC, a suspended VM, a descheduled container: the thread stops and the clock does not. The client resumes holding a lease that expired while it was not running. On pauses alone, 1004 unlocked writes are applied and 637 destroy work.
This is the class the expiry check actually fixes, and the fix is real: re-read the lease immediately before writing and 1004 becomes 19, a 98.1% reduction. Worth saying plainly, because the usual telling of this story treats the check as useless. It is not useless. It is 98% of the way there and it costs a 25.7% abort rate, and it is still not zero, because a pause can land in the window between the check and the send. No margin covers that window; a margin protects the interval it measures, and this one comes after.
The clock step breaks nothing on its own
An NTP correction jumping the clock backwards is the failure everyone expects to be dangerous. On its own it produces zero unlocked writes under all three postures. It cannot: the client’s arithmetic is wrong, but nothing is asking it a question whose wrong answer breaches anything.
What the step does is silently cancel the safety margin — 22 times across the combined run, a check that would have aborted instead sent. That is its entire contribution, and it only exists when there is a margin to cancel.
The corollary is worth its own sentence, because it is the opposite of the usual advice. Steady clock drift is not the hazard. A rate error of 1000 ppm — a bad clock — costs 3 ms on a 3-second lease; 200 ppm on a ten-second lease costs 2 ms. Any margin absorbs that. The hazard is the discontinuity, which is why lease arithmetic belongs on a monotonic clock and why “we run NTP” is not an answer to this question.
Network delay needs no race at all
The write leaves while the lease is valid and lands after it is not. On delay alone: 51 unlocked writes, identical under all three postures — the check was right when it ran, and fencing lets them through because no higher token has written. And only 20 races in the whole run, across 48 of 150 trials.
That is the case the standard diagram cannot draw. There is no second client in the picture, nobody’s lease overlaps anybody’s, and the resource still takes a write from a process holding nothing. Mutual exclusion was never the property that was failing. (None of these 51 destroy work, which is the other half of the point: an unlocked write is not automatically a lost one.)
The margin is not a knob, it is a dial between two failures
The expiry check’s only parameter is how much slack it insists on. Turning it up does work, and it charges for it (80 trials each):
| margin | aborted | still destroying work? |
|---|---|---|
| 0 ms | 25.7% | yes |
| 1000 ms | 50.6% | yes |
| 2000 ms | 81.4% | yes |
| 2500 ms | 94.4% | no |
The abort rate rises monotonically and the corruption does reach zero — at a margin that throws away 94% of all writes. At that setting a client may only enter a critical section it can finish in a sixth of its lease, which is to say the lease is no longer doing anything. There is no value of this dial that is both safe and useful, and that is the honest shape of the trade: the check converts data loss into unavailability at a poor exchange rate, and cannot convert all of it.
The precondition nobody states
Fencing’s zero has a hypothesis attached: the token must be monotonic. Switch the lock service to one that fails over without consensus — it comes back having lost the tail of its counter and reissues numbers it has already issued — and the same 150 trials say:
- work destroyed: still 0. The comparison is still a correct comparison.
- writes refused although their lease was live: 348. A new holder is handed a token the resource has already surpassed and is locked out of its own lock. Safety held; availability did not.
- two clients writing under the same token, each overwriting the other’s committed work: 406.
That last one is the interesting failure, because the resource did everything
right. >= passes in both directions when the numbers are equal, so the
comparison sees nothing wrong and there is nothing in the data to notice
afterwards. The number stopped being unique and the mechanism stopped existing,
quietly. (> closes it and costs you every second write inside one lease.)
Which is the actual reason the argument about Redlock is an argument about consensus rather than about timeouts. A lock service can be wrong about who holds the lock and fencing will cover for it. It cannot be wrong about what number it last handed out.
What’s on screen
Time runs left to right; the playhead is a frontier, so nothing to its right is drawn.
- lock svc — the leases as the lock service records them, one bar per grant, labelled with its token. The amber tick is the instant the lease dies, whether or not its holder is listening.
- each client — the bar is the client’s belief that it holds the lock.
Solid while that belief is true; red-hatched the moment it outlives the
lease, which is the subject of the whole piece. Grey hatching is a stopped
thread with its duration on it. An amber caret is a clock jumping backwards. A
dot under the bar is the expiry check; it says
abortwhen the client declines to write. - the resource — a staircase of the highest token it has ever served. Every write is plotted at its own token, so a mark below the line is exactly a write that fencing would refuse — and under the unprotected resource you watch it land anyway. Green applied, red applied-while-unlocked, blue ✕ refused.
- the red wash across all lanes is a race: two clients believing at once.
- the scoreboard carries the same race under all three postures at all times, so the comparison never requires you to remember a previous run.
Notes
src/fence.jsis framework-free and has no rendering in it: a lock service, a clock per client that is allowed to be wrong, a resource that decides, and one tick loop over integer milliseconds.sweepandmarginSweepproduce every table above.- Every number in this write-up is read off the running demo by
scripts/screenshot-demo.mjs, which checks 61 claims on every build. Several of them are deliberately checks that something is nonzero or identical, because those are the claims that would rot silently. - A client that aborts is held for exactly as long as a client that writes. A
real one would let go sooner; paying that small fiction is what makes
overlapsa control instead of a variable, and the control is the piece. - Pauses are only injected while a client holds a lease, since a pause outside a critical section cannot matter. Work is drawn as a fraction of the lease, up to 0.95, so a section never overruns its own lease — whatever breaches a lease here is one of the three sources and not a client that simply took too long.
- Releasing is done by token comparison too: a client whose lease expired and was regranted must not unlock the stranger now holding it. Same bug as the write, one layer up, same fix.