workshop private

← all creations

Catch Up

viz · created 2026-09-30

Replica lag, simulated rather than diagrammed, so that read-your-writes and monotonic reads can be counted separately instead of blurred into 'eventual consistency'. Six routers replay one offered workload. Session affinity makes monotonic reads exactly zero and does nothing at all for read-your-writes; LSN routing does exactly the opposite; requiring both closes both. The bill is the finding: the correct router keeps its zero by quietly sending reads to the primary — 11% of them on average, 83% in its worst 100 ms, and 4790 reads per second at the moment a heavy record freezes every replica at once, against a primary rated for 4000.

architecturesimulationinterview-prepcanvas

The usual drawing of read scaling is a fan: writes to the primary, reads out to the replicas, an arrow labelled async replication in between. Nothing in it is wrong and nothing in it says what a session is allowed to read, because the arrow is not instantaneous and the diagram never puts a number on it. What actually happens is that a session writes at the primary, reads a moment later from a replica, and is shown a world in which its own write has not happened.

So this is a simulator rather than an animation. An offered workload — when each write arrives, which session issued it, how long that session waits before each of its reads — is drawn once per seed. Then the same workload is replayed under six routers. Across 60 seeds the write stream is bit-identical and the read count is identical, so nothing below is a comparison between two different days.

Two guarantees, and each cheap fix buys exactly one

The thing worth counting separately is that “stale read” is two different bugs.

16 trials, 460,800 reads, five replicas, 600 writes/s offered:

routerblind to its own writeworld went backwardsreads on the primaryworst 100 msread p99
any replica25.08%14.74%0.00%0%4 ms
session affinity25.27%zero0.00%0%4 ms
primary for 500 ms after a write5.32%8.36%54.90%100%4 ms
LSN ≥ my writezero3.42%10.91%81%4 ms
LSN ≥ my write and everything seenzerozero11.39%83%4 ms
wait for a replica with bothzerozero2.05%56%404 ms

Read the first two rows together, because they are the whole point.

Session affinity is exactly zero on monotonic reads and buys nothing at all on read-your-writes — 25.27% against 25.08%, marginally worse than routing at random. It is zero for a structural reason: one replica’s applied position only ever increases, and a session’s reads are sequential, so it cannot be shown anything it has already been shown. And it is useless for the other guarantee for an equally structural reason: your write went to the primary, and your replica is a different machine that has not got it yet. Pinning a session to a replica does not pin it to its own data.

LSN routing does precisely the opposite. Require a replica that has applied your last write and read-your-writes is exactly zero, in every trial — a write is served only by a node whose applied position already passed it, so it is not a rate that got small, it is a quantity the comparison will not permit. And monotonic reads are still broken, 3.42%, because a session that happened to be shown record 9,000 by a fast replica is still allowed onto a replica sitting at 8,400: nothing in ”≥ my own write” mentions what it has been shown. Requiring both closes both, and costs 0.5 points of primary share.

Two orthogonal guarantees, two cheap fixes, each buying exactly the one the other misses. “Eventual consistency” names neither of them.

The middle row is the one people ship

Reading the primary for a window after each write is the fix that gets written without a design document, and it is the only row that is expensive and still wrong: 54.9% of every read taken off the replica fleet, and 5.32% of reads still blind, because a time window is a guess about a quantity you are not measuring. When lag exceeds the window the guarantee silently lapses.

The dial has no good setting (600 writes/s, 8 trials):

windowreads on the primaryblind to own writewent backwards
0 ms0.00%24.57%14.10%
100 ms19.43%15.82%15.15%
250 ms36.89%9.77%12.00%
500 ms54.74%4.98%8.10%
1000 ms74.31%1.81%4.08%
2000 ms90.48%0.31%1.17%
4000 ms98.71%0.00%0.06%

It does reach zero. It reaches zero at a setting that has taken 98.71% of the reads off the replicas, which is to say it reaches zero by not having replicas. Everything in between is paying most of the cost for part of the guarantee.

And note row two. A 100 ms window makes monotonic reads worse than having no window at all — 15.15% against 14.10%. It shows the session the front of the log and then takes it away, which is a violation the session would never have seen if nobody had tried to help.

The control for all of this: set replication to instantaneous, where there is nothing to protect against and every router scores zero. The window still sends 54.9% of the reads to the primary. It cannot tell, because it never looks.

What the bill actually is

Every router here is a router, not a replica, so when no replica qualifies the read has to go somewhere. That is the number nobody configures and no dashboard shows, and it moves with load (6 trials per point):

offered writes/slag meanlag p99LSN router: reads on primaryworst 100 mspeak reads/s at the primary
200116 ms860 ms7.84%77%460
600315 ms2910 ms11.49%83%1600
1200891 ms7437 ms18.92%92%3760
16001148 ms9234 ms25.87%95%4790
20001507 ms10246 ms35.51%98%6690

The primary in this simulation is one server rated at 4000 ops/s. At 1600 writes/s the correct, provably-zero-violation router asks it for 4790 reads/s in its worst 100 ms while it is also taking 1600 writes/s — 6390 offered against 4000. Nothing in the configuration says “send 4790 reads per second to the primary”. It is the arithmetic of a correctness rule meeting a lag distribution, and it arrives as a capacity event.

The average is not the number that hurts you. 11.49% mean against an 83% worst bin: the capacity plan reads the first and the outage is caused by the second.

Why the burst is a burst: replicas fail apart and together

Replicas fall behind for two different reasons and only one of them is diluted by buying more of them.

worldblind reads (any replica)lag meanlag p99LSN router on primaryworst 100 ms
replication is instant0.00%0 ms0 ms0.00%0%
identical replicas, nothing wrong0.88%11 ms14 ms0.84%5%
unequal replicas, one straggler that stalls6.67%49 ms547 ms0.59%5%
…plus heavy records25.08%320 ms2640 ms11.39%83%

A straggler is an independent failure: one replica is slow, the others are not, and a router only ever needed one healthy replica. A heavy record — one 200,000-row statement, one schema change — is a correlated one, because it is in the log, and every replica replays the same log, serially, while the primary committed it once with a parallel plan and hundreds of backends. (That asymmetry is the modelling assumption doing the most work here, and it is the one real systems actually have.)

So the decisive experiment is fleet size. 8 trials, 2 replicas to 6:

worldblind reads, 2 → 6 replicasLSN router’s worst 100 ms, 2 → 6
identical replicas0.86% → 0.87%5% → 5%
straggler13.67% → 5.85%3% → 3%
straggler + heavy records42.67% → 22.85%83% → 83%

Three readings, and the first is the one to take away.

With identical replicas, tripling the fleet changes the blind-read rate by 0.01 points. Replicas are read capacity. They are not freshness, and nothing about adding them makes the log arrive sooner. The only reason the middle row improves is that a bigger fleet dilutes one sick member — which is a real benefit and an entirely different one.

And the correlated burst does not move at all. 83% at two replicas, 83% at six. When the reason no replica qualifies is a record every replica has to replay, there is no Nth replica that is doing any better, so the fallback load is exactly the same however many you have bought. This is the answer to “we’ll add read replicas”: it buys throughput, it does not buy the guarantee, and it does not buy down the fallback spike.

The other bill: the primary is also committing the writes

Everything above gives primary-bound reads their own path. Let them share the one server with the write path instead (1400 writes/s, 8 trials):

routingreads on primarywrite latency meanwrite p99primary offered
any replica0.00%0.32 ms1 ms1400/s
primary for a window63.98%155.90 ms404 ms~4090/s, saturated
LSN ≥ my write & seen20.04%8.46 ms154 ms~2240/s

A read consistency mechanism, with no write path in it anywhere, moved mean write latency from 0.32 ms to 155.90 ms — 487×. The window posture pushes the primary past its rating on reads that a replica could have answered if anyone had asked it whether it could.

But the honest version of this has a second half. The loop does not run away. Mean replica lag goes 1018 ms → 1000 ms when contention is switched on: the delayed commits do not measurably feed back into lag, because lag here is dominated by replay cost rather than by when the record was committed. The correct story is not a spiral. It is a conversion, at a bad exchange rate: your read load becomes write latency, which is a queue you never put it in.

Waiting, and what it costs

The last router does not fall back to the primary. It waits for a replica to catch up, and gives up only after a timeout (6 trials):

wait budgetreads on the primaryreads that waitedtimeoutsread meanread p99
0 ms11.49%0.00%04.03 ms4 ms
100 ms7.58%10.27%1310612.88 ms104 ms
400 ms2.09%8.14%361722.48 ms404 ms
800 ms0.09%7.34%15723.98 ms544 ms
1600 ms0.00%7.31%023.95 ms534 ms

It works, in the sense it claims: correctness stays exactly zero at every setting and the primary is eventually left entirely alone. What it does is move the bill from the primary’s capacity to the reader’s tail, and the shape is the dangerous part. Between 0 ms and 1600 ms the mean read goes 4.0 ms → 24.0 ms, six-fold, while the p99 goes 4 ms → 534 ms, 130-fold. On a graph of mean read latency almost nothing happened. The p99 has become the replica lag distribution’s tail, which is to say your read latency is now set by your slowest replica’s worst checkpoint, which is a thing you do not control and probably do not measure.

What’s on screen

Time runs left to right in a five-second window; the playhead is a frontier, so nothing to its right is drawn.

Notes