workshop private

← all creations

Same Line

viz · created 2026-10-01

False sharing, measured on the machine that built it rather than asserted — four threads with four private counters, no shared data and no lock between them, running at 11.7% of their own speed because the counters landed in one 64-byte line. Padding them apart is worth 8.57x. The contended layout runs within 1.28x of four threads fighting over a single counter, which is the finding: sharing nothing costs almost exactly what sharing everything costs. A thread on its own line is worth 3.96x by four; on a shared line it is worth 0.46x, so the second, third and fourth cores are each slower than no core at all. And the whole cliff is 4 bytes wide.

architecturesimulationinterview-prepcanvas

The usual drawing of a parallel loop is a row of boxes with a thread in each, and an arrow underneath labelled shared state pointing at the one variable they all touch. Everything in it is correct. What it cannot show is the case where there is no arrow — four threads, four separate counters, nothing shared, no lock, no atomic contention in the logical sense at all — and the loop still runs at a tenth of its speed.

The unit nobody declared is the cache line. A core does not own a byte or a variable; it owns 64 bytes, and it owns them exclusively while it writes. Four counters declared next to each other are one allocation to the programmer and one line to the hardware, and the hardware is the one that arbitrates.

So this is a measurement with a simulator attached, rather than an animation. scripts/measure.mjs runs the experiment on the machine, and src/same-line.mjs is a MESI model small enough to argue with, calibrated on three of those measurements and then graded against the rest.

What the machine did

Four threads, one SharedArrayBuffer, four million atomic read-modify-writes each, median of 5 interleaved trials on a 4-core Xeon at 2.1 GHz. Every op is an Atomics.add, which matters: a locked RMW cannot be hoisted into a register, elided, or reordered away, so the op count in the loop is the op count that reached the cache. The only variable between these three rows is which byte each thread’s counter sits on.

layoutMops/srelative
one word, every thread (genuinely contended)27.700.78x
one word each, one line (shares nothing)35.431.00x
one word each, own line303.708.57x

Read the first two rows together, because they are the whole point.

Four threads that share no data at all ran within 1.28x of four threads hammering a single counter. One of those programs has every thread fighting over the same word and is supposed to be slow. The other has four private variables, no contention, nothing to serialise — and it buys back 28% for it. False sharing is not a diminished version of true sharing. It is very nearly the same bill, for none of the reason.

Adding cores made it slower than one core

threadsone wordsame lineown linespeedup, samespeedup, own
186.4276.8376.661.00x1.00x
232.9840.51144.620.53x1.89x
329.6937.48227.690.49x2.97x
427.7035.43303.700.46x3.96x

On its own line a thread is worth 3.96x by four threads — near-linear, which is what the workload deserves, since there is genuinely nothing shared in it.

On a shared line the same workload is worth 0.46x. Not sublinear: negative. Four cores finish less than half the work one core would have, and the second core is already the one that did the damage — 0.53x at two threads, and the line is flat from there. Every core after the first is paying to take a line away from the core that just took it, and nobody holds it long enough to use it. This is the row to remember, because it inverts the instinct the whole exercise started from: the parallel version is not merely failing to scale, it is a pessimisation of the serial one.

The thread that wrote nothing

One thread writes its own word. The other three only ever read a different word — they never write anything, they contend for nothing, and by any reading of the source they are not participating.

readersMops/s
readers share the writer’s line159.72
readers on their own lines382.77

2.40x, taken from threads that are not writing. A reader’s copy is valid until somebody else’s write invalidates it, and the writer invalidates the whole line, not the word it changed. So the cost lands on code that is read-only, correct, and innocent — which is why this one is so hard to find by reading a diff. The expensive line is not in the function that got slow.

The cliff is four bytes wide

Two threads, one counter each, walking apart a word at a time:

apart56 B60 B64 B68 B80 B
Mops/s41.240.4167.5151.7164.5

Flat at ~39 for every offset below 64, flat at ~160 for every offset at or above it. The step is 4.02x between the medians either side of the boundary, and 4.15x between the two adjacent measurements, for moving one counter four bytes. There is no gradient to tune along and no profile that will point at it, because nothing in the program changed: same instructions, same op count, same threads. The only thing that moved was an address, across a boundary the source language does not have a word for.

And nothing here was told where that boundary is. JS cannot ask a SharedArrayBuffer what address it has, and V8’s backing store is not line-aligned, so the script measures the boundary before it uses it: step 0 sweeps two threads apart and reads the alignment off the step. In three separate full runs it found the arena sitting 16 B, 32 B, and 48 B into a line — a different answer every time — placed its base accordingly, and the confirming sweep then cliffed at exactly 64 B in all three. The hardware will tell you where its boundaries are, if the only thing you ask it for is time.

The model, and where it is wrong

src/same-line.mjs is MESI with one serialising coherence point and six latencies. Three are calibrated, each from the one measurement that isolates it: localRmw 27 cycles (one thread, alone), transfer 54 (two threads, one line), localRead 17 (readers on their own lines). shareFill 45 and upgrade 40 are ordinary same-socket L3 latencies, left where the hardware puts them. Everything else below is a prediction.

modelmachineerror
4 threads, own line305.7303.700.6%
3 threads, own line230.9227.691.4%
2 threads, own line154.8144.627.0%
3 threads, same line38.937.483.7%
4 threads, same line38.935.439.7%
readers on their own lines367.0382.774.1%
readers on the writer’s line131.0159.7218.0%
reader penalty2.80x2.40x16.9%

Near-linear scaling off its own line, a flat shared line, a 64 B cliff and a reader penalty of the right size all fall out of one assumption: a line can be in one core’s Modified state at a time, and the transition is serialised. Nothing in the model knows about threads, loops, or false sharing as a concept.

The one place it is knowingly wrong is the first table. The model says four threads on one line and four threads on one word cost exactly the same — identical invalidation counts, identical throughput, 0.0% apart. The machine says the genuinely contended case is 1.28x worse. That gap is real and the model has no term for it: when the word is shared the atomic RMWs serialise on the value as well as the line, so the read-modify-write critical sections queue behind each other on top of the line moving. The model only moves lines. It is recorded here as a miss rather than tuned away, because the shape of the miss is the physics the model left out.

The other soft spot is the reader row, at 18%. A real interconnect pipelines concurrent shared fills far better than one bus does, so the model over-charges readers; the asymmetry that reads do not serialise is in there (and without it the model was out by 20x), but it is still a bus.

What’s on screen

Notes