2026-07-15
Ask a hardware engineer "how fast is your logic?" and they won't say picoseconds — they'll say FO4s. The fanout-of-4 delay is the propagation delay of a standard inverter driving four identical copies of itself. It's the industry's favorite unit of delay because it normalizes away process, voltage, and temperature: 10 FO4 at 180nm and 10 FO4 at 5nm both mean "ten inverters of work."
Why four? Logical effort theory says the optimal fanout for a chain of inverters (minimizing total delay) is roughly e ≈ 2.7. But four is close enough, easy to lay out on a test structure, and gives a "typical" load — most gates in real designs drive somewhere between 2 and 6 fanouts. Not too lightly loaded (dominated by intrinsic delay), not too heavily loaded (dominated by wire).
The rule of thumb: FO4 delay in picoseconds ≈ 0.5 × (drawn gate length in nm). At 45nm, that's ~22 ps. At 14nm, ~7 ps. At 5nm (effective), ~2 ps. Voltage scaling and finFET geometry break this a bit at bleeding-edge nodes, but it's still the mental math engineers use.
How it shapes design: A pipeline stage's clock period, expressed in FO4s, tells you what kind of animal you're building:
Concrete example: Suppose you're targeting 4 GHz on a 7nm process. Cycle time = 250 ps. FO4 ≈ 3.5 ps. Budget = 250 / 3.5 ≈ 71 FO4 per cycle — but subtract clock skew (~5 FO4), setup time (~2 FO4), flip-flop clock-to-Q (~3 FO4), and jitter margin (~4 FO4), and you're left with about 57 FO4 of pure combinational logic. If your critical path has 12 gates averaging 5 FO4 each, you're already at 60 — over budget. Time to retime, pipeline, or restructure.
The beauty of FO4 is portability. A block that closes at 20 FO4 on 28nm will close at 20 FO4 on 7nm — the numerator and denominator both shrink together. That's why architects reason in FO4 during early exploration, long before a physical library exists.
