Slack
The Spare Capacity That Looks Like Waste Until the Day It Does Not
Introduction
Every organisation contains things that are not being used. An empty hospital bed. A warehouse holding more parts than this month needs. An engineer whose calendar has gaps in it. A second supplier who currently supplies nothing. A shipping route nobody sails because the shorter one is open. Cash sitting in an account earning less than it could earn somewhere else.
Each of these is easy to argue against. You can put a number on what the idle thing costs. You can put no number at all on what it prevents, because what it prevents has not happened. So these things get removed over time, one defensible decision at a time, by people who are being sensible. Then something goes wrong, and the thing that would have absorbed it is not there.
The name for the unused capacity is slack. It is one of the few concepts that explains failures across completely unrelated systems: a hospital in winter, a car factory without chips, a shipping lane with no usable alternative, a person with no savings and a broken car. The page below walks through what slack does, the arithmetic that makes it necessary, why it is removed anyway, and the honest case on the other side. It is not an argument that more slack is always better. It is an argument that the trade-off is real, that most people misjudge which way it runs, and that the misjudgement has a specific mathematical shape.
The Arithmetic Almost Nobody Runs
Start with the least intuitive part, because everything else follows from it. When a system with variable demand runs closer to its capacity, the delays inside it do not grow in proportion. They grow far faster than that, and the growth accelerates as you approach the limit.
The simplest version of this comes from queueing theory, the branch of mathematics that describes things waiting in line. For a basic queue with random arrivals, the average number of items in the system - waiting plus being served - works out as the utilisation rate divided by one minus the utilisation rate. That formula is worth converting into ordinary numbers, because the numbers are startling. At 50% utilisation, on average one item is in the system. At 80%, four. At 90%, nine. At 95%, nineteen. At 99%, ninety-nine.
Read that sequence again with an eye on the gaps. Going from 50% to 80% utilisation - a large gain in efficiency by any manager's reckoning - triples the queue. Going from 90% to 95%, a much smaller gain, doubles it. Going from 95% to 99% multiplies it by five. The last few percentage points of capacity, the ones that look like the easiest remaining savings, are the ones that cost the most.
This is not a quirk of one formula. It is a general property of systems where demand varies and capacity does not. Human intuition about this is close to linear. Asked what happens when a system at 90% capacity is pushed five percentage points higher, most people expect a little more trouble. The real answer is that the queue roughly doubles. The intuition is wrong in the same direction every time, which is why the mistake is systematic rather than random.
The clearest empirical version of this comes from hospitals. Alan Bagust, Michael Place and John Posnett published a simulation in the British Medical Journal in 1999 that has been cited in bed-planning ever since. Their finding was that the risk of running out of beds becomes visible once average occupancy passes roughly 85%, and that a hospital averaging 90% or more should expect regular shortages and periodic crises. The number is not magic and it varies by setting. What matters is the shape: there is a level of fullness beyond which a system stops absorbing variation and starts transmitting it.
Notice what that implies about a hospital running at 85% occupancy. Fifteen per cent of its beds are empty on an average day. An accountant looking at that line sees fifteen per cent waste. The accountant is not being stupid: the empty beds genuinely cost money and genuinely treat nobody. But they are not idle in the sense of doing nothing. They are doing the only job that empty beds can do, which is to exist on the day when more people arrive than usual.
Efficiency and Resilience Are One Lever
Organisations usually talk about efficiency and resilience as two separate goals, to be pursued in parallel by different initiatives. The framing is comfortable and mostly wrong. For a fixed level of capability, they are opposite ends of a single lever, and the lever is how much of your capacity is committed.
Efficiency, operationally, means having less of everything than the worst case requires. Fewer beds than the worst flu week. Less inventory than the longest supplier delay. Fewer staff than the busiest shift. Less cash than the deepest downturn. Each reduction is a real gain in ordinary conditions and a real loss in unusual ones, and it is the same reduction. You cannot remove the cost of holding a buffer while keeping the protection the buffer provides, because the protection is the holding.
This is why "we need to be both more efficient and more resilient" is a sentence that sounds like strategy and functions as an instruction to do nothing. It can be made true, but only by changing capability rather than committing more of it - better forecasting, faster reconfiguration, substitutable parts, cross-trained staff, suppliers who can be switched on quickly. Those are genuine ways to buy resilience without holding as much idle stock. They are also expensive, slow to build, and much harder than the thing organisations usually do instead, which is to declare both goals and quietly optimise for the one that shows up in this quarter's numbers.
Why the Buffer Goes Anyway
If slack is so useful, its steady disappearance needs explaining. The explanation is not that decision-makers are foolish. It is that the incentives around slack are asymmetric in every direction at once.
The cost is measured and the benefit is not. What the spare capacity costs appears in an account every month. What it prevents appears nowhere, because prevented events leave no record. Anyone proposing to cut it arrives with a number. Anyone defending it arrives with a hypothetical. In most organisations that argument has one likely ending.
Success is indistinguishable from over-provision. A buffer that is working looks exactly like a buffer that was never needed. Ten quiet years are evidence for cutting, right up until the eleventh. This is the same structure that makes infrastructure maintenance chronically underfunded, and it produces the same result for the same reason.
The ratchet only turns one way. Cutting a buffer is a decision someone makes and gets credit for. Restoring one is a decision nobody makes, because the case for it is the same hypothetical that lost the argument last time, now weakened by the years of quiet that followed the cut. So the level drifts downward in steps and returns upward only after a failure, briefly, and usually not all the way.
Competition removes it whether or not anyone decides to. This is what makes slack a hard problem rather than a management failing. A firm that holds spare capacity carries a cost its competitors do not. In a market with thin margins it is undercut by the firm that holds none - not eventually, but every quarter, until it either matches them or loses. The prudent firm is punished for prudence in all the years when nothing happens, which is most years. No single firm can hold the line alone. That makes this a collective-action problem rather than a decision.
That last point matters for what can be done about it. If slack disappears because managers are short-sighted, the fix is better managers. If it disappears because competition removes it from anyone who holds it, better managers change nothing, and the only mechanisms that work are the ones that apply to everyone at once: regulation, mandated reserves, public provision, or an industry agreement that survives the temptation to defect. Which of those is appropriate depends on how bad the failure is when it comes, and reasonable people put that line in different places.
The Same Shape in Different Materials
Slack is easier to recognise once you notice it takes different physical forms while doing the same job. Each form is held for a different reason and cut by a different argument, but the function is identical: absorbing variation so that it does not propagate.
Inventory is slack in the form of stuff. It absorbs the gap between when supply arrives and when demand appears. Spare capacity is slack in the form of machines or beds or seats, absorbing peaks in demand. Staffing above the average requirement absorbs illness, turnover, and the days when everything happens at once. Cash is slack in the most liquid form, absorbing anything that can be solved with money. Time in a schedule absorbs the tasks that run long, which is most of them.
Two forms deserve separate mention because they are the least visible. Redundant paths - a second supplier, an alternative route, a backup system - cost almost nothing to hold and everything to lack. That is why they survive longer than other forms, and why they so often turn out to have quietly atrophied when finally tested. Attention is slack in the form of unallocated human capacity. An organisation where everyone is fully committed has nobody free to look at the new problem, which is why fully-loaded teams handle surprises so badly.
The forms are partly substitutable, which is where real design decisions live. Cash can replace inventory if suppliers can deliver fast enough. Flexibility can replace spare capacity if the machines can be switched between products. Good forecasting can replace some of everything. This substitution is genuinely valuable and it is the honest answer to the accountant, but it has a limit that is often forgotten: every substitution depends on something else continuing to work. Cash only replaces inventory while the supplier exists and the ship sails.
What Happened When It Was Tested
The argument above is theoretical until systems are pushed, and several were pushed hard in a short period. The cases are worth reading together, because the same shape appears in industries that share nothing else.
Hospitals in the pandemic. Health systems in most developed countries had spent decades reducing bed numbers, for reasons that were partly good - shorter stays, better day surgery, care moved out of hospitals. The reduction went further than the clinical improvements alone justified, because occupancy is a visible number and empty beds are a visible cost. Systems entered 2020 running at or above the level where variation stops being absorbed, and had almost nothing left to give when a large correlated shock arrived.
Car makers and chips. When demand collapsed in early 2020, automakers cancelled semiconductor orders, exactly as just-in-time practice instructs. Chip makers reallocated that capacity to consumer electronics, where demand was rising. When car demand recovered, the automakers went back and found the queue was long and they were at the end of it. Plants stood idle for want of parts costing a few dollars. The decision to cancel was locally correct in every respect except one: it assumed the ability to return, and that ability was not theirs to grant.
Chokepoints at sea. A single grounded container ship closed the Suez Canal for six days in 2021 and the effects were felt for months, because there was no slack anywhere in the chain to absorb the delay. The larger test came in 2026, when Iran closed the Strait of Hormuz to normal commercial traffic. Cargo rerouted around the Cape of Good Hope at a cost of roughly ten to fourteen extra days per voyage, and the alternative pipelines carried part but not all of the volume. The redundant routes existed. They had simply not been maintained at the scale that would have made them a real substitute.
The common feature is worth stating plainly: in none of these cases did anyone decide to remove the safety margin. There was no meeting, no memo, no moment at which the risk was weighed and accepted. Each step improved a number that someone was accountable for, and the accumulation was nobody's job to watch.
The Honest Case Against Slack
Everything above can be turned into an argument for holding buffers everywhere, and that argument is wrong. Slack has real costs, and the people who cut it are usually responding to something true.
The cost is not notional. Capital held as inventory is capital not doing anything else. A hospital bed that stays empty was still built, staffed and heated. Money spent on a buffer is money not spent on the thing the organisation exists to do, and in a health system that trade-off is measured in treatments not given. "Hold more slack" is never free and is sometimes paid for by the same people it is meant to protect.
Slack hides problems as well as absorbing them. This was the original insight behind just-in-time manufacturing and it has not stopped being true. Inventory conceals unreliable suppliers, because you never notice the late delivery. Spare staff conceal a broken process, because someone always picks up the slack. Taiichi Ohno's argument at Toyota was that lowering the water level reveals the rocks, and that revealing them is how they get removed. A system with generous buffers everywhere can be comfortably mediocre for a very long time.
Buffers in the wrong place do nothing. Slack only helps where the constraint actually binds. Inventory in front of a machine that is never the bottleneck is pure cost, and the intuition that "more is safer" produces exactly this outcome when applied without looking. Eliyahu Goldratt's work on constraints makes the case that most buffers in most organisations are protecting steps that need no protection, while the one step that does need it has none.
And some slack is not slack at all. Unused capacity that could not be brought to bear in a crisis is just cost wearing a useful name. Reserve equipment that has not been tested, backup suppliers with no live contract, standby staff without current training - these appear in the plan and fail on the day. The distinction between a buffer and a decoration is whether anyone has recently checked.
Taken together, these are not a refutation. They are the reason the sensible question is never "more or less" but "how much, where, and who pays." A system with no slack fails on its first bad day. A system with slack everywhere fails slowly, by being outcompeted, and never learns what is wrong with it.
How Much Is Enough
There is no general answer, but there are questions that narrow it, and they are more useful than a target number would be.
How variable is the demand? Buffers exist to absorb variation, so the size of the variation sets the size of the buffer. A system with steady, predictable load needs very little. A system where a bad week is three times an average week needs a great deal, and no amount of forecasting will change that if the variation is genuinely random rather than merely unmeasured.
How long does it take to add capacity? This is the question most often skipped. If more can be brought online in hours, little needs to be held. If it takes a decade - power stations, hospitals, trained specialists, semiconductor fabrication plants - then the buffer has to cover the entire lead time, because during that period the only capacity available is the capacity that already exists.
How correlated are the failures? A buffer sized for independent problems is useless against a shock that hits everything at once. Two suppliers in the same industrial park are one supplier. Backup generators below the flood line are no backup. This is the failure mode that most often defeats plans that looked adequate: the redundancy was real but it shared a hidden dependency with the thing it was meant to replace.
And what does failure actually cost? A retailer that runs out of a product loses a sale. A hospital that runs out of beds loses something else. Where the downside is recoverable, running lean is reasonable and the arithmetic favours it. Where the downside includes outcomes that cannot be undone, the calculation changes completely, and it changes in a way that expected-value reasoning handles badly - which is the same asymmetry that makes ruin a special case rather than just a large loss.
Working Practically With Slack
The same reasoning applies to a life, and the personal version is easier to act on than the organisational one because there is no committee. An emergency fund is slack in its purest form: money doing nothing. That is the entire point, and it is also why the fund is the first thing people convert into something that looks more productive. Unscheduled time is the same idea applied to a calendar. A week booked to capacity has no room for the thing that goes wrong, and something goes wrong most weeks.
The utilisation arithmetic applies to people as directly as it applies to hospitals. Someone running at 95% committed does not have 5% left; they have a queue that doubles when anything unexpected arrives. This is why capable people sometimes appear to collapse suddenly rather than gradually. Nothing changed in their capability. The load crossed the point where the curve turns.
Three habits follow, and they are worth more than any general resolution to be prepared. Notice which of your buffers has quietly been spent; the spending is gradual and nobody announces it. Check that the backups are real. A skill unused for five years and a contact unspoken to for three are decorations, not options. And when you are asked to justify the unused thing, remember that you are arguing against a specific number with a hypothetical. Losing that argument repeatedly is how a normal bad week turns into a crisis.
Slack is the part of a system that is doing nothing, and the part that determines what happens on the worst day. The arithmetic that governs it is not intuitive - the last few percentage points of capacity are the expensive ones, not the cheap ones - and the incentives around it are asymmetric at every level, from the manager with a budget to the firm facing a competitor who holds none. None of that argues for buffers everywhere. It argues for knowing where yours are, what they are for, and how much of them you have already spent without deciding to.


