Derating

On capacitors, queues, and the strange conviction that people are the one component in the universe rated for continuous operation at maximum
There is a small, boring, deeply unglamorous act of engineering humility that happens millions of times a day inside the objects you depend on. Somewhere in the phone charger on your desk sits a capacitor stamped 50V. The circuit it lives in never exceeds 30V. Nobody made a mistake. Nobody over-ordered. A grown adult with a degree deliberately paid for 50 volts of capability and then chose, on purpose, to use 60 percent of it.
This is called derating, and it is the difference between things that work and things that work in the lab.
The logic is not complicated. A rating is not a target. A rating is the edge of the map, the place where the manufacturer stops promising and starts shrugging. It is a failure threshold. Operate a component at its rated maximum and you have not achieved efficiency, you have achieved a coin flip with extra steps.
So the engineer backs off. Electrolytic capacitors get run at 50 to 60 percent of rated voltage, partly for margin against transients and partly because the leakage and the ripple current heat the thing from the inside and every 10 degrees Celsius of extra core temperature roughly halves its life. Resistors get half their rated power. Semiconductors get a junction temperature ceiling well below the datasheet number. Wires in a bundled harness get derated for current because their neighbors are also warm, and warmth is contagious in a way that will become uncomfortably relevant in about four paragraphs.
Nobody argues about this. No procurement manager has ever demanded to know why we are wasting 20 volts. No one runs a quarterly review asking whether the capacitors could be pushed a little harder this cycle, given the ambitious roadmap.
And then those same organizations staff human beings at 100 percent utilization, publish it in a spreadsheet as a virtue, and express genuine, wide-eyed astonishment when the whole apparatus turns out to be brittle.
The one equation you actually need
Here is the awkward part for anyone who wants to treat this as a soft, cultural, wellbeing-adjacent conversation. It is not one. It is arithmetic, and it is the same shape as the arithmetic in the electronics.
Take the simplest possible model of a person or a team receiving work - a single server queue, arrivals random, service times random. This is the M/M/1 queue, the hydrogen atom of queueing theory. Define utilization as
rho = arrival rate divided by service rate
Then the average time a unit of work spends in the system, from arrival to done, is
Flow time = service time divided by (1 minus rho)
Stare at that denominator. That is the whole essay. As rho approaches 1, the denominator approaches zero, and flow time approaches infinity. Not "gets worse." Not "degrades gracefully." Goes asymptotic. The mathematics does not contain a polite region near the top.
The multiplier, tabulated, so you can put it on a slide and watch the room go quiet.
| Utilization | Flow time multiplier - 1 divided by (1 minus rho) | Average items waiting | Cost of one more percentage point |
|---|---|---|---|
| 50 percent | 2.0x | 1.0 | tiny |
| 60 percent | 2.5x | 1.5 | small |
| 70 percent | 3.3x | 2.3 | noticeable |
| 80 percent | 5.0x | 4.0 | expensive |
| 85 percent | 6.7x | 5.7 | painful |
| 90 percent | 10.0x | 9.0 | brutal |
| 95 percent | 20.0x | 19.0 | absurd |
| 98 percent | 50.0x | 49.0 | fictional |
Two readings of this table matter.
The first is the obvious one. A two-day task takes about seven days at 70 percent utilization and about twenty days at 90 percent. Same task. Same people. Same skill. The only thing that changed is how full the pipe was, and the pipe fullness is the thing your planning process optimizes for.
The second reading is the one that wins arguments. The marginal cost of utilization is the derivative of that multiplier, which is 1 divided by (1 minus rho) squared. At 70 percent, one extra percentage point of loading costs you about 0.11 service times of delay. At 95 percent, one extra percentage point costs about 4.0. That is roughly a 36-fold difference in price for the identical purchase. You are buying the same percentage point of "efficiency" at thirty-six times the cost and calling it discipline.
And now the bad news about knowledge work
The table above is optimistic, which is a fun sentence to write.
M/M/1 assumes exponentially distributed service times. Real knowledge work is worse behaved than that. Kingman's approximation generalizes the result -
Queue wait is approximately (rho divided by (1 minus rho)) times ((arrival variability squared plus service variability squared) divided by 2) times mean service time
The middle term is the variability penalty, and it equals 1 for the well-mannered M/M/1 case. Knowledge work is not well mannered. A support ticket takes twenty minutes or eleven days. A "small refactor" is either an afternoon or a geological epoch. A squared coefficient of variation of 3 or 4 on service time is entirely ordinary, which multiplies your queueing delay by roughly 2 to 2.5 on top of everything in the table.
So a team at 90 percent utilization doing genuinely variable work is not sitting at 10x. It is somewhere north of 20x, with a distribution whose tail is long enough to swallow a quarter.
This is the exact analogue of ripple current in the capacitor. The average voltage was fine. The variation is what cooked it.
One consolation - pooling helps. Multiple interchangeable servers sharing one queue tolerate higher utilization than a single server, because someone is usually free. Which is a precise mathematical argument that a team of eight generalists sharing work can safely run hotter than eight specialists each guarding a private inbox. Every act of specialization, every "only Priya knows the billing service," converts your forgiving multi-server queue back into a set of single-server queues and drags you back to the harsh column of the table. Specialists need more derating, not less. Bundled conductors, heating each other, all over again.
The utilization number in your spreadsheet is already a lie
Before anyone defends 100 percent, it is worth noticing that 100 percent is almost never what is being measured.
Count a year honestly. Start with 261 weekdays. Remove around 11 public holidays, 20 days of vacation, 5 days of sick leave, and perhaps 10 days of training, hiring loops, all-hands, and organizational weather. You are at roughly 215 days, which is 82 percent of nominal before a single line of the plan has been written. Then remove the standing meetings, the code review, the incident that was not your incident, the onboarding of the new hire, the interview panel, the context switching tax that is invisible precisely because everyone pays it. Twenty percent is a generous estimate and a conservative one at the same time.
So when a plan is loaded to 100 percent of nominal capacity, the actual arrival rate against actual service capacity is somewhere around 130 to 150 percent. And when rho exceeds 1, queueing theory stops offering a flow time at all. The queue simply grows without bound. Which, if you have ever looked at a backlog that only ever gets longer no matter how hard everyone works, is not a metaphor. It is a diagnosis.
Here is that arithmetic in the form most likely to survive contact with a planning meeting. Team of six, sprint of ten working days.
- Naive capacity - 6 times 10, which is 60 person-days, a number with the seductive cleanliness of a lie
- Minus vacation and one sick day - say 5 days gone, leaving 55
- Minus ceremonies, reviews, and standing meetings at 15 percent - leaving about 47
- Minus the support and on-call rotation, one person half-time - leaving about 42
- Minus interrupt work and unplanned defects, historically 15 percent of the remainder - leaving about 36
The honest capacity is 36 person-days. Planning 60 is not ambition, it is planning at 167 percent of a system whose behavior becomes undefined past 100. Planning 36 is not sandbagging. Planning around 25 to 28, which is 70 to 78 percent of the honest number, is spec compliance.
The objections, and the numbers that answer them
"We cannot afford to pay people to be idle."
You are not paying for idle, you are paying for margin, and you already do this everywhere else without complaint. The fire extinguisher in the corridor has a utilization of zero percent and you have never once considered selling it. More usefully - do the cost arithmetic. Six engineers at a loaded cost of 150 thousand is 900 thousand per year, so the 30 percent you are calling waste is about 270 thousand. Now price the delay. If a single initiative is worth 40 thousand per week of margin or advantage, and derating cuts its lead time from nine weeks to three, that one project returns 240 thousand. Two such projects a year and the margin has paid for itself with change left over for the offsite nobody wants.
"Other teams manage at 95 percent."
They manage the number. They do not manage the queue. Ask for lead time distributions rather than utilization, and specifically ask for the 85th percentile. Utilization is an input that is easy to measure and flattering to report. Lead time is an output that is hard to fake and tells you what the customer actually experiences. Optimizing the input while the output rots is the management equivalent of tuning a radio by staring at the volume knob.
"If we build in slack it will just get absorbed."
Correct, and this is the strongest objection in the list, so it deserves a real answer rather than a joke. Slack that is undefended does get eaten, by scope creep, by Parkinson's law, by the well-meaning stakeholder who noticed a gap in a calendar. The answer is not to abandon derating but to give the margin a job. In electronics the derated headroom is not doing nothing, it is absorbing transients. In a team, the headroom absorbs incidents, and when there is no incident it does the work that compounds - the test coverage, the documentation, the mentoring, the tooling, the thing that everyone agrees is important and nobody has ever been given a day for. Explicitly named improvement capacity is defensible. Unnamed emptiness is not.
"But this feels like an excuse for a slower team."
The derated team is faster. That is the entire and only point. Lower utilization, shorter lead times, more work delivered per quarter, because throughput was never the bottleneck - waiting was. Little's Law puts it in five characters. Work in progress equals throughput times lead time. If you want shorter lead times and you cannot magically increase throughput, there is precisely one lever left, and it is the one everyone is emotionally attached to pulling in the wrong direction.
Where high utilization is actually fine
Intellectual honesty requires admitting that the 70 percent heuristic is a heuristic, not a physical constant, and that it comes from the interaction of variability with the cost of delay. Two conditions genuinely relax it.
Low variability - if arrivals are smooth and service times are nearly identical, the Kingman penalty collapses and you can run hot. This is why a well-tuned assembly line hits utilization that would destroy a product team. It is also why "we should operate like a factory" is usually said by people who have never met either a factory or their own defect rate.
No cost of delay - if the queue is genuinely asynchronous and nobody cares when an item emerges, then long waits are free and you should run near capacity. Batch data processing overnight. Fine. Almost nothing involving a customer, a competitor, or a colleague waiting on an answer qualifies.
Neither condition describes your roadmap. Both conditions are frequently claimed to describe your roadmap.
The vent
The thing about an abused electrolytic capacitor is that it does not fail politely. Internal pressure builds, the scored top opens, the electrolyte leaves, and you get the bulging can that anyone who has ever repaired a cheap motherboard recognizes on sight. It worked, and worked, and worked, right up until the moment it did not, and afterwards the failure looks obvious and cheap to have prevented.
The human parallel writes itself, and I am not going to be coy about it because the parallel is not decorative, it is structural. Sustained operation above rated conditions accelerates degradation in a way that is invisible on the dashboard and catastrophic in the field. The people who leave are rarely the ones who complained. They are the ones who absorbed variance for three years while the plan assumed they were a component with no thermal limits.
What the datasheet would say
If people came with datasheets, and mercifully they do not, the recommended operating conditions would say something like - continuous duty at 70 percent of rated capacity, with transient excursions to 90 percent tolerated for periods measured in weeks rather than quarters, adequate cooling required between excursions, life expectancy halves for every sustained increment above nominal.
We would follow it, because we follow it for the 30 cent part.
So the next time someone asks why the plan only fills 70 percent of the sprint, do not reach for the language of wellbeing, or balance, or sustainability, all of which are true and none of which survive a budget meeting. Reach for the table. Point at the denominator. Explain that you are not being generous, you are being compliant, and that the alternative is a design that any competent engineer would reject on sight if only it were made of aluminum and electrolyte rather than people with mortgages.
The capacitor gets its 20 volts of dignity. Ask for yours.