In April 1967, Gene Amdahl stood up at a computer conference in Atlantic City to make an unfashionable argument.
For over a decade, he noted, people had been claiming that single-computer design had reached its limits and that real progress required interconnecting many machines to work cooperatively. Amdahl, then at IBM in Sunnyvale, thought this was wrong. His paper was three pages long and titled "Validity of the Single Processor Approach to Achieving Large Scale Computing Capabilities".
He lost the argument. Parallel computing is how essentially everything now works.
But the reasoning he used against it turned out to be correct, and it is now the standard tool for planning the thing he was arguing against — which is a rare fate for an idea and a useful clue about what it is actually good for.
The argument
Take any piece of work and split it in two: the fraction that can be done in parallel, and the fraction that cannot. Preparation, coordination, the step that must wait for the previous step, the decision only one person can make.
Now add processors. The parallel fraction gets faster. The sequential fraction does not — by definition.
Amdahl’s conclusion: for work with a sequential fraction s, the maximum achievable speedup is bounded by 1/s, no matter how many processors you throw at it.
The numbers are unforgiving. If 10% of the work is inherently sequential, infinite parallelism buys you a tenfold speedup and not one unit more. At 20% sequential, the ceiling is five. At 5%, twenty. The ceiling does not depend on how many processors you have. It depends only on the part you cannot parallelize.
And Amdahl’s original point — the one that made this an argument against multiprocessors rather than a guide to them — was that in the problems he cared about, the sequential fraction was large.
The uncomfortable shape of the curve
What makes this operationally interesting is not the ceiling. It is how fast you approach it.
With 10% sequential work, two processors give you about 1.8x. Four give about 3.1x. Eight give about 4.7x. Sixteen give about 6.4x. A hundred give about 9.2x, against a theoretical maximum of 10.
So the first doubling is nearly free and the ninetieth is nearly worthless. Every increment costs the same and returns less than the one before, and there is no point at which this stops being true. There is only the point at which someone notices.
This is why capacity decisions feel fine right up until they do not. Nobody makes an obviously bad call. Each addition is defensible on the previous one’s evidence.
It is not really about processors
The argument holds for anything divisible into parallel and sequential portions, which is why it keeps reappearing.
Add engineers to a program and the parallel work — independent components, separate services — goes faster. The sequential work does not. Architectural decisions, integration, the single approval, the one person who understands the legacy interface: none of that speeds up. Amdahl gives you the ceiling before you hire.
Add GPUs to a training or inference workload and the same structure applies. Data preparation, loading, checkpointing, synchronization between nodes and the human decisions between runs form a sequential fraction that additional accelerators do not touch. A cluster running at a fraction of its theoretical throughput is usually not misconfigured. It is obeying Amdahl.
Add parallel workstreams to a transformation program and the pattern repeats, with the sequential fraction consisting of steering committees, dependencies on a single data migration, and decisions only one executive can make.
In each case the useful move is the same, and it is counterintuitive: stop measuring the resource and start measuring the sequential fraction. That number, not the headcount or the node count, sets what is achievable.
The honest caveats
Amdahl assumed a fixed problem size, which is a real limitation and the basis of most subsequent argument with him. If the problem grows as resources grow — more data, larger models, bigger simulations — the arithmetic changes, and that is broadly why massive parallelism succeeded despite his objection.
He also assumed the parallel fraction parallelizes perfectly with no coordination overhead. Real systems have coordination cost, which means real speedups fall short of his ceiling rather than exceeding it. The bound is optimistic.
And the sequential fraction is not fixed by nature. It is partly a design choice. Some of what looks inherently sequential is only sequential because of how the work was organized — which is the productive question, and the one the arithmetic forces you to ask.
The fraction is the whole argument
Most scaling decisions are argued in terms of the resource: more engineers, more GPUs, more parallel workstreams. Amdahl says the resource is the wrong subject. The binding constraint is the fraction of work that cannot be divided, and that fraction is usually invisible on the plan because nobody itemizes waiting.
The practical test: before adding capacity, estimate what proportion of the elapsed time is genuinely sequential. If it is a fifth, you are buying a fivefold ceiling, and if the plan promises more than that the plan is wrong regardless of how much you spend.
Then ask the better question. Which parts of that sequential fraction are sequential because of physics, and which are sequential because of a decision somebody made about how to organize the work? Reducing s by a few points does more than doubling the resource ever will — which is the whole argument, and the reason a paper written to defend the single processor is still worth reading fifty-nine years after it lost.
Source: Gene M. Amdahl, "Validity of the Single Processor Approach to Achieving Large Scale Computing Capabilities", Proceedings of the AFIPS Spring Joint Computer Conference, vol. 30, Atlantic City, 18-20 April 1967, pp. 483-485. The speedup figures above are calculated from the bound stated in that paper.
Related Reading
- GPU Cloud & AI Compute
- Kubernetes Platforms
- AI Model Serving & Inference Platforms
- Cloud Infrastructure & IaaS
- Technical debt was never about bad code. That is the whole point of the metaphor.
- Knight Capital lost $460 million in 45 minutes because of code it stopped using in 2003
- Site Reliability Engineering: Implementing SRE Principles in the Enterprise
- Cloud Migration Playbook: From Workload Assessment to Production
Independent. No sponsorships. Unsubscribe anytime.