Topicscpu
CPUs and the End of Free Speed: Why Processors Stopped Getting Faster
A sourced starter collection on the central processor — the two scaling laws that governed it, why one of them broke in 2005 and clock speeds stalled, what multiple cores bought and what Amdahl's law capped, and how the CPU became a component of somebody else's rack.
- Assertions
- 18
- Sources consulted
- 25
- Read in full
- 11/25
- Cited as evidence
- 11
- Disputed
- 2
25 sources sit behind this page — including any that arrive with a concept this page shares with another collection. 11 were retrieved and read in full, and only those can back an assertion. 1 could not be retrieved, and 13 were surfaced and deliberately set aside. Every one of them is named in the register below, with the reason in view. How we source this.
Timeline newest first · evenly spaced, not to scale
Our own synthesis, written to orient you — not evidence. Every factual statement here is asserted and sourced further down this page.
For about thirty years processors got faster for free, and then they stopped. Understanding why requires separating two rules that are constantly confused. Moore's law is an observation about counting: Gordon Moore noted in 1965 that the number of components per integrated circuit had been doubling every year, and revised it in 1975 to roughly every two years. It says nothing about speed. Dennard scaling, from a 1974 paper co-authored by Robert H. Dennard, is the one that promised speed: as transistors shrink, their power density stays constant, so power use stays in proportion to area. Smaller meant faster and cooler at the same time.
Dennard scaling broke around 2005. Leakage current and threshold voltage do not scale with size, so power density rises as features shrink, and the industry hit what it named the power wall. Intel cancelled the Tejas and Jayhawk processors in 2004. Clock frequency has stagnated between 4 and 6 GHz ever since, with per-CPU power settling near 100 watts. Moore's law kept going; the free speed did not. Everything since is an answer to that one failure.
The first answer was more cores. Three walls forced it — the memory wall, the widening gap between processor and memory speed; the ILP wall, the difficulty of finding enough independent work inside one instruction stream; and the power wall. Dual-core parts became commonplace in personal computers in the late 2000s. But Amdahl's law caps what that buys: the gain is bounded by the fraction of the program that can actually run in parallel, and most programs are stubbornly sequential.
Which is the precise reason the accelerator exists. Matrix multiplication, the inner loop of every neural network, is close to embarrassingly parallel — one of the rare workloads for which adding arithmetic units really does add performance. A GPU is a processor that spends its transistor budget on those units instead of on the caches and control logic a CPU needs to make sequential code fast. The whole AI hardware industry sits in the gap Amdahl's law left open.
So the processor's role changed. NVIDIA's Grace CPU Superchip carries 144 Arm Neoverse V2 cores and LPDDR5X memory at up to 1 TB/s, joined to the accelerator by a 900 GB/s NVLink-C2C link — an Arm design in a data centre x86 owned for two decades, and one whose headline specification is the bandwidth of its link rather than its compute. By the Vera Rubin announcement of March 2026 the CPU is simply one of seven coordinated chips, and the rack pairs 72 accelerators with 36 processors. The CPU is no longer the machine. It is a part of one.
Whether Moore's law itself still holds is disputed, and this page leaves it disputed. Intel's chief executive said in 2015 that the cadence had slowed toward two and a half years. In September 2022 NVIDIA's chief executive declared the law dead and Intel's said the opposite days later — two interested parties whose companies' strategies depend on opposite answers.
Most readers know processors through a laptop, and the parts that run data centres differ in kind. A mainstream Ryzen 9000 desktop chip tops out at 16 cores on two memory channels. AMD's Turin server generation reaches 128 cores — or 192 of the denser variety — on twelve channels, and Venice is stated at up to 256 on sixteen, with 128 PCIe lanes so one socket can host several accelerators, a fleet of NVMe drives and a fast network card at once. Intel names the same list for what separates Xeon from Core: ECC memory, more cores, more PCIe lanes, more RAM, larger caches, and reliability features.
ECC is the difference most buyers never see and the one that matters most. Server memory carries extra check bits — 72 bits per word for 64 of data, nine chips on a DIMM side instead of eight — so a single-bit flip is corrected and a double-bit flip is at least noticed. Whether that is worth paying for depends on how often memory goes wrong, and a 2009 Google study measured between 25,000 and 70,000 errors per billion device-hours per megabit: roughly one error per gigabyte every 1.8 hours. Desktop platforms historically went without.
Server processors also reached core counts no single die could hold, and solved it the way accelerators later would. AMD's Rome generation in 2019 rebuilt the processor as eight small 7nm compute chiplets around one cheap 14nm input/output die — leading-edge silicon only where it earns its cost. The difference in motive is worth noting: AMD split the die to make a large processor economic; NVIDIA split Blackwell because the reticle limit left it no choice.
This is a starter collection and shorter than the memory and GPU pages it connects to. What is missing is named rather than implied: Dennard's 1974 paper is behind a subscription wall and is cited here only through a tertiary account, Moore's own 1965 paper was not retrieved, Arm's Neoverse page returned an error shell so the licensor's side of the Grace design is unsourced, and the cache article read here gives no cycle-latency figures, so none are quoted. AMD's and Intel's own product pages remain largely marketing; the generation tables used here come from tertiary accounts rather than from the manufacturers.
Figures
Every number below is asserted and sourced elsewhere on this page.
Cores per socket, when the clock stopped rising
Maximum cores in one AMD EPYC socket by generation. Frequency has been flat since 2005; this is where the performance went instead.
Maximum cores per socket
Naples 2017
Rome 2019
Milan 2021
Genoa 2022
Turin 2024
Venice 2026
Wikipedia's Epyc article, read directly — a tertiary account of specifications AMD publishes only as scattered marketing. The 2026 Venice figure is a stated specification for a generation at the start of its life, not a shipped measurement.
Where server and consumer parts diverge
Memory channels per socket. This is the difference that is hardest to see on a spec comparison and hardest to work around: many cores are useless if you cannot feed them.
Memory channels per socket
Ryzen desktop
Threadripper
EPYC Genoa/Turin
EPYC Venice
Wikipedia's Ryzen and Epyc articles, both read directly. Consumer and Threadripper figures from the Ryzen article; server figures from the Epyc article. Stated platform specifications, not measurements.
Two laws, one of which stopped
Clock frequency stopped rising when Dennard scaling failed. The band since 2005 is a ceiling, not a trend — the figure shows the stated stagnation range rather than any single part.
Clock frequency, GHz (stated stagnation band since 2005)
Band, low
Band, high
Wikipedia's Dennard scaling article, read directly. It states clock frequency has stagnated at 4–6 GHz and CPU power at around 100 W TDP since 2005; the two columns show the low and high ends of that stated range, not measured products.
Concepts
The vocabulary this subject is built from, and what we can show about each.
Dennard Scaling
processDennard scaling, from a 1974 paper co-authored by Robert H. Dennard, held that shrinking transistors keeps power density constant; it broke down around 2005 because leakage current and threshold voltage do not scale, stalling clock frequency at 4–6 GHz and per-CPU power near 100 W.
1 source1 retrieved & read
- SupportsRetrieved & readDennard scaling
DRAM Cell (1T1C)
componentshared from another collection — see its own page for what it assertsECC Memory
memory technologyECC memory adds check bits — 72 bits per word for 64 of data, nine DIMM chips instead of eight — to correct single-bit and detect double-bit errors; a 2009 Google study measured 25,000–70,000 errors per billion device-hours per megabit, about one per gigabyte every 1.8 hours.
1 source1 retrieved & read
- SupportsRetrieved & readECC memory
Graphics Processing Unit (GPU)
componentshared from another collection — see its own page for what it assertsHigh Bandwidth Memory (HBM)
memory technologyshared from another collection — see its own page for what it assertsLeading-Edge Process Node
processshared from another collection — see its own page for what it assertsMoore's Law
otherMoore observed in 1965 that components per integrated circuit had been doubling annually, revising it in 1975 to roughly every two years; the law counts transistors rather than speed, a distinction that only became visible once Dennard scaling failed.
1 source1 retrieved & read
- SupportsRetrieved & readMoore's law
Moore's law's continuation is disputed by interested parties: Intel's chief executive said in 2015 the cadence had slowed toward two and a half years, and in September 2022 NVIDIA's chief executive declared the law dead while Intel's contradicted him within days.
1 source1 retrieved & read
- SupportsChallengesRetrieved & readMoore's law
Multicore and Amdahl's Ceiling
componentThe move to multiple cores answered three walls — memory, instruction-level parallelism and power — with dual-core parts mainstream in personal computers by the late 2000s, but gains are bounded by Amdahl's law and the parallel fraction of the workload.
1 source1 retrieved & read
- SupportsRetrieved & readMulti-core processor
Server CPU Chiplet Scaling
packaging technologyAMD's EPYC line went from a four-die 32-core Naples in 2017 to eight 7nm chiplets around a 14nm I/O die at 64 cores in Rome (2019), 96 on 12-channel DDR5 in Genoa (2022), 128–192 in Turin (2024) and a stated 256 cores on 16 channels in Venice (2026).
2 sources2 retrieved & read
- SupportsPrimary evidenceRetrieved & readAMD EPYC Server CPUs
- SupportsRetrieved & readEpyc
Server Versus Consumer Processors
componentA mainstream consumer Ryzen desktop part reaches 16 cores on two memory channels; AMD's server generations reach 128–192 cores on twelve channels and up to 256 on sixteen, with 128 PCIe lanes — and Intel names ECC, core count, PCIe lanes, RAM capacity, cache size and RAS as what separates Xeon from Core.
The Cache Hierarchy
componentThe L1/L2/L3 cache hierarchy exists to hide the memory wall — a modern CPU can execute hundreds of instructions in the time taken to fetch one cache line from main memory — and spending transistors on caches rather than arithmetic units is the exact inverse of the GPU's bargain.
1 source1 retrieved & read
- SupportsRetrieved & readCPU cache
The Host CPU in an AI System
componentNVIDIA's March 2026 Vera Rubin platform announces the Vera CPU as one of seven coordinated chips, with the NVL72 rack pairing 72 GPUs to 36 CPUs — the host processor shipping as a component of a rack-scale product rather than as the machine itself.
1 source1 retrieved & read
- SupportsPrimary evidenceRetrieved & readNVIDIA Vera Rubin Opens Agentic AI Frontier
NVIDIA's Grace CPU Superchip pairs 144 Arm Neoverse V2 cores and LPDDR5X memory at up to 1 TB/s with a 900 GB/s NVLink-C2C link to the accelerator — an Arm design in an x86 data centre, specified around the bandwidth of that link rather than around compute.
1 source1 retrieved & read
- SupportsPrimary evidenceRetrieved & readNVIDIA Grace CPU
Timeline
What actually happened, in order, with sources.
- 8 United States
Where this topic’s events took place, as far as our sources establish it. Events with no single location — a standards publication, say — and events we have not yet attributed are both counted as unattributed rather than omitted.
Mar 16, 2026
The CPU Ships as One of Seven Chips
NVIDIA's 16 March 2026 Vera Rubin announcement presented the Vera CPU as one of seven coordinated platform chips, with the NVL72 rack pairing 72 GPUs to 36 CPUs — the processor announced as a component rather than as a product.
1 source1 retrieved & read
- SupportsPrimary evidenceRetrieved & readNVIDIA Vera Rubin Opens Agentic AI Frontier
2026
Two Hundred and Fifty-Six Cores in a Socket
AMD's Venice generation is stated at up to 256 cores per socket on sockets SP7/SP8 with sixteen DDR5 channels and 128 PCIe Gen6 lanes — eight times the cores of the 2017 first generation, assembled from many small dies rather than one large one.
1 source1 retrieved & read
- SupportsRetrieved & readEpyc
Sep 2022
Two Chief Executives Disagree in Public About Moore's Law
In September 2022 NVIDIA's chief executive declared Moore's law dead and Intel's chief executive publicly contradicted him within days — two interested parties whose strategies depend on opposite answers, recorded here as an unresolved dispute.
1 source1 retrieved & read
- SupportsChallengesRetrieved & readMoore's law
2019
The Server Processor Becomes Several Dies
AMD's EPYC Rome in 2019 rebuilt the server processor as eight 7nm compute chiplets around a 14nm I/O die, doubling maximum cores to 64 — separating high-yielding leading-edge compute from cheap mature I/O, five years before accelerators went multi-die for a different reason.
1 source1 retrieved & read
- SupportsRetrieved & readEpyc
2004
Intel Cancels Tejas and Jayhawk at the Power Wall
Intel cancelled the Tejas and Jayhawk processors in 2004 on hitting the power wall; from 2005 clock frequency stagnated at 4–6 GHz and per-CPU power settled near 100 W.
1 source1 retrieved & read
- SupportsRetrieved & readDennard scaling
1975
Moore Revises the Period to Two Years
At the 1975 IEEE International Electron Devices Meeting Moore revised the doubling period from one year to approximately two, a compound growth rate near 41%; Carver Mead popularised the name shortly afterwards.
1 source1 retrieved & read
- SupportsRetrieved & readMoore's law
1974
Dennard States the Scaling Rule
A 1974 paper co-authored by Robert H. Dennard stated that shrinking transistors keeps power density constant so power stays proportional to area — the rule that made each process generation faster at no thermal cost.
1 source1 retrieved & read
- SupportsRetrieved & readDennard scaling
1965
Moore Counts the Doubling
Gordon Moore observed in 1965 that components per integrated circuit had been doubling every year and projected at least another decade of it — an observation about component count and cost, not about speed.
1 source1 retrieved & read
- SupportsRetrieved & readMoore's law
Source register
All 25 sources behind this page — what we read, what we tried to read and could not, and what we looked at and set aside, with the reason in view for each. A concept shared with another collection brings its own references with it, so some entries here were surfaced for a neighbouring topic rather than this one.
- Cited as evidence
- 11
- Tried, could not read
- 1
- Surfaced, set aside
- 13
Cited sources 11 distinct links
Original publisher links. Files open on the publisher’s site; we do not host copies. A linked document is not an additional source or an independent verification.
- AMD EPYC Server CPUs ↗
AMD
- CPU cache ↗
Wikipedia
- Dennard scaling ↗
Wikipedia
- ECC memory ↗
Wikipedia
- Epyc ↗
Wikipedia
- Moore's law ↗
Wikipedia
- Multi-core processor ↗
Wikipedia
- NVIDIA Grace CPU ↗
NVIDIA
- NVIDIA Vera Rubin Opens Agentic AI Frontier ↗
NVIDIA · Published 2026-03-16
- Ryzen ↗
Wikipedia
- Xeon ↗
Wikipedia
Tried, could not read1
We attempted these and were refused or served nothing. Nothing on this page rests on them; they are published so the gaps are checkable rather than invisible.
Surfaced, set aside13
These came up while researching and were deliberately not used. We do not claim to have read them — each is listed with why it was passed over, so the shape of the survey is visible and not just its conclusions.
- 1970 1-Kbit DRAM (Intel, U.S.A.)
- Arm Neoverse V3 processor documentation
- Back to the Future: The 1103 Commercial DRAM has Landed
- Cramming more components onto integrated circuits (Moore, 1965)
- Design of Ion-Implanted MOSFET's with Very Small Physical Dimensions (Dennard et al., 1974)
- Dynamic Random Access Memory (DRAM) Explained — All About Semiconductor
- Intel Xeon Processors
- Memory lane
- NVIDIA Blackwell Architecture Explained: B200, GB200 & PCB Design Impact
- NVIDIA GPU History: GeForce 256 to Vera Rubin
- Robert H. Dennard of IBM Invents DRAM
- Robert H. Dennard, National Inventors Hall of Fame Inductee
- What is CUDA? Parallel programming for GPUs
Coverage & limits
What this page does and does not claim.
Fifth packet. Seeded so that the memory and GPU collections have somewhere to point when they reach for the processor, and expanded in a second pass to treat the server-versus-consumer split as a first-class subject rather than a footnote. Sixteen sources were consulted: eleven were retrieved and read, one was attempted and returned an error shell, and four were surfaced and set aside with a stated reason. The spine of the topic — Dennard scaling, Moore's law, the multicore transition, ECC, the cache hierarchy and the EPYC and Xeon generation tables — rests largely on tertiary encyclopedia accounts rather than primary papers or manufacturer documentation, which is the honest weakness of this collection and the first thing a later pass should fix: Dennard's 1974 IEEE paper is paywalled and was not circumvented, and Moore's 1965 paper was surfaced but not retrieved. AMD's own EPYC page was read in a browser and turned out to be almost entirely comparative marketing — its claims against Intel Xeon and AWS Graviton are AMD's own, made on tests AMD chose, and none is asserted here. Arm's Neoverse page returned a navigation shell, so the licensor's side of the Grace CPU is unsourced. The cache source gives no per-level cycle latencies, so none are quoted. One disagreement is published rather than resolved: NVIDIA's and Intel's chief executives took opposite public positions on whether Moore's law is dead within days of each other in September 2022, and both lead companies whose strategy depends on the answer. Not yet editor-reviewed; every assertion reads as reported.
Source check, 2026-09-17. Numeric-presence checks passed for 18 assertions using available source text, which may be cached. This is not verification of their meaning. What this check does and does not prove →
- Not editor-reviewed unless labelled. Assertions marked Reported are assembled from the sources shown and have not yet been checked by an editor. Only Primary source and Corroborated mean a human verified them.
- Disagreements are preserved, not resolved. Where sources conflict, both accounts appear and the assertion is marked Disputed.
- Retrieval status is disclosed per source. A source we could not open is never counted as evidence for an assertion.
This page is also available as structured data: /api/v1/topics/cpu