An AI accelerator is not a faster computer. It is a computer that spent its transistors differently. NVIDIA's own documentation puts the bargain plainly: a GPU is “designed such that more transistors are devoted to data processing rather than data caching and flow control”. A CPU spends its area on caches and control logic so one chain of dependent instructions runs as quickly as possible; a GPU spends the same area on arithmetic units and covers memory latency by keeping enough independent work queued. That is the whole reason the same part is extraordinary at training a neural network and unremarkable at running an operating system.
The category was named on 11 October 1999, when NVIDIA launched the GeForce 256 — 17 million transistors on a TSMC 220 nm process — and defined a GPU as a single-chip processor with integrated transform, lighting, triangle setup and rendering engines handling at least 10 million polygons per second. Its real innovation was moving geometry work off the CPU. In November 2006 CUDA made that parallel engine addressable in a general-purpose language instead of only through the graphics pipeline. Then in 2012 a paper by Krizhevsky, Sutskever and Hinton reported ImageNet top-5 error of 18.9% from a 60-million-parameter network, crediting “a very efficient GPU implementation” for making the training tractable. Six years of available capability, and then a demand curve that has not stopped.
The hardware answered by dedicating silicon to one operation. The V100 of May 2017 introduced tensor cores, which multiply two 4x4 FP16 matrices and add a third — the inner loop of essentially all deep learning. Every generation since has changed the precision rather than the operation: TensorFloat-32 and bfloat16 with the A100 in 2020, an FP8 transformer engine with the H100 in 2022, 4-bit floating point with Blackwell in 2024. FP8 became usable across vendors because NVIDIA, Arm and Intel published a shared format in September 2022 — E4M3 for weights and activations, E5M2 for gradients — rather than each inventing their own.
Then physics intervened. A lithography scanner can only project so large an image in one exposure: 26 mm by 33 mm at the 0.33 numerical aperture of current EUV tools, and 26 mm by 16.5 mm at high-NA, because those optics are anamorphic. NVIDIA's die area sat between 610 and 826 square millimetres across the P100, V100, A100 and H100 while transistor count rose from 15.3 to 80 billion — all of the growth came from process density, against a wall the area could not cross. Blackwell went around it: 208 billion transistors as two reticle-limited dies joined by a 10 TB/s link and presented as one GPU. At that point accelerator design stops being a logic problem and becomes a packaging problem.
Nobody makes one of these alone. NVIDIA states in its annual report that it uses “a fabless and contracting manufacturing strategy” covering wafer fabrication, assembly, testing and packaging, and names the suppliers: wafers from TSMC and Samsung, memory from SK hynix, Micron and Samsung, CoWoS for packaging, and Hon Hai, Wistron and Fabrinet for assembly and test. In the quarter ended 30 June 2026 TSMC reported that 77% of its wafer revenue came from 7-nanometer and below. The coupling is tight enough to be visible in the calendar: on 16 March 2026 NVIDIA announced the Vera Rubin platform and Micron announced volume production of HBM4 built for it — the same day, because the accelerator generation and the memory generation are specified and qualified together.
The money is concentrated at both ends. NVIDIA's fiscal 2026, ended 25 January 2026, produced $215.9 billion of revenue against $60.9 billion two years earlier, with $193.7 billion of it from Data Center and $120.1 billion of net income — from a company that owns no factory. In the same year one direct customer was 22% of total revenue and another was 14%. Three memory makers and one leading foundry at one end; a handful of buyers at the other. The two quarters since have run $81.6 billion and $96.2 billion, with guidance of $108.0 billion issued on 26 August 2026 — a projection, not a result.
Competition now comes from two directions at once. AMD launched the Instinct MI400 series on 23 July 2026, competing at rack scale rather than by the card. And the largest buyers are building their own: Google documents its seventh-generation TPU at 192 GiB of HBM per chip and pods of 9,216 chips, and Amazon announced Trainium3 in December 2025 as its first 3-nanometer part. TrendForce projected in January 2026 that ASIC-based AI servers would reach nearly 28% of shipments against 69.7% for GPU-based systems — a forecast published before the year it describes, and treated here as one.
Since 7 October 2022 the performance of these chips has been a regulated quantity. The US rule creating ECCN 3A090 controls circuits that combine 600 GByte/s of aggregate input-output with a bit-length-times-TOPS product of 4,800 or more, naming GPUs, TPUs and neural processors as in scope. What that costs is on the record: an April 2025 licence requirement for the H20 produced a $4.5 billion charge, later licences yielded about $60 million of revenue, and by August 2026 NVIDIA's own guidance assumed no Data Center compute revenue from China at all.
What is missing is stated beside the claims rather than hidden. NVIDIA has not published per-GPU HBM4 capacity, bandwidth or FP4 throughput for Rubin, and the figures circulating for those come from resellers and conference coverage, so none is recorded here. AMD gives a rack total of 3 AI exaflops but still no HBM4 capacity, aggregate bandwidth or precision split. Intel publishes no node, memory or bandwidth for Gaudi 3. TSMC's revenue-by-platform split lives in a slide deck that could not be retrieved. One report places the Vera Rubin full-production announcement at a CES keynote rather than the March newsroom post, and that discrepancy is published rather than resolved. And the widely-quoted claim that NVIDIA holds 70, 80 or 92 percent of this market is asserted nowhere here, because no source read for this collection states one with a denominator you could check.