Topicsgpu

GPUs and AI Accelerators: How the Chips That Train AI Are Designed, Made and Sold

A sourced reference collection on the AI accelerator — what a GPU actually is and why it is fast, how the reticle limit forced it onto two dies, the four industries that have to deliver before one ships, what the money and the customer concentration look like in the filings, and how export control turned chip performance into a regulated quantity.

Assertions
44
Sources consulted
68
Read in full
28/68
Cited as evidence
28

68 sources sit behind this page — including any that arrive with a concept this page shares with another collection. 28 were retrieved and read in full, and only those can back an assertion. 3 could not be retrieved, and 37 were surfaced and deliberately set aside. Every one of them is named in the register below, with the reason in view. How we source this.

Background

Our own synthesis, written to orient you — not evidence. Every factual statement here is asserted and sourced further down this page.

An AI accelerator is not a faster computer. It is a computer that spent its transistors differently. NVIDIA's own documentation puts the bargain plainly: a GPU is “designed such that more transistors are devoted to data processing rather than data caching and flow control”. A CPU spends its area on caches and control logic so one chain of dependent instructions runs as quickly as possible; a GPU spends the same area on arithmetic units and covers memory latency by keeping enough independent work queued. That is the whole reason the same part is extraordinary at training a neural network and unremarkable at running an operating system.

The category was named on 11 October 1999, when NVIDIA launched the GeForce 256 — 17 million transistors on a TSMC 220 nm process — and defined a GPU as a single-chip processor with integrated transform, lighting, triangle setup and rendering engines handling at least 10 million polygons per second. Its real innovation was moving geometry work off the CPU. In November 2006 CUDA made that parallel engine addressable in a general-purpose language instead of only through the graphics pipeline. Then in 2012 a paper by Krizhevsky, Sutskever and Hinton reported ImageNet top-5 error of 18.9% from a 60-million-parameter network, crediting “a very efficient GPU implementation” for making the training tractable. Six years of available capability, and then a demand curve that has not stopped.

The hardware answered by dedicating silicon to one operation. The V100 of May 2017 introduced tensor cores, which multiply two 4x4 FP16 matrices and add a third — the inner loop of essentially all deep learning. Every generation since has changed the precision rather than the operation: TensorFloat-32 and bfloat16 with the A100 in 2020, an FP8 transformer engine with the H100 in 2022, 4-bit floating point with Blackwell in 2024. FP8 became usable across vendors because NVIDIA, Arm and Intel published a shared format in September 2022 — E4M3 for weights and activations, E5M2 for gradients — rather than each inventing their own.

Then physics intervened. A lithography scanner can only project so large an image in one exposure: 26 mm by 33 mm at the 0.33 numerical aperture of current EUV tools, and 26 mm by 16.5 mm at high-NA, because those optics are anamorphic. NVIDIA's die area sat between 610 and 826 square millimetres across the P100, V100, A100 and H100 while transistor count rose from 15.3 to 80 billion — all of the growth came from process density, against a wall the area could not cross. Blackwell went around it: 208 billion transistors as two reticle-limited dies joined by a 10 TB/s link and presented as one GPU. At that point accelerator design stops being a logic problem and becomes a packaging problem.

Nobody makes one of these alone. NVIDIA states in its annual report that it uses “a fabless and contracting manufacturing strategy” covering wafer fabrication, assembly, testing and packaging, and names the suppliers: wafers from TSMC and Samsung, memory from SK hynix, Micron and Samsung, CoWoS for packaging, and Hon Hai, Wistron and Fabrinet for assembly and test. In the quarter ended 30 June 2026 TSMC reported that 77% of its wafer revenue came from 7-nanometer and below. The coupling is tight enough to be visible in the calendar: on 16 March 2026 NVIDIA announced the Vera Rubin platform and Micron announced volume production of HBM4 built for it — the same day, because the accelerator generation and the memory generation are specified and qualified together.

The money is concentrated at both ends. NVIDIA's fiscal 2026, ended 25 January 2026, produced $215.9 billion of revenue against $60.9 billion two years earlier, with $193.7 billion of it from Data Center and $120.1 billion of net income — from a company that owns no factory. In the same year one direct customer was 22% of total revenue and another was 14%. Three memory makers and one leading foundry at one end; a handful of buyers at the other. The two quarters since have run $81.6 billion and $96.2 billion, with guidance of $108.0 billion issued on 26 August 2026 — a projection, not a result.

Competition now comes from two directions at once. AMD launched the Instinct MI400 series on 23 July 2026, competing at rack scale rather than by the card. And the largest buyers are building their own: Google documents its seventh-generation TPU at 192 GiB of HBM per chip and pods of 9,216 chips, and Amazon announced Trainium3 in December 2025 as its first 3-nanometer part. TrendForce projected in January 2026 that ASIC-based AI servers would reach nearly 28% of shipments against 69.7% for GPU-based systems — a forecast published before the year it describes, and treated here as one.

Since 7 October 2022 the performance of these chips has been a regulated quantity. The US rule creating ECCN 3A090 controls circuits that combine 600 GByte/s of aggregate input-output with a bit-length-times-TOPS product of 4,800 or more, naming GPUs, TPUs and neural processors as in scope. What that costs is on the record: an April 2025 licence requirement for the H20 produced a $4.5 billion charge, later licences yielded about $60 million of revenue, and by August 2026 NVIDIA's own guidance assumed no Data Center compute revenue from China at all.

What is missing is stated beside the claims rather than hidden. NVIDIA has not published per-GPU HBM4 capacity, bandwidth or FP4 throughput for Rubin, and the figures circulating for those come from resellers and conference coverage, so none is recorded here. AMD gives a rack total of 3 AI exaflops but still no HBM4 capacity, aggregate bandwidth or precision split. Intel publishes no node, memory or bandwidth for Gaudi 3. TSMC's revenue-by-platform split lives in a slide deck that could not be retrieved. One report places the Vera Rubin full-production announcement at a CES keynote rather than the March newsroom post, and that discrepancy is published rather than resolved. And the widely-quoted claim that NVIDIA holds 70, 80 or 92 percent of this market is asserted nowhere here, because no source read for this collection states one with a denominator you could check.

Figures

Every number below is asserted and sourced elsewhere on this page.

Three fiscal years

NVIDIA's total revenue as reported in its annual report. The company owns no factory; this is what a design plus a purchase order was worth.

Revenue, billions of US dollars

$60.9bn

FY2024

$130.5bn

FY2025

$215.9bn

FY2026

NVIDIA Form 10-K for the fiscal year ended 25 January 2026, filed 25 February 2026. All three figures are reported results from the same filing's comparative tables, not projections.

What the revenue actually is

NVIDIA's fiscal 2026 revenue by end market. Data centre compute alone is three quarters of the company; gaming, the business the company was built on, is 7%.

Total $215.9bn

  • Data Center — compute$162.4bn
  • Data Center — networking$31.4bn
  • Gaming$16.0bn
  • Professional Visualization$3.2bn
  • Automotive$2.3bn
  • OEM and other$0.6bn

NVIDIA Form 10-K for the fiscal year ended 25 January 2026. Reported figures using NVIDIA's own end-market definitions, from the revenue-by-specialized-markets table.

Two quarters and a projection

The most recent reported quarters, and the guidance issued alongside the second. The third column has not happened.

Revenue, billions of US dollars

$81.6bn

FQ1-27

$96.2bn

FQ2-27

$108.0bn

FQ3-27 guide

NVIDIA quarterly results releases of 20 May 2026 and 26 August 2026. The first two columns are reported results; the third is NVIDIA's own guidance of $108.0bn plus or minus 2%, issued 26 August 2026 and stated to assume no Data Center compute revenue from China.

The die that stopped growing

Four generations of NVIDIA data-centre silicon. Area flattens against the reticle limit after 2017 while transistor count rises from 15.3 to 80 billion — and then the next generation needed two dies.

Die area, square millimetres

610 mm²

GP100 (2016)

815 mm²

GV100 (2017)

826 mm²

GA100 (2020)

814 mm²

GH100 (2022)

Die areas from the Wikipedia articles on the Pascal, Volta, Ampere and Hopper microarchitectures, each read directly. Tertiary sources, reporting manufacturer specifications; measured die areas, not projections.

Where the leading foundry's wafers go

TSMC's wafer revenue by process node in the quarter ended 30 June 2026. Nodes that barely existed five years ago are 77% of it.

Total 100% of wafer revenue

  • 2 nm3%
  • 3 nm30%
  • 5 nm33%
  • 7 nm11%
  • Older than 7 nm23%

TSMC second-quarter 2026 earnings release, filed with the SEC on 16 July 2026. Reported percentages of total wafer revenue. The release states these figures had not been approved by the board of directors when published. The 'Older than 7nm' share is the remainder of the disclosed mix.

The buyers moved home

Share of NVIDIA's revenue by the headquarters location of its direct customers. In two fiscal years the non-US share fell from nearly half to under a third.

52%
48%

4 pts

FY2024

59%
41%

18 pts

FY2025

69%
31%

38 pts

FY2026

The gap opens from 4 to 38 percentage points in two years.

NVIDIA Form 10-K for the fiscal year ended 25 January 2026. The filing discloses the share from customers headquartered outside the United States (48%, 41%, 31%); the United States series is its complement. Revenue is attributed to a direct customer's headquarters, which the filing warns is not where the equipment ends up.

Concepts

The vocabulary this subject is built from, and what we can show about each.

Accelerator Customer Concentration

other

NVIDIA's fiscal 2026 revenue by customer-headquarters location was $149,617 million United States, $42,345 million Taiwan, $19,677 million China including Hong Kong and $4,299 million other, within total revenue of $215,938 million; the non-US share fell from 48% to 41% to 31% across fiscal 2024-2026, and the company estimates 76% of Data Center revenue from Taiwan-headquartered customers went to end customers in the US and Europe.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

NVIDIA reported that in fiscal 2026 one direct customer represented 22% of total revenue and another 14%, against a flatter fiscal 2025 pattern of 12%, 11% and 11%, and states its revenue is concentrated among a limited number of both direct and indirect customers.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

Accelerator Die Scaling

process

NVIDIA data-centre die area sat between 610 and 826 mm-squared across the P100, V100, A100 and H100 while transistor count rose from 15.3 to 80 billion — growth came from process density until the reticle limit forced the next generation onto two dies.

ReportedSupported by the sources below, not yet editor-reviewed.
4 sources4 retrieved & read

Accelerator Power Envelope

other

Accelerator power rose from a 400 W TDP on the A100 in 2020 to 700 W on the SXM5 H100 in 2022, and by Blackwell the specified unit is a liquid-cooled rack rather than a card — which is what ties the accelerator roadmap to data-centre and grid design.

ReportedSupported by the sources below, not yet editor-reviewed.
3 sources3 retrieved & read

AI Accelerator Market Structure

other

NVIDIA's filing names AMD, Huawei and Intel as accelerator competitors, Alibaba, Alphabet, Amazon, Baidu, Huawei and Microsoft as cloud companies designing their own AI hardware, and nine companies in networking; AMD announced the Instinct MI400 series and the Helios rack on 23 July 2026, with the MI455X for frontier AI and the MI430X at up to 288 TFLOPS FP64.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

TrendForce forecast on 20 January 2026 that 2026 AI server shipments would grow over 28% year on year, with GPU-based systems at 69.7% of shipments and ASIC-based systems near 28%, and top-five North American cloud capital expenditure rising 40% — all projections, not measured outcomes.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

AMD previewed the Helios rack at up to 3 AI exaflops built on Instinct MI455X, EPYC 'Venice' and Pensando 'Vulcano' parts, while Intel's Gaudi 3 connects over standard Ethernet rather than a proprietary fabric — the clearest structural difference between the merchant alternatives and the incumbent.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

NVIDIA reported fiscal 2026 revenue of $215.938 billion, up 65%, of which Data Center was $193.737 billion ($162.361bn compute, $31.376bn networking); net income was $120.067 billion at $4.90 per diluted share and gross margin fell to 71.1% from 75.0%, with Blackwell stated as the majority of Data Center revenue.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

NVIDIA reported $81.6 billion of revenue in the quarter ended 26 April 2026 ($75.2bn Data Center) and $96.2 billion in the quarter ended 26 July 2026 ($89.0bn Data Center, 75.0% gross margins, $2.46 GAAP diluted EPS), and guided to $108.0 billion plus or minus 2% for the following quarter assuming no Data Center compute revenue from China.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

AI Accelerator Numeric Formats

other

The September 2022 FP8 proposal — a joint NVIDIA, Arm and Intel paper — defines E4M3 and E5M2 encodings, with E4M3 deliberately departing from IEEE 754 to buy dynamic range, recommends E4M3 for weights and activations and E5M2 for gradients, and reports training quality matching 16-bit formats.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

AI Chip Export Controls

other

NVIDIA's filing records export restrictions from August 2022 onward affecting the A100 and H100, an April 2025 licence requirement for the H20 that produced a $4.5 billion charge, roughly $60 million of licensed H20 revenue in August 2025, a February 2026 H200 licence under which no revenue had been generated, and forward guidance in August 2026 assuming no Data Center compute revenue from China.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

A rule published 25 October 2023 (88 FR 73458, effective 17 November 2023) replaced ECCN 3A090's original structure with tests on 'total processing performance' and a new 'performance density' — 4,800 TPP, or 1,600 TPP with density 5.92 — explicitly to stop buyers aggregating many smaller chips, and extended the requirement to Country Groups D:1, D:4 and D:5.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

The Bureau of Industry and Security's rule published 13 October 2022 created ECCNs 3A090 and 4A090, controlling integrated circuits that both reach 600 GByte/s of aggregate bidirectional I/O to non-volatile-memory circuits and a bit-length-times-TOPS product of 4,800 or more, with a note naming GPUs, TPUs, neural processors and FPLDs as in scope.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

AI Compute Demand

other

The March 2022 Chinchilla paper found large language models were significantly undertrained and that training tokens should double with every doubling of model size — redirecting the field from larger models toward more training data, which is more compute either way.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

AI Rack Power Density

componentshared from another collection — see its own page for what it asserts

Chiplet

componentshared from another collection — see its own page for what it asserts

CoWoS (Chip-on-Wafer-on-Substrate)

packaging technologyshared from another collection — see its own page for what it asserts

CUDA

other

NVIDIA introduced CUDA in November 2006 as a general-purpose parallel computing platform and programming model, making the GPU's parallel engine addressable directly rather than only through graphics operations.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

NVIDIA's annual report names CUDA as one of the fundamental building blocks of a unified architecture spanning its end markets, and lists software support and API conformity among the principal competitive factors in its market.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

Custom AI ASIC

component

Amazon announced Trainium3 on 4 December 2025 as its first 3nm AI chip, with up to 144 chips per Trn3 UltraServer and company claims of up to 4.4x the compute and 4x the energy efficiency of Trainium2 UltraServers.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

Google documents its seventh-generation TPU7x at 2,307 TFLOPs BF16 and 4,614 TFLOPs FP8 per chip, 192 GiB of HBM at 7,380 GB/s, 1,200 GB/s of bidirectional inter-chip interconnect, and a maximum pod of 9,216 chips.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

Graphics Processing Unit (GPU)

component

NVIDIA's own documentation states the architectural bargain: a GPU devotes more transistors to data processing rather than to data caching and flow control, trading single-thread speed for parallel throughput and hiding memory latency with queued independent work.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

NVIDIA defined the GPU at the GeForce 256's launch on 11 October 1999 as a single-chip processor with integrated transform, lighting, triangle setup/clipping and rendering engines processing at least 10 million polygons per second — 17 million transistors on a TSMC 220 nm process, and the first consumer part to move geometry work off the CPU.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

High Bandwidth Memory (HBM)

memory technologyshared from another collection — see its own page for what it asserts

Leading-Edge Process Node

processshared from another collection — see its own page for what it asserts

Reticle Limit

process

The reticle limit is the largest image a scanner can project in one exposure: 26 mm by 33 mm at 0.33 NA EUV, halving to 26 mm by 16.5 mm at 0.55 NA because high-NA optics are anamorphic — so finer resolution is bought by giving up die area.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

Cerebras builds a single wafer-scale processor of 46,225 square millimetres carrying 4 trillion transistors and 900,000 cores — roughly fifty times the area a lithography scanner can expose in one exposure — as an alternative to joining many reticle-limited dies.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

Blackwell, announced 18 March 2024, is two reticle-limited dies joined by a 10 TB/s chip-to-chip link and presented as one GPU — 208 billion transistors on a custom TSMC 4NP process — because a single die had reached the size the optics allow.

ReportedSupported by the sources below, not yet editor-reviewed.
2 sources2 retrieved & read

Tensor Core

component

A tensor core multiplies two small matrices and adds a third — in the 2017 V100, two 4x4 FP16 matrices plus an FP16 or FP32 accumulator — and every generation since has changed the numeric precision rather than the operation: TF32 and bfloat16 in 2020, FP8 in 2022, FP4 in 2024.

ReportedSupported by the sources below, not yet editor-reviewed.
4 sources4 retrieved & read

The Fabless Accelerator Supply Chain

other

NVIDIA's annual report states it uses a fabless and contracting manufacturing strategy for all phases including wafer fabrication, assembly, testing and packaging, and names TSMC and Samsung as foundries, SK hynix, Micron and Samsung for memory, CoWoS for packaging, and Hon Hai, Wistron and Fabrinet for assembly and test — with the supply chain mainly concentrated in Asia.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

In the quarter ended 30 June 2026 TSMC reported 2nm at 3% of wafer revenue, 3nm at 30%, 5nm at 33% and 7nm at 11%, with 7nm-and-below totalling 77%, on revenue of US$40.20 billion at a 67.7% gross margin.

ReportedSupported by the sources below, not yet editor-reviewed.
1 source1 retrieved & read

Timeline

What actually happened, in order, with sources.

Coverage
  • 15 United States
  • 2 Unattributed

Where this topic’s events took place, as far as our sources establish it. Events with no single location — a standards publication, say — and events we have not yet attributed are both counted as unattributed rather than omitted.

  1. Aug 26, 2026

    NVIDIA Reports $96.2 Billion in a Quarter

    otherUnited States

    NVIDIA reported $96.2 billion of revenue for the quarter ended 26 July 2026, up 106% year on year with $89.0 billion from Data Center at 75.0% gross margins, and guided to $108.0 billion plus or minus 2% while assuming no Data Center compute revenue from China.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  2. Jul 23, 2026

    AMD Launches the Instinct MI400 Series

    technology generation milestoneUnited States

    AMD launched the Instinct MI400 series on 23 July 2026 — the MI455X for frontier AI within the Helios rackscale solution and the MI430X for sovereign AI and HPC at up to 288 TFLOPS of hardware FP64 — establishing a second merchant supplier competing at rack scale.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  3. Mar 16, 2026

    NVIDIA Announces the Vera Rubin Platform

    technology generation milestoneUnited States

    NVIDIA announced the Vera Rubin platform on 16 March 2026 as seven coordinated chips — Vera CPU, Rubin GPU, NVLink 6 Switch, ConnectX-9 SuperNIC, BlueField-4 DPU, Spectrum-6 Ethernet switch and Groq 3 LPU — with the Vera Rubin NVL72 of 72 GPUs and 36 CPUs available from partners in the second half of 2026.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  4. Mar 16, 2026

    Micron Puts HBM4 for Vera Rubin into Volume Production

    production milestoneUnited States

    Micron announced high-volume production of HBM4 for NVIDIA's Vera Rubin on 16 March 2026 — 36GB in a 12-high stack at more than 2.8 TB/s and pin speeds above 11 Gb/s, with volume shipment begun in the first calendar quarter of 2026 — the same day the accelerator platform was announced.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  5. Dec 4, 2025

    Amazon Announces Trainium3

    technology generation milestoneUnited States

    Amazon announced Trainium3 on 4 December 2025 as its first 3nm AI chip, with up to 144 chips per Trn3 UltraServer and company-claimed gains of up to 4.4x compute and 4x energy efficiency over Trainium2 UltraServers.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  6. Apr 2025

    H20 Licence Requirement Produces a $4.5 Billion Charge

    otherUnited States

    In April 2025 the US government required a licence for NVIDIA's H20 and any circuit matching its memory or interconnect bandwidth for export to China and D:5 countries; NVIDIA took a $4.5 billion charge in Q1 fiscal 2026 and later generated approximately $60 million of H20 revenue under August 2025 licences.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  7. Mar 18, 2024

    Blackwell Goes to Two Reticle-Limited Dies

    technology generation milestoneUnited States

    NVIDIA announced Blackwell on 18 March 2024 — 208 billion transistors on custom TSMC 4NP as two reticle-limited dies joined by a 10 TB/s link, with FP4 inference, fifth-generation NVLink at 1.8 TB/s per GPU, and the liquid-cooled 72-GPU GB200 NVL72 rack.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  8. Nov 17, 2023

    Export Thresholds Rewritten Around Performance Density

    otherUnited States

    A rule effective 17 November 2023 rewrote ECCN 3A090 around 'total processing performance' and a new 'performance density' test — explicitly to stop buyers aggregating many smaller chips — and extended the licence requirement to Country Groups D:1, D:4 and D:5.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  9. Oct 7, 2022

    United States Controls Advanced Computing Chips by Performance

    otherUnited States

    The Bureau of Industry and Security's advanced-computing rule took effect from 7 October 2022, creating ECCNs 3A090 and 4A090 and defining a control threshold of 600 GByte/s aggregate I/O combined with a bit-length-times-TOPS product of 4,800 — making accelerator performance a regulated quantity.

    ReportedSupported by the sources below, not yet editor-reviewed.
    2 sources2 retrieved & read
  10. Sep 12, 2022

    NVIDIA, Arm and Intel Jointly Propose an FP8 Interchange Format

    standards milestone

    A 12 September 2022 paper from NVIDIA, Arm and Intel proposed the FP8 interchange format of E4M3 and E5M2 encodings, with E4M3 departing from IEEE 754 to extend dynamic range, recommending E4M3 for weights and activations and E5M2 for gradients.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  11. Mar 2022

    H100 Brings HBM3, the Transformer Engine and 700 Watts

    technology generation milestoneUnited States

    Hopper was revealed in March 2022: the H100 at 80 billion transistors on 814 mm-squared of TSMC N4, up to 80 GB of HBM3 at 3 TB/s, the first transformer engine dropping FP16 to FP8, and a 700 W SXM5 thermal design power.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  12. May 14, 2020

    A100 Adds TF32, Partitioning and a 400-Watt Envelope

    technology generation milestoneUnited States

    The A100, announced 14 May 2020, put 54.2 billion transistors on 826 mm-squared of TSMC N7 — effectively at the reticle limit — with third-generation tensor cores adding TF32 and bfloat16, partitioning into up to seven instances, and a 400 W thermal design power.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  13. May 2017

    Volta V100 Introduces Tensor Cores

    technology generation milestoneUnited States

    The V100, announced May 2017, put 21.1 billion transistors on 815 mm-squared of TSMC 12 nm FinFET with HBM2 at 900 GB/s and introduced tensor cores — fixed-function units multiplying two 4x4 FP16 matrices and adding a third.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  14. 2012

    AlexNet Wins ImageNet on a GPU Implementation

    technology generation milestone

    The 2012 AlexNet paper reported top-1 and top-5 ImageNet error rates of 39.7% and 18.9% from a 60-million-parameter network, crediting a very efficient GPU implementation of convolutional nets for making training tractable.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  15. Nov 2006

    NVIDIA Introduces CUDA

    technology generation milestoneUnited States

    NVIDIA introduced CUDA in November 2006, making the GPU's parallel compute engine addressable as a general-purpose programming platform rather than only through the graphics pipeline.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read
  16. Oct 11, 1999

    NVIDIA Ships the GeForce 256 and Names the Category

    technology generation milestoneUnited States

    NVIDIA launched the GeForce 256 on 11 October 1999 — 17 million transistors on a TSMC 220 nm process, with an integrated transform and lighting engine that moved geometry work off the CPU — and marketed it with the definition of GPU the industry adopted.

    ReportedSupported by the sources below, not yet editor-reviewed.
    1 source1 retrieved & read

Source register

All 68 sources behind this page — what we read, what we tried to read and could not, and what we looked at and set aside, with the reason in view for each. A concept shared with another collection brings its own references with it, so some entries here were surfaced for a neighbouring topic rather than this one.

Cited as evidence
28
Tried, could not read
3
Surfaced, set aside
37
Cited sources 28 distinct links

Original publisher links. Files open on the publisher’s site; we do not host copies. A linked document is not an additional source or an independent verification.

Tried, could not read3

We attempted these and were refused or served nothing. Nothing on this page rests on them; they are published so the gaps are checkable rather than invisible.

Surfaced, set aside37

These came up while researching and were deliberately not used. We do not claim to have read them — each is listed with why it was passed over, so the shape of the survey is visible and not just its conclusions.

Coverage & limits

What this page does and does not claim.

Fourth packet built to the reference-collection template rather than a news window, and the first covering the compute side of the AI supply chain. Sixty-five sources were consulted: twenty-eight were retrieved and read, three were attempted and could not be read, and thirty-four were surfaced and deliberately set aside with a stated reason each. The strongest material is regulatory — NVIDIA's fiscal 2026 annual report, TSMC's filed quarterly release, and the Federal Register texts of both the October 2022 advanced-computing rule and its October 2023 revision — all read through ordinary browser access to public pages after the automated fetcher was refused. Those two rule texts together correct a widely-repeated misstatement of the 2022 thresholds and show the control being rewritten around performance density once hardware had been designed against it. Vendor performance multiples are attributed to the vendor throughout and never treated as measurements; no market-share percentage is asserted anywhere, because no source read here states one with a traceable denominator. One retrieved figure — A100 memory bandwidth — was discarded as an evident misread and the omission is disclosed rather than papered over. Remaining named gaps: Rubin's per-GPU memory and compute figures, which NVIDIA has not published; AMD's per-GPU and aggregate Helios figures, which neither AMD document states; Intel's Gaudi 3 node, memory and bandwidth, likewise unpublished; and TSMC's revenue-by-platform split, which lives in an unretrieved slide deck. Geographically this collection reaches the United States, Taiwan, South Korea, Japan and China across twenty-seven years. Not yet editor-reviewed; every assertion reads as reported.

Source check, 2026-09-17. Numeric-presence checks passed for 44 assertions using available source text, which may be cached. This is not verification of their meaning. What this check does and does not prove →

  • Not editor-reviewed unless labelled. Assertions marked Reported are assembled from the sources shown and have not yet been checked by an editor. Only Primary source and Corroborated mean a human verified them.
  • Disagreements are preserved, not resolved. Where sources conflict, both accounts appear and the assertion is marked Disputed.
  • Retrieval status is disclosed per source. A source we could not open is never counted as evidence for an assertion.

This page is also available as structured data: /api/v1/topics/gpu