The AI Chip Supercycle and Custom Silicon
A "chip" is the small piece of silicon that does the actual math inside a computer, and modern AI runs on specialized chips called accelerators that are built to crunch the huge number of calculations a model needs.
What's happening now
As of mid-2026 the whole semiconductor industry is forecast by IDC to hit roughly 1.29 trillion dollars in revenue, up about 53 percent in a single year from 842.8 billion in 2025, with AI data-center chips as the main engine. The sharpest shift is toward custom chips: research firm TrendForce projects custom ASIC shipments from cloud providers will grow 44.6 percent in 2026 versus 16.1 percent for merchant GPUs, pushing custom-chip servers to about 27.8 percent of AI servers. The marquee in-house chips are Google's TPU v7 "Ironwood," Amazon's Trainium 3, Microsoft's Maia 200, and Meta's MTIA, and both Amazon and Google now report passing one million deployed units of their own silicon. On June 5, 2026, Broadcom (which co-designs many of these chips) guided to about 56 billion dollars in fiscal-2026 AI revenue, up roughly 1.8 times year over year from about 20 billion dollars. Underneath it all, the memory that feeds these chips, High Bandwidth Memory (HBM), is sold out: SK Hynix, Samsung, and Micron are racing to mass-produce HBM4, demand is growing well over 70 percent in 2026, and the companies warn shortages could persist into 2027 and beyond. Nvidia still holds roughly 70 to 75 percent of the AI accelerator market, but custom silicon is the fastest-growing slice.
What it is
For years almost all of those chips were graphics processors (GPUs) bought from a single company, Nvidia. The AI Chip Supercycle is the current, unusually large and sustained surge in spending on these chips, and "custom silicon" is the newer twist: the biggest cloud companies are now designing their own in-house AI chips instead of only buying off-the-shelf ones, to cut cost and avoid depending on one supplier.
The axis is whether hyperscaler custom silicon will structurally erode Nvidia's AI accelerator dominance or merely optimize a narrow workload slice while Nvidia keeps widening its lead across training, software, and interconnect.
Nvidia Fortress: the moat is widening, not narrowingCustom silicon proves Nvidia's thesis rather than threatening it, because hyperscalers self-supply a narrow inference slice while Nvidia keeps sole control of frontier training, the CUDA ecosystem, and the rack-scale interconnect that no ASIC vendor replicates.
The Fortress case rests on a vertically integrated platform whose switching costs compound every year even as percentage share modestly contracts. CUDA carries more than 5 million active developers and tuned libraries, so switching is an institutional and human-capital gap, not a benchmark gap. Custom ASICs solve one narrow problem, inference cost at hyperscale for fixed architectures, while leaving Nvidia in control of frontier training, programmable experimentation, and NVLink, NVSwitch, and InfiniBand interconnect that no ASIC vendor replicates end to end. The hyperscalers prove the thesis: Google, Amazon, and Microsoft deploy their own ASICs and accelerate Nvidia purchases at the same time, with GPU-based systems accounting for roughly 60 percent of AWS AI server build-out in 2026 and Microsoft confirmed as an early adopter of Vera Rubin NVL72 while it deploys its own Maia chips. Absolute revenue is the tell: Nvidia data center revenue grew 68 percent year over year to $193.7 billion in FY2026, and at GTC 2026 Huang announced combined Blackwell and Vera Rubin orders reaching $1 trillion through 2027, with Vera Rubin in full production delivering 5x inference and 3.5x training performance over Blackwell. A roadmap that cadence outruns any fixed ASIC program.
Jensen Huang (Nvidia CEO); institutional analysts at Morgan Stanley, Goldman Sachs, and Bernstein who have held overweight ratings through 2025-2026; research firms Silicon Analysts and ABI Research; long-term bulls including Wedbush analyst Dan Ives
Hyperscaler Breakaway: custom silicon will meaningfully erode Nvidia's shareAs inference becomes the dominant AI workload by volume and the cost gap turns into a fiduciary obligation, hyperscaler ASICs will capture the most margin-rich slice and bend Nvidia's inference share and pricing power downward.
This camp does not claim wholesale displacement; it claims the economic center of gravity is shifting from training to inference, the workload where custom silicon enjoys a decisive and compounding cost advantage, and the five largest buyers of Nvidia hardware have built the design partnerships and deployment scale to route a growing majority of inference spend away permanently. Inference already represents about two-thirds of AI compute cycles and is projected to reach 75 percent by 2030. The evidence of decoupling is concrete: Google has deployed Ironwood TPU clusters at 9,216-chip scale delivering 42.5 exaflops, with the program 30x more power-efficient per FLOP than its 2018 baseline, and AWS disclosed over 500,000 Trainium2 chips in production as of late 2025. Broadcom, the design partner behind Google TPU, Meta MTIA, Microsoft Maia, and the OpenAI and Anthropic programs, reported $8.4 billion in AI semiconductor revenue in Q1 FY2026, up 106 percent year over year, with a $73 billion committed backlog and a target above $100 billion in AI revenue for FY2027. Custom ASIC shipments are growing at 44.6 percent CAGR versus 16.1 percent for merchant GPUs, and New Street Research projects Nvidia's inference share could fall to 20 to 30 percent by 2028.
Google (TPU), AWS (Trainium, Inferentia), Microsoft (Maia), Meta (MTIA); Broadcom as primary ASIC co-design partner; analysts at New Street Research, SemiAnalysis, Bloomberg Intelligence, and Morgan Stanley
Supercycle Skeptic: the supercycle is real on revenue but mispriced on share and marginTotal AI spend keeps growing, but custom in-house silicon is already eating the highest-margin inference economics Nvidia would otherwise own, and the current valuation prices in monopoly rents that captive ASICs structurally foreclose.
The Skeptic argues the supercycle framing is right on the revenue numerator and wrong on the margin and share denominator, and that error is what makes it dangerous as a thesis. Nvidia's four largest customers, Google, Amazon, Microsoft, and Meta, account for roughly half of its data center revenue, and each has committed billions to proprietary accelerators precisely because Nvidia pricing power destroys cloud margin at scale. Google has deployed over one million TPU units internally and now sells TPU capacity externally; Amazon's Trainium and Inferentia have crossed a $20 billion annualized run rate with customers reporting 50 percent lower cost per inference versus GPU equivalents; Meta is scaling MTIA across three billion users while raising 2026 capex to $125 to $145 billion. The shift is asymmetric: inference now represents about two-thirds of compute cycles, and fixed-function silicon delivers 3 to 5x better performance per watt there than a training-era GPU. TrendForce projects custom ASIC server shipments growing 44.6 percent in 2026 against 16.1 percent for GPU systems, reaching 27.8 percent of AI server shipments, while New Street projects Nvidia inference share falling from over 90 percent to 20 to 30 percent by 2028. With Nvidia's market cap above $4 trillion pricing in sustained dominance, a 20 to 30 point inference migration implies a $15 to $25 billion annual revenue reduction in its most defensible segment.
Jay Goldberg (Seaport Research); Meta VP of Engineering Yee Jiun Song; Bloomberg Intelligence infrastructure analysts; TrendForce semiconductor research; New Street Research AI hardware analysts; value-skeptic investors
Bottleneck Realist: real inference erosion, no structural displacementCustom silicon will take meaningful inference share but cannot break Nvidia's dominance, because the flexibility gap, the concentration of who can afford ASICs, the annual roadmap, and the CUDA software moat all hold across the full stack.
The Realist accepts custom silicon is real and growing fast yet sees structural coexistence rather than displacement, resting on four interlocking claims. First, the flexibility gap is not closing: once taped out, an ASIC is frozen, so in an era of annual architecture shifts the GPU remains the only compute that can absorb a novel model on the day it ships, and training still represents roughly one-third of AI compute cycles. Second, the programs are concentrated and partial even inside their owners' walls: Microsoft Maia runs about 30 percent of Azure AI workloads as of late 2025, leaving 70 percent on Nvidia; AWS Trainium delivers roughly 70 percent of GPU efficiency and is positioned as cost-disruptive rather than performance-leading; Meta MTIA handles recommendation workloads but cannot satisfy full LLM inference, which is why Meta still buys Nvidia at scale. Third, the annual roadmap is a weapon: Nvidia ships Blackwell in 2025, Vera Rubin in H2 2026, Rubin Ultra in 2027, and Feynman in 2028, while custom ASICs need 18 to 24 month design cycles plus software maturation, so they perpetually chase a moving target, and Vera Rubin claims 10x higher inference throughput and 10x lower cost-per-token versus Blackwell. Fourth, CUDA spans more than 20 years of optimization and 5 million active developers; ROCm 7 delivering 3.5x better inference than prior versions is an acknowledgment of how far behind it has been, not proof the gap is closed. Net result: share erodes at the inference margin, modestly, while training, frontier research, and diverse enterprise workloads stay overwhelmingly Nvidia's.
SemiAnalysis (Dylan Patel); Gartner analysts Gaurav Gupta and Alvin Nguyen; Counterpoint Research (Gareth Owen); Moor Insights analyst Matt Kimball; Futurum Group analyst Brendan Burke
This is a genuinely contested, forward-looking forecast rather than a settled question, and notably all four camps agree on the same verified facts: Nvidia's absolute revenue is climbing steeply (data center revenue up 68 percent to $193.7 billion in FY2026) while custom ASIC shipments are growing roughly three times faster than merchant GPUs (44.6 percent versus 16.1 percent). They diverge only on interpretation. The strongest agreed-upon evidence is that the metric you choose decides who is right: by absolute revenue, frontier-training, and interconnect dominance, the Fortress and Realist cases are well supported today; by percentage share of inference and direction of travel, the Breakaway and Skeptic cases carry the clearest quantitative momentum (the 44.6 versus 16.1 percent CAGR, Broadcom up 106 percent year over year, and New Street's projection that Nvidia's inference share could fall to 20 to 30 percent by 2028). The honest synthesis is that Nvidia almost certainly keeps growing in absolute terms and keeps the training crown for now, while losing meaningful percentage share of inference, so the live disagreement is not whether share shifts but whether that shift compresses Nvidia's margins and pricing power enough to end the monopoly narrative. The Breakaway and Skeptic camps lean on independent shipment data and named-analyst risk warnings; the Realist and Fortress camps lean on the annual roadmap cadence and the still-partial adoption of ASICs inside hyperscaler walls, both well documented in the cited evidence.
Broadcom co-designs the custom chips for hyperscalers like Google and OpenAI, so its sharply rising AI revenue is the clearest proof that the move away from buying only Nvidia GPUs is real and large, not just talk.
TrendForce's 44.6 percent versus 16.1 percent split is the headline number behind this trend: it quantifies how fast hyperscaler in-house silicon is taking share from the GPU model that defined the early AI boom.
AI accelerators are useless without High Bandwidth Memory to feed them, so a multi-year HBM shortage shows the supercycle is now constrained by the whole supply chain, not just by how many chips can be designed.