YouTubeFeed

Dylan Patel — The single biggest bottleneck to scaling AI compute

Summary

Dwarkesh Patel interviews his roommate Dylan Patel, CEO of SemiAnalysis, for a deep technical dive into the three major bottlenecks constraining AI compute scaling: logic (semiconductor fabrication), memory, and power. Dylan walks through the entire supply chain from ASML’s EUV lithography machines down to the economics of individual GPU depreciation, providing specific numbers at every level. The big four hyperscalers (Amazon, Meta, Google, Microsoft) are on track to spend roughly $600 billion in combined CapEx, which translates to over a trillion dollars when you include the rest of the supply chain.

The central thesis is that by 2028-2029, the ultimate constraint on AI compute scaling falls to the lowest rung of the supply chain: ASML’s EUV lithography tools. ASML currently ships about 60 EUV machines per year and is ramping to 70-80, but each machine takes 18 months to manufacture and requires components from suppliers who are themselves capacity-constrained. Dylan reveals that ASML’s machines are so complex that a single EUV source fires 50,000 tin droplets per second, hitting each one three times with a laser at nine G’s of force, to produce the extreme ultraviolet light that patterns transistors. He argues that ASML has been conservative about capacity expansion because the semiconductor industry has historically punished companies that over-invested during booms.

The conversation covers the memory crunch in detail — HBM (high-bandwidth memory) demand is outstripping supply so badly that DRAM prices have tripled, adding $150 to iPhone costs and potentially causing Apple to lose its position as TSMC’s largest customer to AI chip makers. On power, Dylan is bullish: he argues the US can scale to 200+ gigawatts using diverse gas generation sources (combined cycle turbines, aero-derivative turbines, reciprocating engines, and more), with over 16 different power equipment manufacturers each capable of tens of gigawatts. Space GPUs, meanwhile, are dismissed as impractical this decade due to the same chip bottleneck plus networking constraints. The episode closes with discussion of Taiwan risk, China’s semiconductor ambitions, robots, and the economics of the “fast takeoff” scenario where AI revenue compounds fast enough to self-fund ever-larger compute investments.

Highlights

”The bottleneck falls to the lowest rung: ASML”

Clip

“Right. So to scale compute further, right, there’s some different bottlenecks this year, next year, but ultimately by 28, 29, the bottleneck falls to the lowest rung on the supply chain.” — Dylan Patel, 37:01

Clip command
yt-dlp --download-sections "*37:01-37:57" "https://www.youtube.com/watch?v=mDG_Hx3BSUE" --force-keyframes-at-cuts --merge-output-format mp4 -o "asml-bottleneck.mp4"

”50,000 tin droplets per second, each hit three times”

Clip

“And each of them are moving at nine G’s in opposite directions. So each of these things is like a wonder and it’s the most complicated machine humanity has ever built.” — Dylan Patel, 50:30

Clip command
yt-dlp --download-sections "*48:00-50:45" "https://www.youtube.com/watch?v=mDG_Hx3BSUE" --force-keyframes-at-cuts --merge-output-format mp4 -o "euv-source-tindrops.mp4"

”Any of these individually will do tens of gigawatts”

Clip

“Any of these individually will do tens of gigawatts and in a whole they will do hundreds of gigawatts.” — Dylan Patel, 1:50:32

Clip command
yt-dlp --download-sections "*110:28-111:10" "https://www.youtube.com/watch?v=mDG_Hx3BSUE" --force-keyframes-at-cuts --merge-output-format mp4 -o "hundreds-of-gigawatts.mp4"

”An H100 is worth more today than 3 years ago”

Clip

“So, when you talk about the Capex of these hyperscalers, right, on the order of $600 billion, and you look across the rest of the supply chain, it gets you to on the order of a trillion dollars.” — Dylan Patel, 1:43

Clip command
yt-dlp --download-sections "*1:43-3:00" "https://www.youtube.com/watch?v=mDG_Hx3BSUE" --force-keyframes-at-cuts --merge-output-format mp4 -o "trillion-dollar-supply-chain.mp4"

”Anthropic saw it before Google”

Clip

“Because Anthropic saw it before Google. And then Google had Nanabanana and Gemini 3, which caused their user metrics to skyrocket and leadership at Google was like oh.” — Dylan Patel, 32:07

Clip command
yt-dlp --download-sections "*32:07-33:00" "https://www.youtube.com/watch?v=mDG_Hx3BSUE" --force-keyframes-at-cuts --merge-output-format mp4 -o "anthropic-saw-it-first.mp4"

”Space GPUs aren’t happening this decade”

Clip

“So space data centers effectively are not limited by, you know, hey we have this energy advantage. It’s actually just limited by the same contended resource. We could only make 200 gigawatts worth of AI chips.” — Dylan Patel, 2:01:57

Clip command
yt-dlp --download-sections "*121:57-123:00" "https://www.youtube.com/watch?v=mDG_Hx3BSUE" --force-keyframes-at-cuts --merge-output-format mp4 -o "space-gpus-not-happening.mp4"

Key Points

  • Hyperscaler CapEx hits $600B, supply chain over $1T (1:43) - Combined spending by Amazon, Meta, Google, Microsoft drives a trillion-dollar ecosystem
  • Anthropic at 2.5 GW, needs 5 GW by year-end (3:00) - Anthropic’s revenue growth has outpaced compute acquisition plans
  • GPU value increases over time (12:00) - Unlike typical IT assets, H100s are worth more now than at launch because the models that run on them generate more revenue
  • GPT-5.4 is cheaper to run than GPT-4 (14:25) - Newer models have fewer active parameters but generate far more revenue per token, increasing GPU economics
  • Nvidia holds 70%+ of N3 wafer capacity by 2027 (25:00) - TSMC’s 3nm capacity is increasingly dominated by AI accelerators over mobile chips
  • Anthropic saw compute demand before Google (32:07) - Anthropic rushed to secure Google TPU capacity; Google’s own demand spiked after Gemini 3 success
  • ASML ships ~60 EUV tools/year (34:54) - Each takes 18 months to build; this is the ultimate bottleneck by 2028-29
  • EUV source: most complex machine ever built (48:00) - 50,000 tin droplets/second hit by laser at 9 G’s; machine is so large it ships on multiple planes
  • Supply chain not AGI-pilled enough (53:38) - Each level down the supply chain builds less than the level above demands; ASML most conservative
  • Old fabs could be repurposed for AI (55:43) - 7nm and older process nodes could theoretically be used for inference chips in a compute crunch
  • China could have working EUV by 2030 (1:07:31) - But manufacturing at volume is a separate challenge from having working tools
  • Memory crunch: DRAM prices tripled (1:24:08) - iPhone memory cost went from $50 to $150; HBM demand is cannibalizing consumer DRAM
  • Apple losing TSMC primacy to AI (2:21:00) - For the first time, Apple won’t be TSMC’s first N2 customer; AI chips take priority
  • 200+ GW achievable from gas generation alone (1:49:32) - 16+ manufacturers of power equipment each capable of tens of gigawatts
  • Space GPUs impractical this decade (2:01:57) - Bottlenecked by the same chip supply plus networking and cooling constraints
  • Fast takeoff via revenue compounding (1:13:59) - AI revenue could compound fast enough to self-fund larger compute investments without requiring AGI beliefs
  • Taiwan risk: destroying fabs would shrink global GDP (2:30:00) - Even with ASML tools elsewhere, losing Taiwan’s process engineers and institutional knowledge is catastrophic

Mentions

Companies

  • ASML (34:54) - Dutch lithography monopoly; ships ~60 EUV machines/year; ultimate bottleneck by 2028-29
  • TSMC (24:00) - World’s largest chip foundry; $100B CapEx over 3 years; conservative allocator favoring stable customers
  • Nvidia (25:00) - Will hold 70%+ of N3 capacity by 2027; launching new Samsung-fabbed robot chip
  • Anthropic (3:00) - At 2.5 GW compute, needs 5 GW; saw demand surge before Google; Claude Code reliability limited by compute
  • Google (31:54) - Sold 1M Ironwood TPUs to Anthropic; demand spiked after Gemini 3 success
  • Apple (26:07) - Losing position as TSMC’s top customer as AI chips take priority at N2
  • SemiAnalysis (0:00) - Dylan’s company; tracks every data center, wafer order, and supply chain bottleneck
  • Meta (22:51) - Adding as much compute capacity this year as their entire prior fleet
  • Samsung (2:28:15) - Fabricating Nvidia’s new robot chip in Texas
  • Micron (1:30:33) - Bought a fab in Taiwan to increase memory capacity during crunch
  • Huawei (2:22:26) - Would have eclipsed Apple as largest TSMC customer if not banned; first with 7nm AI chip
  • Crusoe (1:50:51) - Building 1.2 GW data center in Abilene for OpenAI with 5,000 workers at peak

Products & Technologies

  • EUV lithography (37:57) - Extreme ultraviolet light used to pattern transistors; $200M+ per machine
  • H100 (0:00) - Nvidia GPU worth more today than at launch due to increasing model revenue
  • HBM (High Bandwidth Memory) (1:16:02) - Critical memory for AI accelerators; supply severely constrained
  • Blackwell (1:03:27) - Nvidia’s latest GPU architecture; performance varies across on-chip vs cross-chip communication
  • GPT-5.4 (14:25) - Cheaper to run than GPT-4 with fewer active parameters but higher revenue per token
  • Claude Code (1:13:14) - Reliability limited by compute constraints at Anthropic

People

  • Jensen Huang (1:39:50) - Nvidia CEO; referenced asking TSMC to expand faster
  • Elon Musk (1:32:41) - Plans for Gigafab/Terafab; bullish on space GPUs; signed Samsung deal for robot chips
  • Michael Burry (12:00) - Argued GPU depreciation is 3 years or less; Dylan disagrees given increasing value
  • Sam Altman (53:24) - OpenAI CEO; knows they need X compute but supply chain builds less

Surprising Quotes

“The depreciation cycle of a GPU — Michael Burry was saying it’s three years or less. But an H100 is worth more today than when it was shipped three years ago.” — Dylan Patel, 12:00

“This year alone, Meta’s adding as much capacity as they had in the entire fleet before.” — Dylan Patel, 22:51

“Constantly we’re told our numbers are way too high, and then when they’re right, they’re like, ‘Oh, well, it was obvious.’” — Dylan Patel, 47:05

“You don’t have to believe in AGI to have the timelines where the US wins.” — Dylan Patel, 1:15:54

“If Huawei was not banned from using TSMC, Huawei would have already eclipsed Apple as the biggest TSMC customer.” — Dylan Patel, 2:24:00

Transcript

Dwarkesh Patel: 0:00 Alright, this is the episode of my roommate teaches me semiconductors.

Dylan Patel: 0:04 Yeah, and there’s going to be a lot of tangents. A lot of tangents.

Dwarkesh Patel: 0:14 Yes. Okay, Dylan is the CEO of SemiAnalysis. Dylan, the burning question I have for you — if you add up the big four: Amazon, Meta, Google, Microsoft, their combined forecasted CapEx is staggering. What’s the number?

Dylan Patel: 1:43 So, when you talk about the Capex of these hyperscalers, right, on the order of $600 billion, and you look across the rest of the supply chain, it gets you to on the order of a trillion dollars.

Dylan Patel: 3:00 Anthropic is at 2.5 gigawatts and 1.5 gigawatts roughly right now. They’re trying to scale to much larger. If you look at what Anthropic has done over the last few months — 4 billion dollar raise, the revenue has gone way crazier than expected.

Dwarkesh Patel: 4:03 Can I ask a question about that? If Anthropic was not on track to have 5 gigawatts by the end of this year, but it needs that to serve both the revenue and the training runs, what does it mean to acquire compute in a pinch?

Dylan Patel: 6:55 To acquire excess compute — there is capacity at hyperscalers, and not all contracts for compute are long-term. There is compute available for shorter-term commitments, but at a premium.

Dylan Patel: 12:00 The depreciation cycle of a GPU — Michael Burry was saying it’s three years or less. But there’s two ways to think about it. An H100 is generating more revenue per token today than when it first shipped, because the models running on it are better. The value of a GPU actually goes up over time as you approach more capable models.

Dylan Patel: 14:25 GPT-5.4 is both way cheaper to run than GPT-4, has fewer active parameters, it’s much smaller. And yet the revenue per token is dramatically higher because the use cases expand so much.

Dwarkesh Patel: 15:53 That’s crazy. If we had genuinely transformative AI models, the value of compute goes to infinity.

Dylan Patel: 19:42 There’s this interesting economics effect called Alchian-Allen, which is the idea that if you increase the fixed cost of different goods, one of which is higher quality, the higher quality one gets consumed more.

Dylan Patel: 22:51 Every year they’re adding incrementally way more capacity than they had previously. This year alone, Meta’s adding as much capacity as they had in the entire fleet before.

Dwarkesh Patel: 25:00 According to your numbers, by ‘27, Nvidia is going to have like 70-plus percent of N3 wafer capacity.

Dylan Patel: 26:07 On 3 nanometer, if we go back to last year, the vast majority was Apple. Apple’s revenue is not growing that fast. And here comes Nvidia, which is growing way faster.

Dylan Patel: 27:30 TSMC is conservative and doesn’t want to ride cycles of growth too hard. You actually want to allocate to the market that is more stable and lower growth rate first because you know it’s real.

Dylan Patel: 32:07 Because Anthropic saw it before Google. And then Google had Nanabanana and Gemini 3, which caused their user metrics to skyrocket and leadership at Google was like, oh.

Dwarkesh Patel: 34:39 Every year the bottleneck for what is preventing us from scaling AI compute keeps changing. A couple years ago it was CoWoS, last year it was power, this year you’ll say it’s something else.

Dylan Patel: 34:54 The biggest bottleneck is compute and for that the longest lead time supply chains are not power or data centers, they’re actually the semiconductor supply chain itself.

Dylan Patel: 37:01 Right. So to scale compute further, there’s some different bottlenecks this year, next year, but ultimately by 28, 29, the bottleneck falls to the lowest rung on the supply chain. ASML ships about 60 EUV tools a year. The problem is each tool takes 18 months to manufacture.

Dylan Patel: 39:00 You take the wafer, you deposit photoresist, which is a chemical that chemically changes when you expose it to light. And then ASML’s machine is the thing that shines light onto this layer at extraordinarily precise patterns.

Dwarkesh Patel: 41:08 Over the last three years, TSMC has done $100 billion of CapEx. How many EUV tools will there be by 2030?

Dylan Patel: 42:16 The entire ecosystem has something like 250 to 300 EUV tools already. And then you stack on 70 this year, 80 next year. By 2030 you might have 600-700 total.

Dylan Patel: 43:49 ASML’s been shipping EUV tools now for roughly a decade, but it only entered mass volume production around 2020. The tool’s not the same — back then it was half as productive.

Dylan Patel: 46:31 ASML has not decided to just go YOLO, let’s expand capacity as fast as possible. The semiconductor supply chain has historically punished companies that over-invested.

Dylan Patel: 47:05 Constantly we’re told our numbers are way too high, and then when they’re right, they’re like, ‘Oh, well, it was obvious.’

Dylan Patel: 48:00 What does the source do? It drops these tin droplets, it hits it three subsequent times with the laser. 50,000 droplets a second. Each of them are moving at nine G’s in opposite directions. Each of these things is like a wonder and it’s the most complicated machine humanity has ever built.

Dylan Patel: 51:00 So large that you’re building it in the factory in Eindhoven, Netherlands, and they’re deconstructing it and shipping it on many planes to the customer.

Dylan Patel: 53:38 You go down the supply chain, everyone’s doing minus one, and in some cases they’re doing divided by two. Because they just don’t believe the demand — they’re not AGI pilled.

Dwarkesh Patel: 55:43 People have been making arguments about specific bottlenecks. But what about using older fabs for inference?

Dylan Patel: 57:56 We potentially do go crazy enough that this happens because we just need incremental compute and the compute is worth the higher cost.

Dylan Patel: 1:06:00 To date, China still does not have an entire indigenous semiconductor supply chain. All of China’s 7nm and 14nm capacity uses ASML DUV tools.

Dylan Patel: 1:07:31 I think they’ll have working tools. I don’t think that they’ll be able to manufacture a bunch yet. There’s having it work and then there’s productionizing it.

Dylan Patel: 1:13:14 Given compute constraints are what’s bottlenecking their growth, the reliability of Claude Code is actually quite low because of compute limits.

Dylan Patel: 1:13:59 The margins are sub 50% at least last reported by The Information. So you’re at like 13, 14 billion dollars of compute that Anthropic is running. The revenue is compounding and self-funding larger compute investments.

Dylan Patel: 1:15:54 Like I don’t think you have to believe in AGI to have the timelines where the US wins.

Dwarkesh Patel: 1:16:02 Okay, let’s go back to memory because I think this is maybe people on Wall Street and people in the industry are understanding how big a deal this is, but generally people don’t.

Dylan Patel: 1:19:43 In many cases these GPUs are not running at full memory capacity. It’s a system design thing — model, hardware, software co-design.

Dylan Patel: 1:24:08 An iPhone has 12 gigabytes of memory. Each gig used to cost roughly three or four dollars, so 50 bucks. But now the price of memory is like triple. Let’s call it $150 per iPhone just for memory.

Dylan Patel: 1:30:33 Some really crazy stuff to get capacity — Micron bought a fab from a company in Taiwan that makes lagging edge chips. Hynix is converting old fabs.

Dwarkesh Patel: 1:42:30 Let me ask about power now. It sounds like you think power can be arbitrarily scaled.

Dylan Patel: 1:43:19 Right now we’re at 20 gigawatts of critical IT capacity. This is an important distinction. When I’m talking about these gigawatt numbers, it’s actual compute power, not total facility power.

Dylan Patel: 1:45:00 It’s not just turbines. There’s medium-speed reciprocating engines, there’s aero-derivative turbines, there’s all these different sources. Like 10 different companies make engines that can power data centers.

Dylan Patel: 1:49:32 We’re tracking over 16 different manufacturers of power-generating things just from gas alone.

Dylan Patel: 1:50:32 Any of these individually will do tens of gigawatts and in a whole they will do hundreds of gigawatts.

Dwarkesh Patel: 1:50:51 Right now in Abilene, the 1.2 gigawatt data center that Crusoe’s building for OpenAI, I think they had like 5,000 people working there at peak. If you turn that into 100 gigawatts, that’s like 400,000 construction workers.

Dylan Patel: 1:51:34 Labor is a humongous constraint in this. People have to be trained. We probably start importing the highest skilled labor.

Dylan Patel: 1:55:00 Permitting-wise, air pollution permits are a challenge, but the current administration has been friendlier to that.

Dylan Patel: 1:57:00 Australia, Malaysia, Indonesia, India — these are all places where data centers are going up at a much faster pace, but currently still 70%+ of AI compute is in the US.

Dylan Patel: 1:57:42 Obviously power’s free in space basically. That’s the reason to do it. But then there’s all the counter arguments.

Dylan Patel: 1:58:22 You’ve tested them all, deconstructed them, put them on a spaceship, put them into space, and then put them online again. That’s months of downtime.

Dylan Patel: 2:01:57 So space data centers effectively are not limited by the energy advantage. It’s actually just limited by the same contended resource. We could only make 200 gigawatts worth of AI chips. It doesn’t matter whether you put them on earth or space.

Dylan Patel: 2:05:18 You can’t run it hotter, you can only run it denser. And the problem is getting the heat out of that dense area means you have to move from standard air cooling to more exotic forms of liquid cooling or even immersion. That’s more difficult in space than on earth.

Dylan Patel: 2:19:07 What would really happen is Nvidia and all these others will say, we’re going to prepay for the capacity and you’re going to expand it for us. And that puts Apple in a tough spot.

Dylan Patel: 2:21:00 At 2 nanometer, actually the first time Apple is not the first customer. The first customer is AI. As we move to A16, the first customer there is not even Apple. It’s AI.

Dylan Patel: 2:22:26 Huawei was the first with the 7nm AI chip. If Huawei was not banned from using TSMC, Huawei would have already eclipsed Apple as the biggest TSMC customer.

Dylan Patel: 2:24:33 There’s a lot of difficulties with VLMs and VLAs that people are deploying on robots. But to some extent you don’t need all the intelligence on the edge — a lot can happen in the data center.

Dylan Patel: 2:27:27 I think Elon recognizes this, which is why he’s going to different places for his chips. He signed this massive deal with Samsung to make his robot chips in Texas.

Dwarkesh Patel: 2:28:34 Final question on Taiwan. If we believe that tools are the ultimate bottleneck, how much of Taiwan’s place could we de-risk by having ASML tools elsewhere?

Dylan Patel: 2:29:00 If you ship out all the process engineers and assuming it’s hot enough to destroy the fabs — no one has all the fabs in Taiwan now, which is a big risk. You’ve drastically slowed US and global GDP. Not just growth — you’ve shrunk GDP massively.

Dylan Patel: 2:30:44 And you’ve got a lot bigger problems, and your incremental AI compute is the least of your concerns at that point.