YouTubeFeed
← All-In Podcast

Open Source Wins, AGI Is Here, and Scorsese's AI Toolkit with CEOs of Cerebras & Black Forest Labs

63:57 110.7K views 2026-07-10 Watch on YouTube ↗

Open Source Wins, AGI Is Here, and Scorsese’s AI Toolkit with CEOs of Cerebras & Black Forest Labs

Summary

Recorded in Versailles, this episode pairs two builder interviews on the physical and creative frontiers of AI. Andrew Feldman, CEO and founder of Cerebras, describes an infrastructure build-out with no modern precedent — data centers the size of football fields drawing more power than mid-sized cities, being erected everywhere from Texas to Kazakhstan. Cerebras is sitting on a $25 billion backlog, and Feldman argues the whole industry is chasing already-booked demand rather than speculating. He frames the shift from prompt-whispering to intent-understanding “reasoning models” as the leap that makes fast inference the center of gravity, and claims his architecture is now breaking Moore’s Law, doubling performance well past 2X every 18 months.

The conversation with Feldman ranges across open source and sovereignty (Cerebras runs GLM, Kimi, Qwen, OpenAI’s models, and sovereign models for G42/MBZUAI and GlaxoSmithKline), the reasonableness of staged model rollouts and government red-teaming, and the recursive “loop maxing” dynamics driving the road to superintelligence. The Besties land on a shared claim — that by any definition from 20 years ago, AGI has already arrived — and Feldman makes an optimistic case for abundance, personalized AI tutors, and a shot at a world where no one’s children die of cancer.

In the second half, Robin Rombach, co-founder and CEO of Black Forest Labs, traces generative media from 64x64-pixel images to multi-minute video and, ultimately, to “physical AI.” A pioneer of latent diffusion (the foundation of Stable Diffusion and his company’s Flux models), Rombach frames AI as a collaborative medium, sharing how Martin Scorsese used the tools to pull a mental picture of an Eastern European village out of his head. He discusses AI-generated sets replacing green screens in film, fan-film creativity and IP licensing for holders like Disney, and the convergence of world models, action prediction, and robotics into a single multimodal system.

Highlights

”More power than the previous 50 years on Earth took”

Data centers bigger than cities

“What we’re talking about now are data centers that are in the next several years going to use more power than the previous 50 years on Earth took.” — Andrew Feldman, 2:17

Clip command
yt-dlp --download-sections "*2:17-3:00" "https://www.youtube.com/watch?v=Y7p4rUCdqi0" --force-keyframes-at-cuts --merge-output-format mp4 -o "Y7p4rUCdqi0-2m17s.mp4"

”We have a 25 billion dollar backlog”

Chasing yesterday's demand

“The irony is unlike many sort of exciting times in technology, they’re trying to capture yesterday’s demand, right? The demand is way outstripping our ability to build data centers and to fill them with hardware. And so, you know, we have a 25 billion dollar backlog.” — Andrew Feldman, 3:42

Clip command
yt-dlp --download-sections "*3:42-4:25" "https://www.youtube.com/watch?v=Y7p4rUCdqi0" --force-keyframes-at-cuts --merge-output-format mp4 -o "Y7p4rUCdqi0-3m42s.mp4"

Breaking Moore’s Law

Way over 2X

“And we crushed it with this chip and we’ve carved out a whole new trajectory. And my view is in the next 18 months we’ll be way over 2X.” — Andrew Feldman, 13:10

Clip command
yt-dlp --download-sections "*13:10-13:41" "https://www.youtube.com/watch?v=Y7p4rUCdqi0" --force-keyframes-at-cuts --merge-output-format mp4 -o "Y7p4rUCdqi0-13m10s.mp4"

”By any definition we had 20 years ago, we’ve hit it”

AGI is here

“By any definition we had 20 years ago, we’ve hit it.” — David Friedberg, 29:49

Clip command
yt-dlp --download-sections "*29:49-30:25" "https://www.youtube.com/watch?v=Y7p4rUCdqi0" --force-keyframes-at-cuts --merge-output-format mp4 -o "Y7p4rUCdqi0-29m49s.mp4"

A shot at no one dying of cancer

The pro side of the ledger

“We have a shot with this technology so that not our children, nor anyone they know dies of cancer.” — David Friedberg, 38:30

Clip command
yt-dlp --download-sections "*38:30-39:06" "https://www.youtube.com/watch?v=Y7p4rUCdqi0" --force-keyframes-at-cuts --merge-output-format mp4 -o "Y7p4rUCdqi0-38m30s.mp4"

Scorsese pulling a mental picture out of his head

Working with Martin Scorsese

“Getting the mental picture of something out of your head and communicating it in a visual way by making these images or the series of images is something that just makes it easier to communicate and convey an idea of what is actually in your head.” — Robin Rombach, 48:26

Clip command
yt-dlp --download-sections "*48:26-49:25" "https://www.youtube.com/watch?v=Y7p4rUCdqi0" --force-keyframes-at-cuts --merge-output-format mp4 -o "Y7p4rUCdqi0-48m26s.mp4"

Key Points

  • The unprecedented scale of the build-out (2:04) - Feldman on the physical enormity of AI data centers, unlike chips or boxes.
  • Data centers drawing more power than cities (2:28) - Football-field-sized buildings being built across the US, Nordics, Middle East, Kazakhstan, and beyond.
  • Chasing booked demand, not speculation (3:42) - A $25 billion backlog; demand is outstripping the ability to build.
  • Token maxing vs. real value (4:32) - Massive experimentation alongside enormous net value, likened to early AWS and shopping Costco.
  • A new class of systems thinker (6:19) - Chamath on how the tools push users to define goals and write requirements documents.
  • From prompt-whispering to understanding intent (7:37) - Feldman on how models now infer what you meant, a huge leap in 24 months.
  • Watching a reasoning model debate itself (9:00) - Friedberg’s trend-hunting job on GLM-5.2 via a BitTensor project with unlimited capacity.
  • Unlimited tokens means unlimited reasoning (10:38) - Run a model for 24-48 hours (15x faster on Cerebras) and get weeks or months of thinking.
  • Reasoning is inference (11:47) - Why fast compute makes reasoning tractable rather than crippling it.
  • Breaking Moore’s Law (13:02) - New architectures have room to beat the 18-month doubling; old GPU architectures don’t.
  • Why hyperscalers build their own chips (15:22) - Nobody likes being dependent; controlling your own destiny after the Intel-dependence lesson.
  • The Ferrari vs. minivan model choice (16:43) - Frontier models for hard problems, rock-solid open source for ordinary G&A work.
  • Sovereignty as a trend (18:00) - Regulated industries want on-prem, domestic, open-source models they can control.
  • The need for more domestic open-source models (18:38) - Right now the world’s choice is OSS 120B or Chinese models.
  • Staged rollouts and government red-teaming (21:26) - Feldman on why phasing a dangerous model, like powerful pharmaceuticals, is reasonable.
  • Palo Alto Networks found unknown bugs (25:11) - Nikesh’s team had to stop everything and patch for six weeks.
  • A massive data breach is inevitable (26:33) - Chamath’s reinsurance analogy: prepare for the black swan you can’t predict specifically.
  • AGI has already been hit (29:29) - Turing tests blown away; any prior definition surpassed.
  • Recursive loops and exponential gains (31:53) - Feldman on “loop maxing,” steep curves, and not knowing where it ends.
  • Shortening the inter-generational gap (36:00) - Human learning moves at generational pace; AI iterates like fruit flies.
  • The ledger weighted toward abundance (38:33) - Sacks on dislocation vs. cures; Feldman on personalized AI tutors for every child.
  • Latent diffusion, the foundation of it all (41:47) - Rombach on compressing images/video/audio into efficient representations for a Transformer.
  • Video training gives implicit physics (43:51) - Intuitive vs. deep-reasoning intelligence converging into a multimodal model with action prediction.
  • AI as a medium, not a replacement (47:41) - Rombach won’t tell Scorsese how to use the model; human-in-the-loop yields the best output.
  • Generative sets replacing green screens (52:23) - A $30M Bitcoin movie that would have cost $150M with physical sets and never been greenlit.
  • World models, action prediction, and robots (54:09) - The same model can make a movie and serve as a brain on a robot.
  • Fan films and IP licensing (1:00:23) - Empowering fans to make their own Star Wars stories under a licensing model.

Mentions

Companies

  • Cerebras (0:00) - Andrew Feldman’s inference-chip company; $25B backlog, breaking Moore’s Law.
  • Black Forest Labs (40:53) - Robin Rombach’s open-source image/video model company, based in Freiburg and San Francisco.
  • OpenAI (3:57) - Named as an insatiable data-center customer; released OSS 120B open-source model.
  • Anthropic (3:57) - Cited as a frontier-model player; Dario’s staged-release decision discussed.
  • Google / Gemini (3:57) - Named among hyperscalers demanding more data centers.
  • Microsoft (3:57) - Wants more data centers.
  • AWS / Amazon (4:51) - The early-cloud experimentation analogy; making their own chips.
  • NVIDIA (15:00) - Jensen’s dilemma over pushing open-source models that compete with customers.
  • Intel (15:22) - The x86 dependence lesson for hyperscalers.
  • Palo Alto Networks (25:11) - Nikesh’s firm found unknown bugs and patched for six weeks.
  • GlaxoSmithKline (19:40) - Runs its own models on Cerebras.
  • G42 / MBZUAI (19:40) - UAE partners running sovereign models on Cerebras.
  • Stability AI (Stable Diffusion) (41:11) - Rombach’s prior work before Black Forest Labs.
  • Disney (58:03) - Major IP holder discussed as a potential model-training partner.
  • Perplexity (28:14) - Automatically suggests next prompts.
  • AppLovin (1:13) - Sponsor.
  • Nasdaq (40:22) - Sponsor.

Products & Technologies

  • Cerebras wafer-scale chip (12:00) - The half-billion-dollar-to-make machine Feldman brought to the interview.
  • Reasoning / inference models (11:47) - Compute-intensive reasoning as the new center of gravity.
  • GLM-5.2 / Z-AI (8:53) - The model Friedberg ran his trend-hunting job on.
  • Hermes agent (8:30) - Agent Chamath asked about.
  • BitTensor (8:54) - Distributed crypto project supplying extra capacity.
  • Kimi (16:08) - Open-source model Jason smart-routes to; runs on Cerebras.
  • Qwen (19:40) - Open-source model family run on Cerebras.
  • OSS 120B (18:38) - OpenAI’s open-source model.
  • Claude / ChatGPT (16:08) - Referenced for token consumption and reasoning generations.
  • Latent diffusion (41:47) - The algorithm Rombach invented that underpins generative image/video/physical AI.
  • Flux (41:11) - Black Forest Labs’ open-source image model.
  • Sora (58:03) - OpenAI’s video product, referenced re: Disney IP licensing.
  • Optimus (34:35) - Robots invoked in the “build me the Palace of Versailles” thought experiment.
  • World / action models (54:09) - Multimodal models for movies, robotics, and computer use.

People

  • Andrew Feldman (0:00) - CEO and founder of Cerebras.
  • Robin Rombach (40:53) - Co-founder and CEO of Black Forest Labs.
  • Martin Scorsese (46:29) - Explored Black Forest Labs’ models to visualize scenes.
  • Sam Altman (10:40) - Saw reasoning coming early; described it on All-In.
  • Ilya Sutskever (10:40) - Cited for early safety warnings and seeing around corners.
  • Dario Amodei (20:35) - Anthropic CEO; staged-release and partisanship discussion.
  • Demis Hassabis (31:53) - Named among those who saw recursive gains early.
  • Elon Musk (30:08) - Cited for driving launch-vehicle costs toward zero.
  • Nikesh Arora (25:11) - Palo Alto Networks CEO who tested a frontier model against his software.
  • Warren Buffett (26:33) - Reinsurance analogy for inevitable breaches.
  • Gal Gadot (52:23) - Described shooting a Bitcoin movie with AI-generated sets.
  • Ridley Scott, Spielberg, George Lucas (50:40) - Directors known for storyboarding, used as an analogy.
  • Thomas Kuhn (37:12) - “Paradigms don’t die, people do.”

Surprising Quotes

“I’ve never far without, you know, when one costs half a billion to make, you bring it everywhere with you.” — Andrew Feldman, 12:08

“The AI starts telling people you’re token maxing and you need to get a little more focused here.” — Chamath Palihapitiya, 6:30

“It opened my eyes just this morning of what a world of unlimited tokens might look like. Because unlimited tokens, I believe, means unlimited reasoning.” — David Friedberg, 9:38

“There’s a shot that our children, none of them nor their people they love will die of cancer.” — David Sacks, 38:33

“The only thing that you could do was like images of 64x64 pixels. Now you can do like multi-minute videos at a high resolution, but it’s not going to stop there.” — Robin Rombach, 53:18

Transcript

Jason Calacanis: 0:00 We are in the race for super intelligence and Andrew Feldman is back and obviously CEO and founder of Cerebras, doing inference chips, pioneered the space, had a successful IPO. We’ve talked about this a couple times. We got to see each other in January at Davos, IPO happens. The boys and I got to sit with you recently at Liquidity.

Andrew Feldman: 0:22 That was fun.

Jason Calacanis: 0:23 That was really fun. Had a great discussion with the boys, but I wanted to deep dive with you about a couple topics. The first one is the build out of AI. We’ve never seen a build out like this since, you know, the Great Wall of China, the pyramids. I mean, it feels like the amount of capital, time and intelligent people on the planet dedicating themselves to the build out of something, I can’t think of anything in our lifetimes, with perhaps before our lifetimes, the war effort. This is a mobilization and a scale that we read about, we hear about, but you’re actually doing it. You have customers who are building data centers and you’re a key piece of that.

Jason Calacanis: 1:51 Maybe you could just enlighten us in 2026. What is Cerebras doing and what is happening with this build out out in Texas? These are some gigantic, gigantic efforts.

Andrew Feldman: 2:04 The size and scope of what is being built, the physical size and scope. Usually when we talk about software or we talk about hardware, we talk about chips or boxes and they don’t have the same sort of physical enormity.

Jason Calacanis: 2:16 Right.

Andrew Feldman: 2:17 And what we’re talking about now are data centers that are in the next several years going to use more power than the previous 50 years on Earth took.

Jason Calacanis: 2:27 Wow.

Andrew Feldman: 2:28 Right. We’re talking about individual buildings the size of football fields that have more power coming into them than mid-sized cities. And they’re being built, they’re being built across the US. They’re being built in Canada. They’re being built throughout the Nordics. They’re being built here in Paris and throughout France and Europe, in the Middle East, in nations that sort of weren’t front and center in anybody’s mind previously. You know, Kazakhstan, Tajikistan are building out, Georgia building out data centers of size, Armenia. Everybody’s sort of…

Jason Calacanis: 3:00 Focus. Every country, and every state obviously in America, feels they need to participate in this and the people who are buying the capacity, the OpenAIs, Anthropics, SpaceX AI, the Googles, they are insatiable right now.

Andrew Feldman: 3:20 Yeah.

Jason Calacanis: 3:22 And they’re building how many years out? When you talk to them, they were ordering chips from Cerebras before you were finished with the chip. They’re putting orders in ahead of time.

Andrew Feldman: 3:42 The irony is unlike many sort of exciting times in technology, they’re trying to capture yesterday’s demand, right? The demand is way outstripping our ability to build data centers and to fill them with hardware. All right, and so, you know, we have a 25 billion dollar backlog.

Jason Calacanis: 3:54 25 billion dollar backlog?

Andrew Feldman: 3:57 And we are not alone in that. The OpenAI, uh-huh, Anthropic, you go through this list of Google wants more data centers, Microsoft wants more data centers, AWS wants more data centers, right? All of these players are not chasing sort of if you build it they will come, they’re chasing the demand is booked. Right? How do we keep them from leaving? Right? And that’s extremely unusual.

Jason Calacanis: 4:25 It’s very unusual. And now we have people who are, you know, we have a term for it, token maxing.

Andrew Feldman: 4:31 Yeah.

Jason Calacanis: 4:32 And there’s a great debate, is this actually creating value? I’m curious where you stand, you know, is it even possible that this much demand could be created if value did not exist? There is clearly massive value happening.

Andrew Feldman: 4:47 Yeah.

Jason Calacanis: 4:48 But there’s also massive experimentation.

Andrew Feldman: 4:51 Oh, for sure. You know, what I, I liken this to when we first started with AWS and it was so good to get around your own IT organization, right, that you told every engineer had go ahead, put on your credit card, sign up. Right? And a lot of it was really useful and some of it was like, God, I wish we didn’t do that.

Jason Calacanis: 5:07 Yeah.

Andrew Feldman: 5:08 And so for sure there’s experimentation. But it doesn’t mean that the net value isn’t enormous. It means some of it is going to go nowhere. And you know, it was the same, I remember when Costco opened up in, in, in the Palo Alto area in 1988 and people used to shop Costco like they shop Safeway. They’d go down every aisle.

Jason Calacanis: 5:30 Yeah.

Andrew Feldman: 5:31 And that’s a horrible way to shop Costco, because you end up with four things you didn’t need and each was 22 dollars, right? And as people got more sort of accustomed to it, you go to the back to get the chicken, 18 cupcakes for the kids’ birthday party, bang, you were out, strategic. And it’s exactly the same. I think at first people opened up and said everybody as much tokens as you want and in enterprises there’s no open loop.

Jason Calacanis: 6:00 Don’t give sort of any resource unconstrained to people and now we’re jumping on and saying, whoa, alright, these guys should have as much as they need, they’re enormously productive. Over here we can use maybe an open source model, maybe a cheaper model over here, and now we’re sort of running like a business.

Chamath Palihapitiya: 6:19 And we’re really seeing a certain type of person emerge who knows how to deploy this technology, systems thinking,

Jason Calacanis: 6:29 Yeah.

Chamath Palihapitiya: 6:30 which developers kind of have innately. CEOs tend to be great strategists and understand systems, but the intelligence is getting so much better every step along the way that I’m watching individuals, typically startup founders but also venture capitalists and associates who work at my venture fund, they start playing with the tool and then the tool starts playing with them. They start to go, oh, I haven’t clearly defined what my goal was. I don’t understand what a system is. I’ve never heard about making a requirements document. And the software’s like, do you have a requirements document? What’s your goal? The AI starts telling people you’re token maxing and you need to get a little more focused here.

Andrew Feldman: 7:13 One of my colleagues 20 years ago, a really smart, smart computer scientist said, computers are really dumb. They do exactly what you tell them.

Chamath Palihapitiya: 7:24 Yeah.

Andrew Feldman: 7:25 And at first, prompting was like that, right? You modified your prompt a little bit and it changed the answer

Chamath Palihapitiya: 7:31 Dramatically.

Andrew Feldman: 7:32 dramatically. And increasingly, it’s understanding what your intent was.

Chamath Palihapitiya: 7:35 Right, right.

Andrew Feldman: 7:37 And if you have a chance to play with Fable or five six from OpenAI, increasingly what you don’t have to get the prompt just right. You don’t have to be a prompt whisperer. Instead, you ask it and it says, well here’s some things and by the way, maybe you want the chart to go two ways, you want it a line and a bar. And it’s like, well that’s exactly what I wanted. I didn’t ask for it, but that is better. And so it’s understanding intent and that’s a huge leap.

Chamath Palihapitiya: 8:04 Which,

Andrew Feldman: 8:05 if we were sitting here two years ago, the idea,

Chamath Palihapitiya: 8:10 we would never have been able to predict in a short 24 months that it would go from being a great summarizer researcher of web results to actually understanding your intent and then providing a solution and abstracting it all from you.

Andrew Feldman: 8:29 That’s right.

Chamath Palihapitiya: 8:30 Which is a very weird thing. I don’t know if you’ve played with the Hermes agent yet?

Andrew Feldman: 8:35 Yeah.

Chamath Palihapitiya: 8:36 Have you played with it yet? I mean, I asked it just this morning and I was given a secret BitTensor project that has the new Z-AI’s model five two and they gave me

Andrew Feldman: 8:53 GLM five two.

Chamath Palihapitiya: 8:54 GLM five two. So somebody in that BitTensor, I think you understand BitTensor, you’ve heard of it, the distributed

David Friedberg: 9:00 crypto project and so they have all this extra capacity. I was— a whisper told me probably some capacity in China that has free energy. Okay, fine. So they gave me unlimited capacity, so I started having it do some really crazy jobs where I was saying, like, every hour I want you to tell me what the trends in the world are that nobody else has identified yet. And you can do whatever you want to do that, but my goal is to be the smartest trend hunter in the world. And I watched what it was doing in the background, and it started debating itself on where it should find things. It said, well, we should probably go to Hacker News and Reddit. And then it was like, yeah, but there’s also social media, and trends tend to manifest on Instagram.

Jason Calacanis: 9:32 That’s a reasoning model. You were watching a reasoning model work out. Isn’t that interesting? I mean, that’s amazing.

David Friedberg: 9:38 Yeah. And it was collapsed, so as a civilian who doesn’t hit the uncollapsed moment, and if you were using ChatGPT 3.5 or you were using 4.0, whatever it was, and you haven’t used this new level of reasoning and inference and unlimited compute essentially, it opened my eyes just this morning of what a world of unlimited tokens might look like. Right. Because unlimited tokens, I believe, means unlimited reasoning.

Andrew Feldman: 10:17 It does.

David Friedberg: 10:18 What does that mean?

Andrew Feldman: 10:21 Yeah, it’s— I mean, you run these for 24 or 48 hours, you get amazing things now. And what if by using Cerebras you were 15 times faster and then you ran it for 24 hours?

David Friedberg: 10:31 Right.

Andrew Feldman: 10:32 Right.

David Friedberg: 10:34 And you got weeks or months worth of thinking.

Andrew Feldman: 10:38 Yeah, and I mean, it is— it is extraordinary.

David Friedberg: 10:40 And I think one of the things is people like Ilya and Sam in the early days were saying this was coming. Right? Right. And I think when you look back, you say to yourself, holy crap, those guys— they knew.

Andrew Feldman: 10:53 Yeah, they could see around the corner.

David Friedberg: 10:54 That’s right. And the rest of us were like, what? I’m not sure.

Andrew Feldman: 11:00 Yeah.

Jason Calacanis: 11:01 When we had Sam on All-In at one point, he— and he said, oh, you know, I’d love to come on at some point. I said, sure, come on. And he was talking about it. He said— I said, what’s next? He said, reasoning. I said, unpack that. What does it mean? He’s like, well, understanding what your intent was, just as you’re saying, and then figuring out a strategy and then maybe talking to other agents and other threads about, like, is this the right thing to do and vetting each other’s work. And I’m like, wow, we have come a long way from guess the next word.

David Friedberg: 11:37 Right, right. Fill the sentence in, you know. Summarize this PDF.

Jason Calacanis: 11:44 Now Cerebras is at the center of this because this reasoning is inference.

Andrew Feldman: 11:47 This reasoning is inference, and it’s computationally intensive. Right? And so fast compute makes this sort of work fast and sort of tractable. It doesn’t cripple it by taking a huge amount of time to get a good answer. And so it’s exactly the fact that this reasoning is inference that makes it so exciting. reasoning, consumes a huge amount of tokens internally, that allows a blisteringly fast machine like ours. And I brought one because—

Jason Calacanis: 12:07 Oh, you got it?

Andrew Feldman: 12:08 I’ve never far without, you know, when one costs half a billion to make, you bring it everywhere with you.

Jason Calacanis: 12:14 And we were—we were tossing this back and forth at Davos. What’s the model number of this one?

Andrew Feldman: 12:19 This was in the first eight or ten.

Jason Calacanis: 12:21 Got it. So this has a special place.

Andrew Feldman: 12:24 This has a special place. I mean, my wife says it’s like I’m a kid with a dirt bike for his eighth birthday, he’s in his bedroom at night, I carry him with me.

Jason Calacanis: 12:32 I mean, when you have, you know, your next party at the house, I highly recommend just a little hors d’oeuvre something.

Andrew Feldman: 12:40 A little hors d’oeuvre.

Jason Calacanis: 12:41 I think it would be like a great bit. It would be a great bit if you had some.

Andrew Feldman: 12:44 That’s right.

David Friedberg: 12:45 But what we’re looking at here is the ability to do that reasoning at scale. And what is Moore’s Law for inference and for Cerebras? Do you have something internally you discuss as every—we’re going to double this every X time period?

Andrew Feldman: 13:02 So all chips prior to us in the processor world followed Moore’s Law.

David Friedberg: 13:06 Got it.

Andrew Feldman: 13:07 Doubling every 18 months.

David Friedberg: 13:09 Got it.

Andrew Feldman: 13:10 And we crushed it with this chip and we’ve carved out a whole new trajectory. And my view is in the next 18 months we’ll be way over 2X.

Jason Calacanis: 13:24 Interesting.

Andrew Feldman: 13:25 And so I think that early in an architecture, you have room to do much better than what was traditionally Moore’s Law. Now if you’ve got a 20-year-old architecture like the GPU, it’s much harder.

Jason Calacanis: 13:41 Right.

Andrew Feldman: 13:42 You have to rely on things like smaller geometry, right, going to the next fab node. But in a newer architecture, you have a huge amount of room still to learn about the work that is being presented and make optimizations that give you huge gains.

Jason Calacanis: 14:03 How do you run the company looking at just being the CEO now in the age of AI? You have $25 billion in demand, you have to—you have to deploy at an just an incredible blistering pace, you have to hire people, you have to create a roadmap. I don’t mean to give you a panic attack here.

Andrew Feldman: 14:26 Yeah.

Jason Calacanis: 14:27 You have to keep up with somebody like OpenAI who’s moving so unbelievably quickly, right? And they’re competitive. You’ve got to keep up, right?

Andrew Feldman: 14:37 Yes.

Jason Calacanis: 14:38 Your hardware, your software, your deployments have to keep up with some of the fastest moving organizations in history. They’re demanding customers.

Andrew Feldman: 14:49 They’re not—they’re not pushovers, for sure.

Jason Calacanis: 14:51 Yeah. And also potentially competitors down the road.

Andrew Feldman: 14:55 Look, I think there is so much demand right now that—that there is no silicon that will go unused. Right.

Jason Calacanis: 15:00 But why— Is OpenAI releasing Jalapeño? Why is Amazon making their own chips? You see this reoccurring trend. Is it a way to let you know, to let Jensen and Nvidia know, ‘Hey, we can do this too, so we need good pricing’? Is it a little bit of a flex that way, or is that the future, that they’re going to be in your business?

Andrew Feldman: 15:22 No, I think nobody likes being dependent. And I think some of the lessons learned by the hyperscalers of the x86 world is they were dependent on Intel. And some of the lessons learned by the GPU makers was they were dependent on a small number of hyperscalers.

Jason Calacanis: 15:46 Yeah.

Andrew Feldman: 15:47 And they wanted more customers, and so they set about to help fund these neo-clouds.

Jason Calacanis: 15:52 Yeah.

Andrew Feldman: 15:53 And so I think mostly it’s about an opportunity to control at least an important part of your destiny. And I think that’s a very reasonable thing. I think you don’t have to sort of make the fastest chip, you just can’t be entirely dependent on other people’s chips.

Jason Calacanis: 16:08 Got it. And that dependency has become a hot topic. I’m not sure if you caught the episodes over the last two weeks, but we’ve been talking over the last year about open source. I’ve been championing that a lot just because I was early into open source and quickly started using Kimi and was like, ‘Wait a second, I’m blowing out my Claude tokens, but this Kimi, I can’t tell the difference.’ And then we started smart-routing it, and suddenly this open source started to figure out reasoning and the gap has suddenly closed this year.

Andrew Feldman: 16:43 Well, I, you know, you don’t want to take your Ferrari to the grocery store, right? There are times you want to drive your fun car, right?

Jason Calacanis: 16:54 Yeah.

Andrew Feldman: 16:54 And there are times you want to throw the kids in and don’t worry if there are Cheerios on the floor.

Jason Calacanis: 17:02 Minivan time.

Andrew Feldman: 17:03 Right, there’s minivan time. And I think that as the sort of sophistication of the user grows, right, you’re going to have hard problems, and those are going to be frontier model problems. They’re going to be OpenAI problems, they’re going to be Anthropic problems, they’re going to be Gemini problems. And behind that, they’re going to be a lot of ordinary problems, right? I mean, if you think about a company, you know how much time is spent cutting things out of workday and getting it in a different cell for… yeah…

Jason Calacanis: 17:26 Right, the cutting and pasting economy is real.

Andrew Feldman: 17:30 That’s right. And this doesn’t need, right, gold-medal math. What this needs is sort of rock-solid open-source capability. And if you think about what… I mean, we’ve been thinking a lot about it in G&A, but a huge amount of G&A, alright, is not invention, right? And you may not need sort of the most sophisticated agents for this. And another card that…

Jason Calacanis: 18:00 turned over recently is some folks uh maybe have concerns with the ambition of the frontier models and maybe sharing their data data leakage and sovereignty of intelligence and they’re saying hey our company is going to choose maybe we’re in a regulated industry finance healthcare HIPAA you know FINRA all kinds of different regulations we need to have this on prem and we want to have it on prem domestically and we’d like an open source version where we have a little bit more control and I think are you seeing that now?

Andrew Feldman: 18:38 We are seeing that for sure and I and I think OpenAI made a good call releasing OSS 120B some months back that was a good open source model but I think in the US we need more domestic open source models we need to give the world a choice right if they want to run open source right now it’s OSS 120B or Chinese models.

Jason Calacanis: 18:56 NVIDIA has some.

Andrew Feldman: 18:58 NVIDIA has seen the same opportunity to push open source models I I think giving them more power might might be sort of something

Jason Calacanis: 19:13 Well I was about to that was you you cut me off at the pass like my understanding was Jensen was like hey we we don’t even want to talk about these open source models we have because our customers we’re now going to be competing with Sam Dario Elon uh Sergey like do we want to be in that position?

Andrew Feldman: 19:31 Right.

Jason Calacanis: 19:32 So but we do need some more champions here and it’s open source so people can fork it uh but that puts you in a more neutral position.

Andrew Feldman: 19:40 That’s right we we run today we run GLM we run Kimi we run the Qwen set of models and we run OpenAI’s models the closed source ones we run models for say GlaxoSmithKline which they wrote and developed um we run models for uh our partner in the UAE G42 and MBZUAI um that are are their models that they designed so we have a a a wide variety.

Jason Calacanis: 20:08 So sovereignty is a trend.

Andrew Feldman: 20:11 Sovereignty is a trend and I think uh the government’s actions with regard to Fable and uh five six um where they said oh whoa let’s think and then we can act um I I think particularly here in Europe was a bit of a wake-up call.

Jason Calacanis: 20:29 And when you saw this going down there’s a layer of partisanship in our country right now it’s pretty fervent Dario is pretty explicitly you know not part of this administration they’ve been very adversarial both sides have have admitted that they’re starting to work it out now so it’s hard I think for us not being in the room with these parties to understand what’s partisanship what’s gamesmanship here but do you believe that what they released was truly dangerous for cyber warfare for cyberattacks and that if you were to rate Dario’s, not communication because he’s a very effervescent communicator, I think is a diplomatic way to say it, but to have a scheduled rolled out release, right, we’ll put aside the government’s control of it, but do you think that is a wise thing for us to do at this point and do you think there was actually a major threat there?

Andrew Feldman: 21:26 So what’s interesting is I hadn’t seen it before. Right, and I think let’s if we just step back and say is it reasonable? I don’t know whether this was the right time, but at a time that a model is sufficiently creative in its thinking that it poses a meaningful threat for the government to say we’d like you to roll it out in steps.

Jason Calacanis: 21:54 Yeah.

Andrew Feldman: 21:55 Right? This doesn’t seem unreasonable to me. Not at all. Right? I mean we do this with powerful pharmaceuticals, right? We’d like, I mean we’re certainly not encouraging seven years of trial and the amount of paperwork and all the garbage that has accrued to the FDA. But with a powerful new technology it certainly doesn’t seem unreasonable to say ‘Hey guys, let’s at least do some red teaming at the government so we know our defenses can block this.’

Jason Calacanis: 22:24 Yeah, have we checked the infrastructure of the country like of the NSA? Have we checked the infrastructure of, right, and can you give us two or three weeks to patch any obvious holes that are found?

Andrew Feldman: 22:36 This doesn’t seem to me an unreasonable thing for the government to ask.

Jason Calacanis: 22:41 Right. We’ve, but we in this very polarized time put on top of it, well oh my god it’s President Trump doing it and then you have to think, well what if it was President AOC or President anybody in between the two extremes.

Andrew Feldman: 22:54 Right. I think the polarization hurts a great deal. It hurts clear thinking, right? It hurts clear thinking and and both sides are going to do some dumb things and some really smart things. Right? And in fact what I’ve found is the people in the government are trying really hard.

Jason Calacanis: 23:13 The rank and file.

Andrew Feldman: 23:14 The rank and file are trying really hard and this is moving fast. And I think that that an ability to to set aside some of the polarization and say ‘how do we do this in a reasonable manner?’ I mean we want Dario and Sam competing like crazy.

Jason Calacanis: 23:32 100%. It’s been awesome to watch.

Andrew Feldman: 23:34 It’s awesome. Yeah, it’s good for the technology, it’s good for it’s good for entrepreneurs to see even with thousands of people this is what what you can continue to achieve.

Jason Calacanis: 23:43 Right, kick Google in the ass, make them get sharper, Amazon start waking up.

Andrew Feldman: 23:47 That’s right, everybody got better because of that. We want that. And we certainly don’t want to become sort of a region where the first thing we want to do is regulate it, right? But as it gets more powerful…

Jason Calacanis: 24:00 The industry really should do a better job of regulating itself perhaps, and it did seem like they were starting that process, but then the communication was lacking maybe.

Andrew Feldman: 24:10 Yeah, it’s uh, you know, I think not only are they racing hard, but they’re inventing this as they go too.

Jason Calacanis: 24:18 Yeah. Right, there’s not a playbook.

Andrew Feldman: 24:20 No. Right, they’re inventing the… we say oh just put on guardrails, well they have to design the guardrails. Right, the guardrails have an impact. Um, you know, one of the things that FAST does is it makes the guardrails less painful. And so that’s what we… you know, we discovered that in the last six weeks. Is that the very guardrails can add time and make it feel slower. And so FAST chips like ours can really help that. But so they’re racing against competition, they’re racing against their own sense of greatness. Right, which is maybe even the biggest driver here. And I think they’re earnest trying to think about how to do the right thing. And all of those are mixed in this bucket. And sometimes you err on one side or other than the other.

Jason Calacanis: 25:11 Yeah, and as you’re saying, this is a first time, right? When with 3.5 came out, it wasn’t like when we’re using ChatGPT 2.5, 3.5, it was taking down network. But in talking to Nikesh from Palo Alto Networks, I asked him, like, hey, well, how would you grade this? And he said, oh, we put it against our software and we found bugs we were not aware of.

Andrew Feldman: 25:34 Oh, yes. It killed them.

Jason Calacanis: 25:35 Yeah, he said we had to stop everything we’re doing and do patches for six weeks.

Andrew Feldman: 25:40 Right. And that’s when you know, right? I mean Nikesh leads maybe the leading security software firm, right? And when it finds in an hour, right, tens of critical opens, you’re like, whoa, this is a powerful tool. And we need to think. And maybe you show it to a group first, right? Maybe you, I don’t know what the right thing is, but…

Jason Calacanis: 26:05 I mean red teaming, and we’ve always had, just when you were releasing the new version of an operating system, you know, when you have your iPhone, you can say I want to be part of the beta.

Andrew Feldman: 26:18 That’s right. Right.

Jason Calacanis: 26:19 And there are like two other betas that you don’t even get the chance to opt into as consumers, and those ones are for security, those ones are for, you know, making sure you don’t lose your data, or data doesn’t disappear, or leak, or corruption, any, any number of these things.

Chamath Palihapitiya: 26:33 I think we can also know that there will be a massive data leak. We know this, right? And it’s like Warren Buffett talked about the reinsurance industry, that you know something bad’s going to happen, you don’t know when, but you gotta save up for it, right? You put money away for it, rainy day, reinsurance. But there will be a tornado, there will be a massive earthquake. I mean we know this, and we can do our best to…

David Friedberg: 27:00 plan, but there’ll be a massive breach, and we’ll have to steel ourselves in advance, and we have to think about it and think about the right response at the time and sort of prepare ourselves for a future that is, in specific, unknown, but in general, we’re pretty sure something’s going to happen.

Jason Calacanis: 27:17 Something will happen, and yeah, it’s typically a black swan, right? I mean, by definition, it’s going to be something we didn’t consider or a question we didn’t know to ask.

David Friedberg: 27:28 Right. But even knowing that there’s some unknown unknowns is a useful place to start.

Jason Calacanis: 27:35 Yeah. What are we not asking ourselves?

David Friedberg: 27:38 With reasoning, the AI is going to be able to tell us, ‘Hey schmuck humans, that’s right…’ By the way, here’s what you’re not thinking about. This is now my closing sentence when I do my prompting is, I need you to make me a prompt that will help me do this trend scouting, for example. And then I always say at the end, please check your work and then tell me what I haven’t considered in terms of my goals and ask me some questions every time you run the job. And that has changed everything because it’s like, ‘I checked my work, by the way, this was incorrect, and I’m wondering, hey, would you like me to also do this?’

Jason Calacanis: 28:14 And some of the tools like Perplexity do that automatically. They give you your next three prompts. But if you give it explicit instructions…

David Friedberg: 28:19 My lord is it good at that. So, you know, over the course of the last 10 years as I was raising money, I thought one of the smarter questions I got at the end of a conversation was where someone asked, ‘What was the smartest question you heard that wasn’t covered by what I asked?’

Jason Calacanis: 28:41 Incredible. Right. Now that’s somebody who’s curious and thinking and humble and trying to sort of use this to get a picture of the space.

David Friedberg: 28:52 And to the extent that you can ask the AI that, and that it can sort of broaden your view, you know, maybe what question should I have asked to be an expert in this?

Jason Calacanis: 29:00 What, what would a PhD level questioner ask of this? Or a gold medal math— I mean, I think those are sort of questions that you know you don’t even know how to ask.

David Friedberg: 29:11 Which, you know, if we start thinking about AGI and superintelligence, you know, they’re just definitions, but they’re important definitions I think to kind of keep in mind, because they’re waypoints.

Jason Calacanis: 29:29 That’s right. And AGI, I think, I suspect you’ll agree with me that we’ve hit it. We just haven’t exactly deployed it fully. We have artificial general intelligence now. It feels like when we’re talking about these reasoning moments and, you know, the ability for it to be as smart as any human. But let’s talk about…

David Friedberg: 29:49 By any definition we had 20 years ago, we’ve hit it.

Jason Calacanis: 29:54 Yes. Right. I mean, if you think about all those Turing tests, blew it away. I mean, you think about any period of time sort of 10 years ago…

Andrew Feldman: 30:00 10, 15, 20, 30, 40, 50 years ago, we any definition we would have previously put forward, we’ve blown past it.

David Friedberg: 30:08 Which goes back to our previous point of like, do we know the questions to ask? 20 years ago, science fiction authors, you know, had their say and we answered all their questions. If they were to look at this today, they’d be like, well, I’m out of questions. Sorry.

Andrew Feldman: 30:25 And that’s where sort of the sort of listening to people who we who sound sometimes like they’re on the fringe. Right? When Ilya was talking eight or ten years ago about the need for safety and you’re like, what? And dead right. Right? When Elon was talking about building rockets and driving the cost to near zero of a launch vehicle, you’re like, what? And there it is. And now you can see. And that’s… I think that’s why it’s really fun to be a technologist now.

David Friedberg: 30:44 Well, and with these tools specifically, you know, we’re talking about building all these tools and then the tools are starting to build themselves in this recursive loop. That’s right. We’re kind of just starting to see people apply loops. In fact, loop maxing became when I was doing my trends… When I did my trend, right, it kept picking up looping and it kept picking up the maxing stuff and it created a buzzword for me, loop maxing. And then it magically people started talking about loop maxing and I was like, wow, this is really weird. It anticipated that this would other humans would come up with this word. But talk a little bit about recursive and then the road to super intelligence and do you have a way, Andrew, that you think about super intelligence and what it will mean for humanity and how we will define it and how we’ll experience it? Yeah.

Andrew Feldman: 31:53 I think let’s begin on on loop maxing or sort of recursive learning. I think what what Sam and Ilya and then later Dario and and Demis saw six years ago or five years ago was that powerful recursive gains are exponential. Right? You get better, you do it again, and if you continue to get gain, the the slope of that curve is so steep.

David Friedberg: 32:19 Yeah.

Andrew Feldman: 32:20 And that um we’re just beginning to see that now. You ask it a question, you learn from the results, you ask it to do it again, the results get better and more information’s added, your answer’s better, you ask it to do again, it covers more material, and these sort of loops are producing sort of not a little bit better answers, but vastly better answers.

David Friedberg: 32:41 Yeah.

Andrew Feldman: 32:42 And that is enormously powerful because we don’t quite know where it ends.

Jason Calacanis: 33:00 Right. But you keep throwing compute at it, I mean how much better does he have to get? You know, we run out of tokens or our budget or… but holy cow, I mean when does the exponential stop or does the answer keep going up and up and up to the right?

David Friedberg: 33:15 Yeah, that’s sort of an enormously interesting intellectual question right now.

Jason Calacanis: 33:20 Yeah, like when do we run out of problems to solve?

David Friedberg: 33:24 Well, that’s right. And when are the the problems no longer sort of intellectual problems and they’re now people problems?

Jason Calacanis: 33:34 Yeah.

David Friedberg: 33:34 Right? How to organize people to to get done what the AI asks for.

Jason Calacanis: 33:38 Right.

David Friedberg: 33:39 I mean, as you know in running your company, a lot of your problems aren’t hard intellectual problems, they’re people working together problems.

Jason Calacanis: 33:46 Yeah, right.

David Friedberg: 33:47 Motivation.

Jason Calacanis: 33:48 Motivation. You spend a lot of time as a leader spraying WD-40 on your team.

David Friedberg: 33:53 Right.

Jason Calacanis: 33:54 Right? Just so so friction is reduced and…

David Friedberg: 33:57 Uh-huh.

Jason Calacanis: 33:58 How do we learn about those from AI? Right? How do we get behavioral insight from from AI? And I think that’s some of the things the world models are going to bring us as they begin to watch human behavior.

David Friedberg: 34:13 Yeah, we didn’t even get to that. This is going to be for another interview, but when these things jump off the screens…

Jason Calacanis: 34:21 Right.

David Friedberg: 34:22 …and they’re in the real world and the recursiveness starts not trying to solve math problems and, you know, humanity’s most difficult ones, but…

Jason Calacanis: 34:31 Hey, you know, there’s an incredible world out here and here’s the Palace of Versailles.

David Friedberg: 34:35 Right.

Jason Calacanis: 34:35 You’re just like now we’re like, make me a new version of Salesforce and we’re like, hey, you know what? I’d like the Palace of Versailles. I’ve got 100 acres somewhere out in Texas or Nevada. I’ll just send a thousand Optimuses out there. Make me the Palace of Versailles.

David Friedberg: 34:51 Right.

Jason Calacanis: 34:52 Sounds fantastical, but the Palace of Versailles would seem fantastical to people who lived a thousand years before it.

David Friedberg: 34:59 And it was fantastical I think to the people who built it.

Jason Calacanis: 35:02 Right?

David Friedberg: 35:02 Right? Even to the builders, I think they were awed at it as they built it.

Jason Calacanis: 35:07 Yeah.

David Friedberg: 35:08 They’re compounding their compounding recursive learning, that’s right, and generations we talked about. You had a really such a great insight of in building this place you had generations of masons.

Jason Calacanis: 35:22 Yeah, I think in in all these large projects, uh-huh, often there were families who were specialists who were used, and you you apprenticed under your father, your uncle, and when you had a project that took 50 or 70 or 100 years, you might have three or four generations of the same family, right, the same stonemason family working on the same structure.

David Friedberg: 35:44 And passing on the learning.

Jason Calacanis: 35:46 And passing on the learnings.

David Friedberg: 35:48 Uh-huh.

Jason Calacanis: 35:48 New innovations.

David Friedberg: 35:49 Right.

Jason Calacanis: 35:50 Which is what we’ve modeled with this new…

David Friedberg: 35:52 That’s right.

Jason Calacanis: 35:52 …models and and what you’re building in the infrastructure.

David Friedberg: 35:55 It’s pretty incredible when you think about it.

Jason Calacanis: 35:57 Pretty cool.

David Friedberg: 35:58 Especially when we’re sitting here in Versailles. And that’s what I mean, I think the… The problem with human learning is um it often moves at the pace of a generation.

David Sacks: 36:08 Huh.

David Friedberg: 36:09 And like elephants and other large mammals, we don’t have generations but every 15 or 20 years. And if you want to move really quickly across generations, you you want them happening more like drosophila, like fruit flies. You want two a day.

Jason Calacanis: 36:25 Yeah. Right. Then you see that in genetics. That’s why we study them in genetics because learning encoded in the DNA you can study over thousands of generations.

David Friedberg: 36:39 And I think what we’re getting is that equivalent in AI. We’re getting sort of learning so quickly over the equivalent of thousands of generations.

Jason Calacanis: 36:43 Yeah Darwin would be in awe of this pace of evolution.

David Friedberg: 36:51 That’s exactly right.

Jason Calacanis: 36:52 You think about it as there was I remember when I was getting my psychology degree and they were teaching us about paradigms and I was like trying to understand how the paradigms shifted. And the professor said to me, Jason, what you have to understand is paradigms don’t die. People do.

David Friedberg: 37:11 That’s right.

Jason Calacanis: 37:12 And Thomas Kuhn, as you say.

David Friedberg: 37:13 Thomas Kuhn, exactly.

Jason Calacanis: 37:14 Freud, and Skinner, and Jung, like it took them dying for the next generation to question it.

David Sacks: 37:19 And that was 20 years. Sometimes 40 years as their students maintained positions of leadership until someone said maybe we could do it differently?

David Friedberg: 37:31 And I think what you’re seeing is this iteration is a shortening of the inter-generation gap and the learning is so fast.

Jason Calacanis: 37:45 It’s always so great to talk to you because one, it’s just intellectually um so your approach to it is so intellectually rigorous, but also with so much p-doom in the world, I feel so good about your being such an optimist about this technology and you’re building it with such thoughtfulness and I think for people who are hearing these horror stories about AI and job loss and everything, they need to understand there are people like yourself who are building this in an incredibly thoughtful way and this is going to be a net benefit for humanity that just is unimaginable, yeah?

David Friedberg: 38:30 We have a shot with this technology so that not our children, nor anyone they know dies of cancer.

David Sacks: 38:33 I mean, say it like that. There will be some dislocation in the economy. Sure. There will be. There was dislocation when cars came and and it was a bad deal to be a guy who shoed horses, right, or build carriages. But you gotta also against that, you know, make your T of the cons and the pros. Yeah. There’s a shot that our children, none of them nor their people they love will die of cancer.

David Friedberg: 38:56 And that’s one. That’s just one thing that we can work on with this technology and we will have great tools…

Andrew Feldman: 39:00 great purchase on. And and I think you begin listing those and then it’s a more thoughtful discussion.

Jason Calacanis: 39:06 Yeah, unlimited energy, unlimited calories, unlimited knowledge, unlimited education, unlimited housing.

Andrew Feldman: 39:12 And how we do it, we imagine imagine sort of we know how to teach children and we don’t do it, right? Aristotle was a tutor to Alexander the Great, Socrates was his tutor. We know that if you give a child a tutor and the tutor modifies the teaching for the child, they learn better.

Jason Calacanis: 39:29 Yeah.

Andrew Feldman: 39:30 That’s not how we do teaching in classes.

Jason Calacanis: 39:32 No, factory farming.

Andrew Feldman: 39:33 That’s right. We teach to some sort of mid-level. Imagine if we built agents that that taught children for their way of learning. Right? And here’s the way we we’ve been doing it the same way for a thousand years and during that entire time we knew how to do it better and we chose not to. Yeah. And and here’s the way we can do it. Put that on the pro side. And so as long as we’re sort of thoughtfully and fairly writing the good and the bad, I think it’ll come out…

Jason Calacanis: 39:59 Well, you got to get out there, Andrew, and keep communicating your version of the world because some folks people see around the corner and they get a little nervous and okay, fair enough. But I think the ledger, as you describe it, is heavily weighted towards abundance.

Andrew Feldman: 40:14 I think it’ll create abundance for sure.

Jason Calacanis: 40:17 Massive abundance. Andrew, pleasure, always a pleasure to talk.

Andrew Feldman: 40:19 I’ll see you in six months for our checkup.

Jason Calacanis: 40:21 That’ll be great. Thank you so much.

Jason Calacanis: 40:53 Robin Rombach is the co-founder and CEO of Black Forest Labs. You are based in Germany, in Black Forest, which is a city in Germany.

Robin Rombach: 41:04 It’s a mountain range actually.

Jason Calacanis: 41:07 A mountain range. Where you grew up.

Robin Rombach: 41:09 Where I grew up, yes.

Jason Calacanis: 41:11 And you are working on open source image and video models. You worked at Stable Diffusion for a little bit, cut your teeth on that, and you’re known for the open source model Flux, and maybe also for some closed source models. Tell us about the business of Black Forest Labs. What is the business and what is the goal?

Robin Rombach: 41:33 100%. One quick addition. We are based in the Black Forest. It’s a town called Freiburg and in San Francisco. We have two offices.

Jason Calacanis: 41:44 Oh and San Francisco, of course, yeah. You’re splitting your time or?

Robin Rombach: 41:47 I’m splitting my time to a certain degree. We yes, started a company two years ago, we and my co-founders, as you said, like we worked on Stable Diffusion in the past before that, we invented like an algorithm called latent diffusion, which is basically like the foundation of all these models. the fundamental algorithm behind all of like generative models that are being deployed for image generation, video generation, even like physical AI now.

Jason Calacanis: 42:03 Yeah.

Robin Rombach: 42:03 Uh, it basically makes use of this principle that you can compress natural data such as images, such as videos, such as audio into a much more like efficient representation and then train a Transformer model on that.

Jason Calacanis: 42:23 Hmm.

Robin Rombach: 42:24 And, um, I mean, this is the stuff where like, you know, like JPEG, MP3, and all of that works. And we basically translated that into like a neural algorithm a few years ago when we were still like PhD students in, in Munich actually. And, uh, then built on like, on top of that we built Stable Diffusion. And then on top of that, uh, yeah, the generative models that we are developing today. And of course, like the technology has advanced. Um, but we are now tackling, I would say, models that are really made for understanding like the whole world around us. Multimodal visual models, pre-trained on images, videos, audio data at the same time. And we are now like entering a new paradigm, which is combining that with uh something that’s called action prediction, such that you can actually use the same model to make images, to make videos, to make audio and to predict actions, which means you can ultimately deploy it on a robot in the real world.

Jason Calacanis: 43:29 Wow. So, from the image to the video, the audio, and then eventually the real world with robotics and a real world model, because if you can make the image, you- and you can train the model, that means by default you understand the world. In order to make a video of the world, you have to understand the world, yeah? And the objects in it?

Robin Rombach: 43:51 I, I think that’s, you know, I think that’s like a really good like way to think about it. Uh, it’s like- it’s like an intuitive way, uh, to interact with the world, right? Like, I would say there’s like these like complementary forms of intelligence ultimately. There’s like intuitive intelligence and then there’s like the deep reasoning layer. Now, ultimately you need for like a kind of like complete form, you need both, um, and you need them to interact. And I think like we have been approaching it more from like the intuitive side. Images is like a very natural way to approach this whole field because it’s not as computationally intensive as let’s say video, right? But now, yeah, I think like we are combining it, it’s converging into like a- a multimodal model. And, uh, yeah, we see like exactly like pre-training on videos gives like implicit understanding of the physics of interactions with the real world, and then you can get stuff like action prediction, like robotics, out of the same model.

Jason Calacanis: 44:50 And with these models and the training, they’ve kind of uh been a limitation in creating videos and creating images where the criticism of… Generative AI is a bit of a slot machine. I give a prompt, it gives me something back. But how did it come up with that? The training data, but you know, maybe I want a different style, maybe I want a different color, maybe I want a different, you know, aesthetic.

Robin Rombach: 45:20 Yep.

Jason Calacanis: 45:21 Has that… how does that problem get solved and do you actually understand what’s happening when the image is being made under the hood?

Robin Rombach: 45:32 Yeah. Yeah, I think, like, ultimately it’s about, like, exposing as many, like, manipulation layers as possible to, like, I don’t know, like a user or developer that builds on top of this model, right? And I think, like, we’ve seen that in the past with, like, in the past, image models, they basically started from simple text-to-image systems.

Jason Calacanis: 45:59 Right.

Robin Rombach: 46:00 Then they’ve expanded into text plus image-to-image systems, which means you could suddenly take an image, like a real image or a generated image, and iterate on that based on a text prompt, like, edit it, modify it, right? And then this expanded into taking multiple images and a text prompt and combining them in a, in a semantic way and producing new content. And the same principle now applies to video, and I think now it becomes actually even more interesting when, like, all of these, like, modalities are actually combined inputs and outputs of the same model.

Jason Calacanis: 46:29 So let’s talk about video. There’s an announcement that you’re working with the greatest director of all time, or living director, Martin Scorsese. We’ll talk about that in a second, yeah?

Robin Rombach: 46:33 Fantastic.

Jason Calacanis: 46:34 Uh, but in a movie, this promise of being able to make a movie in which the camera angle, the sound, could be something that a Martin Scorsese would be proud to release to his fans. How close are we, and maybe tell us a little bit about this partnership, the technology being able to make an actual movie like Goodfellas or a scene from Goodfellas, versus where it is today, where you can make interesting five or ten second clips, and then maybe, oh, people struggle making ten of them and then they use some post-editing software to put them together, but you immediately understand this is not that. It’s not a movie. It’s AI slop, it’s kludgy, it doesn’t pass the uncanny valley.

Robin Rombach: 47:41 Well, I think it’s important and that’s at least, like, the view that we have, is that these AI models, they are a medium, right? They, we don’t want to set, like, any way of how they are supposed to be used. We don’t want to tell anyone, especially not someone like Martin Scorsese, how he is supposed to use this model. Like, he is one of the, like, greatest filmmakers ever. It was insane… sitting in the same room with him multiple times and actually him seeing like exploring our models like as like one of the like core researchers behind it was like just an insane feeling, right? And at the same time I’m also like a big fan.

Jason Calacanis: 48:12 So you sat in a room with Martin Scorsese and showed him your tools.

Robin Rombach: 48:16 Exactly, yeah.

Jason Calacanis: 48:18 And what was his reaction? What did he key off of? What was the thing that he found most inspiring or interesting?

Robin Rombach: 48:26 Um, I think it was really this idea of like, he has clearly a vision in his head of like a scene or a scenery where like maybe a new movie will be shot. And he’s trying to explore that and kind of like we basically looked at the scenery of like a village in Eastern Europe somewhere and he was describing it, we saw some outputs, we iterated on the outputs, um, and ultimately I think, and that’s what he said in the end, is like the like getting like the mental picture of something out of your head and communicating it in a visual way by making like these images or the series of images, um, is something, yeah, that just makes it like easier to communicate and convey like an idea of like what is actually in your head. And I think that’s like one of the like very interesting and powerful ways to use this technology. And I think ultimately…

Jason Calacanis: 49:25 It’s to get the inspiration, to get the vision out of his head onto an image.

Robin Rombach: 49:31 Yeah, I mean like language ultimately is like a little bit of a lossy communication medium, right? Um, it’s also interpreted in different ways but then visual information is so rich, so rich, like an image or video, there’s so much signal in it. And it’s just like another way of communicating. And I think that’s like one of the beautiful things that this technology ultimately enables. And I think like to your question of making like full movies with, I don’t know, like a video generation model for example. I’m not sure if that is like the ultimate goal. Maybe it’s like interesting to plug this into like some kind of agentic workflow and make like a very long video and I think that’s really cool to explore, but I think ultimately like the real interesting use cases they come when you have like a human in the loop who iterates and uses it as a medium. And I think this is, this is at least like the perspective that that I take, um, that makes it interesting and that this is most often when the most interesting outputs arrive or are actually being made.

Jason Calacanis: 50:31 The brainstorming production level is so obviously a huge win for generative AI.

Robin Rombach: 50:37 Yeah, you can parallelize your brainstorming basically.

Jason Calacanis: 50:40 Yeah, and I like that, parallelize your brainstorming. And they have an analogy for this, they do storyboards. And some of the great directors, Ridley Scott of Alien and Gladiator, was known for making his own. I also believe Spielberg was also liked to sketch Raiders of the Lost Ark and some of these. George Lucas was known for collaborating with many amazing artists, um, even making miniatures, even making… making storyboards for the Star Wars franchise he had those people on full-time helping him with that so that’s the obvious place to start but if we look at startups startups always want to try to figure out how to do something cheaply and people used to make a launch video for their startup for you know $100,000 $250,000 so they’d take their $10 million venture raise and spend 250,000 on a launch video I’ve seen with a lot of the startups I’m investing in now they’ll just spend a week or two working with um a director to make a launch video you’ve probably seen this trend and I’m sure people use flux and some of your models for this have you seen this?

Robin Rombach: 51:44 Yeah of course yeah

Jason Calacanis: 51:45 Yeah and what’s your take on that because that feels like the early stage of storytelling you’re trying to communicate a product or service in a fun engaging punchy 30-second 90-second way yeah?

Robin Rombach: 51:58 I mean like again like I think we support these like exploration based on these tools right and I think like ultimately it’s great to see like all different kind of like I don’t know like launch videos products being built on top of like the same kind of like base model or the same technology and I think that’s what’s making it so interesting and also so powerful

Jason Calacanis: 52:23 Yeah and what else are people using the technology for I understand there’s a Bitcoin movie coming out instead of using a green screen in this Bitcoin movie um I was talking to Gal Gadot you know the woman who played the actress who played Wonder Woman Gal Gadot she I was talking to her at an event and she was telling me um it was the Breakthrough prize Yuri Milner’s event and she was telling me she just did a Bitcoin movie and they did it on a sound stage without green screens but all the actors just worked in like a sound stage and then all of the scenery behind them was being done by generative AI that’s the real movie that’s a $30 million budget movie she said it would have cost 150 million if they had to build sets and the film would have never been greenlit are you starting to see people use that in production not just in the backend in the ideation phase but actually in production yet with your tools?

Robin Rombach: 53:18 Um yeah we see some use cases like that in production I think like high-end film production is kind of like the one of the like most demanding use cases and I think I’m glad that it’s being explored but I also yeah really want to uh like it’s I think it’s important to see that this technology is like on a trajectory and it’s improving it’s improving rapidly I don’t know if I look back at like where we started like a few years ago when I was doing my PhD in this field like the only thing that you could do was like images of 64x64 pixels now you can do like multi-minute videos right at like a high resolution but it’s like it’s not going to stop there it’s going to continue to improve and I think like then it’s going to unlock like even mostly like high-end use cases, but I think the main thing…

Jason Calacanis: 54:03 How soon, before we get to that, yeah.

Robin Rombach: 54:05 How to predict, I think, how to predict and I think ultimately…

Jason Calacanis: 54:07 Couple of years.

Robin Rombach: 54:09 Ultimately I think you still want to have like the tool that enables this human-in-the-loop kind of of course yeah production workflow right? But I think when I look at multimodal generative models as a whole, I think what really excites me is you can use the same kind of AI model to make a movie and deploy that as a brain on a robot. And I think this is like this is so interesting and I don’t know like there’s like some thoughts around trying that in the digital world right, which would be for example computer use has remains to be seen if that is actually something that works or not, but I think like the technology is so powerful and so versatile and it’s just just moving into that in all the talk around like world models, world action models, all of that, it’s basically all the same and I think that’s what’s making it so interesting and what I find like most exciting.

Jason Calacanis: 55:14 So do you believe that the technology will be used to analyze or primarily to analyze real world like here is a video of somebody you know making a sandwich, now we have the robot study it and make the sandwich, or do you think there’ll be a lot of synthetic data made that then the robots will just study the synthetic or they’re going to just in some way innately know based on all this massive amounts of training data?

Robin Rombach: 55:31 Um, I think it’s a combination of prediction, right? And prediction is a way of you can think about it as simulation, as generation. It’s predicting actions, which is you have to understand the input, the visual inputs in order to actually predict a reasonable next action. Um, and it’s about perception. It’s like you can only do that if you understand, if you perceive the content, then you can only, I don’t know, like transform it into a new piece of content or predict an action or describe what you actually see in that scene and the combination of all of that is I think, yeah, is I think what’s driving it. It’s not a single one of them, it’s a combination of these thoughts.

Jason Calacanis: 56:24 Right. And what’s the best way to get that training data? Do you need to have people put on glasses, get a first person perspective, have them put on gloves so you have that you know fidelity of understanding, hey this glass is moving, I’m pouring this glass, I’m putting ice into it you know and here is how that’s works and the splashing and the condensation water so I can pick it up and not drop it because it’s wet on the outside, or is it going to be just, hey, take the corpus of YouTube videos and the robots know exactly what to do because they’ll find a thousand videos of people pouring drinks?

Robin Rombach: 57:00 I mean, ultimately, I think you would want to go to a place where you could like prompt a robot in context, right? As you can do with like a language model, basically just tell it, hey, go and I don’t know pick up this glass with the, I don’t know, orange juice or whatever it is.

Jason Calacanis: 57:14 Make a cocktail?

Robin Rombach: 57:15 Yeah, exactly. We’re not there yet, but I think this is like one of the goals. And I think like how these models are deployed currently is there’s like a lot of like different hardware, different robots that are running in factories, that all have like some different kind of action representation that you need to kind of tune the models towards, right? So in practice, what you do is you have like all this like visual understanding in the models, and then you need only a very little bit of like a few hours of like fine-tuning data to adjust the model on that specific task. And I think the goal would be to kind of move away from that towards like as much in-context as possible, but it is still a little bit of a research problem, I think that has—

Jason Calacanis: 58:03 Open source has kind of having a moment right now. We’ve been discussing it on the podcast a whole bunch recently, and people are also talking about sovereignty. You have companies that own incredible IP libraries. I mentioned Star Wars before; Disney owns an incredible library. What should your advice be to a company like Disney? Should they take your open-source software, train their own models, or work with you to train their own models to control it, and then, ‘Hey, this is our IP.’ They’ve already made a point of working with ChatGPT and saying, ‘Hey, you can and cannot use certain characters.’ In fact, OpenAI had a relationship with them that’s—for Sora, that’s no longer happening, but they officially licensed on the output some characters. So how do you think about those major IP holders? What’s your advice to them? Are you in discussions with them? We know about the Martin Scorsese or Tour DL, but how do you think about content libraries?

Robin Rombach: 58:58 I think it is—look, I think like the most interesting use cases of this, like if you think about like content creation, is in generating something, making something that hasn’t been there before, right? That’s the fundamental, like, interesting aspect of this technology. And then I think, yeah, when it comes to IP, what we implement, for example, on like our public-facing tools, is you cannot generate certain IP with these models, right? And I think that’s something that is a sensible approach. And then, yes, we do work with certain IP holders to develop models together with them, some of them based on our open-source models, some of them based on like our more powerful proprietary models. But I think that is like a very like attractive value proposition.

Jason Calacanis: 59:45 What’s the vision there? What do you think that will look like for consumers in another couple of years? What would potentially happen when you open up Disney+?

Robin Rombach: 59:56 I mean, that’s a good question. I’m not in Disney, right? So it’s up to them to decide that. But I think we want to enable them to build all kinds of— the stuff that they, they, they envision. And I think we can support them, we can support, like, other companies in that space to, I don’t know, integrate the technology in the best possible way. I think, like, one of the very interesting angles of it is that it is, it’s becoming much faster, it’s becoming more interactive. I can envision, like, a whole bunch of, like, very interesting interactive content creation tools that you could host on Disney Plus or elsewhere.

Jason Calacanis: 1:00:23 I think the most interesting thing I’ve seen in this regard is fan films. Right. So there’s a category before generative AI, fan fiction. People would write their own Star Wars story. Then there came fan films, where people would dress up as Jedi Knights and record their own films. And George Lucas said, as long as you’re not doing it commercially or not selling it, I give you permission to go make Jedi movies. And they even released how-tos on how to make a lightsaber or, you know, sound files of, like, how to make the lightsaber sound. Now, people are taking the stories that haven’t been told from the Star Wars universe, and they’re recreating them using AI. And for the fans, they’re becoming quite popular on YouTube. Star Wars Stories Untold is, I think, the biggest one. It’s getting millions of views per video already. And I think that’s really the future is letting the customer base pay a licensing fee or pay a fee, maybe rent software, or maybe based on the output, and let them be creative with the characters, let them make their own stories. Um, and you could be in a unique position to empower them.

Robin Rombach: 1:01:40 Well, 100%. I think, like, if you find, like, a model that works for, like, the IP owners, but then also can enable, like, these super, like, creative customization use cases, I think that’s great. Yeah, I mean, like, I mean, like, for myself, like, I when I read a book or whatever, like, watched a movie, I had, like, so many, like, ideas how it could be done differently or this could have happened, right? And, like, this is just, like, so nice that you can actually enable people to visualize these ideas.

Jason Calacanis: 1:02:07 Yeah, it’s going to be incredible. Continued success with it. You have an office in San Francisco, you’re hiring people, yeah? You’ve raised a bunch of money.

Robin Rombach: 1:02:20 We do, yeah. We just raised a bunch of money. We just crossed 100 people. We’re hiring in Germany and in San Francisco.

Jason Calacanis: 1:02:23 Fantastic. Who are you looking for? What’s the right type of person, the right type of skill?

Robin Rombach: 1:02:25 Yeah. Um, on the one hand, we are always looking for researchers who have experience in large-scale model training, um, experience in diffusion model training, score-matching training. We’re looking for engineers who want to be working with the customers to, you know, develop these, like, customized visual AI solutions or, for example, with, like, an IP owner, like, develop these models jointly with them. Um, we are looking for engineers who have experience in just, like, large-scale compute infra, managing that, um, and making sure that the training runs, runs smoothly, that we can…

Andrew Feldman: 1:03:00 Maximize our MFU and all of that, and we are looking for people who have interest in, you know, like, getting the technology out there in the hands of people.

David Friedberg: 1:03:12 Yeah, the- the full deployment of this, there’s just so many great ideas and so many great partners for you. I think you’re going to, with the open source specifically, you know, it seems like the corporates really want to have some additional level of control, but they also need the frontier models or your proprietary ones for some of those refined features. So I think you have a very bright future ahead.

Andrew Feldman: 1:03:29 100%, 100%, exactly.

David Friedberg: 1:03:31 100%, 100%, exactly.

Jason Calacanis: 1:03:32 Alright, continued success. Thank you so much for doing the show. A pleasure.

Andrew Feldman: 1:03:36 Thank you so much. Thank you so much for having me.