How OpenAI Is Rewriting Its Future
How OpenAI Is Rewriting Its Future
Summary
Kevin Roose and Casey Newton break down OpenAI’s strategic reset week: a rewritten Microsoft partnership that frees OpenAI to work with other cloud providers, a new $50 billion expansion with Amazon’s Bedrock platform, a scaled-back Stargate buildout, and a pivot toward an $8/month ChatGPT Go subscription. They also dig into the federal trial in Oakland where Elon Musk is suing OpenAI for “looting a charity,” and what the lawsuit reveals about petty grudges driving the AI industry.
Dr. Adam Rodman returns to discuss the remarkably rapid adoption of AI in medicine. He explains how AI scribes and Open Evidence have gone from novel to commodity in under two years, walks through his “green light / yellow light / red light” framework for patients using ChatGPT, and weighs in on Mayo Clinic’s RedMod system that can detect pancreatic cancer up to three years before diagnosis. The biggest concern, he says, isn’t safety — it’s de-skilling of doctors-in-training.
Finally, David Duvenaud joins to discuss Talkie, an experimental “vintage LLM” trained exclusively on data from before 1931. The project explores whether machines can learn to forecast the future by first being asked to predict from a known past, and reveals how period-specific language models surface both beautiful early-20th-century prose and the uglier prejudices of that era.
Highlights
”The AGI clause is gone”
“The old agreement said that basically once OpenAI reached AGI, Microsoft would stop getting certain revenue share payments. But under the new agreement, OpenAI will keep sharing revenue with Microsoft until 2030, no matter what benchmarks they hit. So the AGI clause is gone.” — Kevin Roose, 1:55
Clip command
yt-dlp --download-sections "*1:55-3:00" "https://www.youtube.com/watch?v=3khLoSNCjWw" --force-keyframes-at-cuts --merge-output-format mp4 -o "agi-clause-gone.mp4"
”Even the biggest companies do not have the resources”
“The story that you just described, Kevin, is one of a world where no one has the resources they need to serve the demand for AI that they have… I just want to, like, posit that as a really important point in understanding what sort of bubble this is, because even the biggest companies do not have the resources that they need to serve that demand.” — Casey Newton, 6:15
Clip command
yt-dlp --download-sections "*6:15-7:28" "https://www.youtube.com/watch?v=3khLoSNCjWw" --force-keyframes-at-cuts --merge-output-format mp4 -o "ai-demand-bubble.mp4"
”It is not okay to steal a charity”
“Elon Musk himself took the witness stand and said quote, ‘This lawsuit is very simple, it is not okay to steal a charity.’ He also said that if OpenAI is allowed to get away with this, quote, ‘it will give license to looting every charity in America.’” — Kevin Roose, 15:10
Clip command
yt-dlp --download-sections "*15:10-16:44" "https://www.youtube.com/watch?v=3khLoSNCjWw" --force-keyframes-at-cuts --merge-output-format mp4 -o "steal-a-charity.mp4"
”AI in medicine is probably the fastest adopted medical technology of all time”
“AI in medicine has gone from, well, depending on how you measure it, it’s probably the fastest adopted medical technology of all time. We went from this being super novel, almost no one used AI tools to just being a routine part of most doctors’ weekly practice.” — Adam Rodman, 25:36
Clip command
yt-dlp --download-sections "*25:36-26:30" "https://www.youtube.com/watch?v=3khLoSNCjWw" --force-keyframes-at-cuts --merge-output-format mp4 -o "fastest-adopted-medtech.mp4"
”Green light, yellow light, red light”
“The red light, what I tell them never to do, is like, ask medical management decisions. Like don’t say my doctor said to do this, is this right? Like I have cancer—God forbid you have cancer—is this the right chemotherapy option? Like a lot of those decisions are so nuanced, take in so much information, those are things that the models don’t do well and they’re so sycophantic, they can convince you that they’re saying the right thing even when they’re wrong.” — Adam Rodman, 34:46
Clip command
yt-dlp --download-sections "*34:46-35:45" "https://www.youtube.com/watch?v=3khLoSNCjWw" --force-keyframes-at-cuts --merge-output-format mp4 -o "red-light-uses.mp4"
”How did universal peace come about?”
“I saw another person ask Talkie, basically, the person told it that it was from the future and would tell Talkie anything it wanted to know about the future. And Talkie’s first question was, ‘How did universal peace come about?’ Which was like the most heartbreaking thing I think I’ve ever read from a large language model.” — Kevin Roose, 1:02:34
Clip command
yt-dlp --download-sections "*1:02:34-1:03:46" "https://www.youtube.com/watch?v=3khLoSNCjWw" --force-keyframes-at-cuts --merge-output-format mp4 -o "universal-peace.mp4"
Key Points
- OpenAI’s strategic reset week (0:30) - Microsoft deal rewrite, Amazon expansion, Stargate changes, ad-supported subscriptions, and the Musk trial all hit in one week.
- Microsoft no longer holds OpenAI to its infrastructure (1:31) - Until this week, OpenAI could only serve models on Microsoft’s maxed-out infrastructure; new deal lets them use any cloud.
- The AGI clause is removed (1:55) - OpenAI will share revenue with Microsoft through 2030 regardless of benchmarks; AGI now “evaluated by vibes.”
- Amazon invests $50 billion in OpenAI (4:46) - OpenAI models and Codex coming to AWS Bedrock; OpenAI displacing Anthropic as Amazon’s preferred model partner.
- The new shape of the AI bubble debate (6:15) - Skepticism has flipped from “no demand for AI” to “they can’t build infrastructure fast enough.”
- Stargate scales back (7:34) - OpenAI halted data centers in UK and Norway, paused Abilene expansion; shifting to leasing third-party capacity ahead of a potential IPO.
- The Go subscription pivot (11:41) - OpenAI projects $8/month ChatGPT Go subs to grow 36x to 112M while $20 Plus subs fall 80%.
- The Musk trial begins in Oakland (15:00) - Only 2 of Musk’s 26 claims survived to trial: unjust enrichment and breach of charitable trust.
- OpenAI’s “looting a charity” defense (16:44) - For-profit business is still controlled by the non-profit foundation; 2017-2018 emails show Musk also wanted a for-profit then.
- AI in medicine — fastest adoption of any med tech (25:36) - Going from novel to commodity in under two years per Dr. Rodman.
- AI scribes and Open Evidence (26:10) - Over 40% of US doctors use Open Evidence; one million queries in a single 24-hour period reported in March.
- AMA: 80%+ of physicians use AI professionally (30:58) - But it’s largely Bring-Your-Own-AI rather than hospital-mandated tools.
- Rodman’s green/yellow/red framework for patients (33:09) - Green: diet/clinic prep. Yellow: exploring symptoms. Red: chemo or treatment decisions.
- Cyberchondria warning for the health-maxers (35:51) - Uploading hundreds of labs to ChatGPT mostly drives anxiety; no evidence of better outcomes yet.
- Mayo Clinic’s RedMod for pancreatic cancer (41:00) - Detects subtle CT changes up to 3 years before a pancreatic cancer diagnosis.
- AI’s biggest gain will be access, not new drugs (42:27) - Rodman: improving routine care for the underserved beats waiting for AI-discovered wonder drugs.
- De-skilling is the real concern (44:38) - Polish polyp-detection study showed skilled doctors lost 6 percentage points of unaided detection ability in 3 months.
- Talkie: a vintage LLM trained only on pre-1931 data (49:23) - Built by David Duvenaud, Nick Levine, and ex-OpenAI’s Alec Radford (GPT-1 author).
- Forecasting is the long-term motivation (52:25) - Goal: build models with century-long forecasting track records to evaluate trustworthiness of machine predictions.
- Data contamination remains a challenge (57:00) - Talkie knows about FDR and Hitler because archival sources have wrong dates, prefaces, and historian footnotes mixed in.
- Talkie produces “beautiful prose by a terrible person” (1:01:25) - Period-accurate racism appears in some responses; team uses a modern model to flag warnings rather than filter training data.
Mentions
Companies
- OpenAI (0:30) - Subject of the trial, strategic reset, and most of the episode’s news.
- Microsoft (1:09) - Still OpenAI’s biggest investor at ~$135B stake; relationship rewritten this week.
- Anthropic (1:07) - Casey’s fiancée works there; previously Amazon’s primary frontier-model partner.
- Amazon / AWS (4:46) - Will invest $50B in OpenAI; offering Codex via Bedrock.
- The New York Times (0:30) - Kevin’s employer; currently suing OpenAI, Microsoft, and Perplexity.
- Perplexity (0:30) - Co-defendant in NYT’s AI lawsuit.
- Google / Google DeepMind (3:57) - Mentioned as alternative cloud provider OpenAI can now use; Demis Hassabis’s AGI claim referenced later.
- Meta (7:34) - Has been poaching senior Stargate figures from OpenAI.
- Tesla (18:00) - Musk wanted to fold OpenAI into Tesla, revealed in trial emails.
- xAI / X (15:10) - Musk’s competing AI company referenced as context for his suit.
- Mayo Clinic (41:00) - Announced RedMod pancreatic cancer detection system.
- Beth Israel Deaconess / Harvard Medical School (24:00) - Dr. Rodman’s affiliations.
- Open Evidence (26:10) - Free medical decision-support tool used by 40%+ of US doctors.
- Epic (30:15) - Largest EHR vendor in the US, integrating native AI features.
- Function Health (35:11) - Premium concierge diagnostic company favored by San Francisco “health-maxers.”
- University of Toronto (51:29) - Where David Duvenaud is an associate professor.
- Financial Times (7:34) - Reported the Stargate scale-back.
- Wall Street Journal (9:00) - Berber Jin reported missed user/revenue targets.
- The Information (11:41) - Reported the ChatGPT Go subscription pivot.
- Netflix (12:00) - Cited as analog for ad-supported cheaper tiers.
Products & Technologies
- ChatGPT / ChatGPT Plus / ChatGPT Go (11:41) - $8 Go tier projected for 36x growth; $20 Plus tier projected to fall 80%.
- ChatGPT Health / ChatGPT for clinicians (23:18) - Healthcare-targeted OpenAI products.
- Codex (OpenAI) (4:46) - OpenAI’s coding model now available on Bedrock; cited for explosive recent growth.
- AWS Bedrock (4:46) - Amazon’s AI platform now offering OpenAI models.
- Azure (3:57) - Microsoft’s cloud, previously OpenAI’s only deployment option.
- Stargate (7:34) - OpenAI’s $500B joint infrastructure project being scaled back.
- Claude / Claude Code (11:11) - Anthropic’s chat and coding tools cited for huge agentic-coding growth.
- Gemini (31:15) - Google’s chatbot, used by some doctors for ad-hoc decision support.
- Talkie / talkie-lm.com (49:23) - Vintage LLM trained on pre-1931 data.
- GPT-1 / GPT-2 / GPT-3 (1:05:08) - Talkie sits between GPT-2 and GPT-3 in size.
- Golden Gate Claude (51:29) - Casey’s favorite weird research model.
- WebMD (31:32) - Casey notes patients have done this for years pre-AI.
- AI scribes (26:10) - Voice-to-note tech that’s gone from novel to commodity in under 2 years.
- RedMod (Mayo Clinic) (41:00) - Detects subtle CT changes up to 3 years before pancreatic cancer diagnosis.
- GLP-1s (43:32) - Rodman’s pick for the “wonder drug” of his career.
- Institutional Books (Harvard Library) (54:00) - Harvard’s scanned-collection dataset used to train Talkie.
- Machine is Mirabilis (1:04:13) - Related project that tried to rediscover special relativity from a 1900 cutoff.
- Apple Watch / Fitbit (33:30) - Wearable data Rodman considers a “green light” use for AI.
People
- Kevin Roose (0:00) - Co-host, NYT tech columnist.
- Casey Newton (0:03) - Co-host, Platformer.
- Elon Musk (0:30) - Suing OpenAI; testified at trial about “looting a charity.”
- Sam Altman (9:00) - OpenAI CEO; reported tension with CFO Sarah Friar over IPO targets.
- Sarah Friar (9:00) - OpenAI CFO at center of internal IPO debate.
- Greg Brockman (15:10) - Among OpenAI co-founders Musk clashed with in original power struggle.
- William Savitt (17:12) - OpenAI’s lead trial counsel.
- Berber Jin (9:00) - WSJ reporter on missed OpenAI targets.
- Dr. Adam Rodman (23:00) - Internal medicine physician at Beth Israel Deaconess, Harvard Medical School.
- David Duvenaud (49:51) - Talkie co-creator, U. of Toronto professor researching AGI governance.
- Nick Levine (49:51) - Co-creator of Talkie.
- Alec Radford (49:51) - Talkie co-creator; lead author of the GPT-1 paper, former OpenAI researcher.
- Demis Hassabis (1:03:46) - Google DeepMind CEO; cited for theory that AGI could rediscover relativity.
- Albert Einstein (1:03:46) - Reference point for AI rediscovery of scientific theories.
- Gavin Leech (1:01:25) - Coined the line “beautiful prose by a terrible person” about Talkie.
Surprising Quotes
“OpenAI wasted no time in its new open marriage with Microsoft. It went back out on the market and found themselves in bed with Amazon.” — Kevin Roose, 4:35
“A shocking percentage of the AI industry is just people who decided they didn’t want to work with Sam Altman and who now have their own companies.” — Casey Newton, 20:22
“It is so weird that we have this technology that is now sort of load-bearing infrastructure for the entire economy, that every business is using to completely reinvent the way that it works. And that out of nowhere, if not specially restrained, it will just start talking about goblins.” — Kevin Roose, 14:01
“These are skilled doctors using a technology and they lose six absolute percentage points of their ability to detect potentially cancer in three months. And then imagine that you’re learning to do it for the first time. Will you ever gain those skills?” — Adam Rodman, 45:00
“In general futurism is a place where people don’t actually often try to predict the future very hard and they more they kind of project their values and if you ask someone what do you think’s going to happen they usually fill in something about what they hope is going to happen.” — David Duvenaud, 1:03:00
Transcript
Kevin Roose: 0:00 I’m Kevin Roose, a tech columnist at the New York Times.
Casey Newton: 0:03 I’m Casey Newton from Platformer.
Kevin Roose: 0:04 And this is Hard Fork!
Casey Newton: 0:05 This week, OpenAI’s big reset. We’ll talk about the company’s new business strategy and its dramatic trial with Elon Musk.
Kevin Roose: 0:12 Then, Dr. Adam Rodman returns to the show to tell us about the latest advances in AI and medicine.
Casey Newton: 0:17 And finally, can an AI made out of very old texts still predict the future? We’re talking about Talkie.
Kevin Roose: 0:30 So Casey, there’s been a lot happening with OpenAI in particular this week. There seems to be something of a major strategic reset happening over there. They’ve got a new deal with Microsoft, an expansion of their deal with Amazon, changes to their Stargate compute strategy, and a new push toward new kinds of ad-supported subscriptions. And of course, they’ve got this big trial with Elon Musk that started this week in Oakland. So let’s get through all of it, but first before we do that, let’s make our disclosures. I work for the New York Times, which is suing OpenAI, Microsoft, and Perplexity.
Casey Newton: 1:07 And my fiancee works at Anthropic.
Kevin Roose: 1:09 Okay, so let’s start this week with this new Microsoft deal. So Microsoft and OpenAI have of course been partners for many years. Microsoft remains the biggest investor in OpenAI. Their stake is valued at about $135 billion. But their relationship has also been strained over the years by various factors, and this week they seem to be sort of consciously uncoupling, or at least rewriting their partnership agreement and allowing OpenAI to be a little bit more promiscuous in who they do deals with.
Casey Newton: 1:31 Yeah, I mean OpenAI just had this real challenge, which was that until this week, they were really only allowed to serve their models on Microsoft’s infrastructure. And one thing we talk about on the show a lot is just that a lot of the big cloud service providers, their infrastructure is just maxed out and Microsoft is one of those. And so for OpenAI’s revenue to grow, they needed to find other ways that they could deliver their services. And so to my mind, that was maybe the most important thing about this deal.
Kevin Roose: 1:55 Yeah, so under this new rewritten version of the Microsoft and OpenAI deal, Microsoft will no longer have to share revenue with OpenAI. The new deal also removes the part of the original agreement that had to do with AGI. The old agreement said that basically once OpenAI reached AGI, Microsoft would stop getting certain revenue share payments. But under the new agreement, OpenAI will keep sharing revenue with Microsoft until 2030, no matter what benchmarks they hit. So the AGI clause is gone.
Casey Newton: 2:44 And I for one will be sad to see it go because I think it was sort of the funniest clause in the entire AI world, right? It was like basically like, ‘Well, if we ever get to a point where OpenAI says the magic word then the entire world changes.’
Kevin Roose: 3:00 And now they’re not allowed to say the magic word anymore. Right. So AGI has been poorly defined for many years and everyone’s got their own definition but it did have this one interesting like contractual stipulation and now even that is off the table. So now we just have sort of AGI as evaluated by vibes. So Casey, what did you make of this loosened partnership between OpenAI and Microsoft?
Casey Newton: 3:25 Well, I think it it seems like probably a good deal for both of them, right? Like there was a moment when it seemed for both companies like being very, very closely aligned was the best thing for both, arguably for a time it was. But with all of the various revenue, compute, and customer needs that both of these companies are now trying to serve, I think it’s benefiting both of them to play the field and sign other partnerships. So my my read on this was like this is basically good for both of them. But what about you?
Kevin Roose: 3:57 Yeah, I think it’s good for both of them. I think it’s a little better for OpenAI. They got most of what they wanted. I think the bigger deal for them is the ability to work with other cloud providers. So now they can work with Amazon or Google Cloud Platform and big corporate customers who use those cloud platforms can now use OpenAI models. They don’t have to go to Azure to do that. I think that allows them to strike these other bigger deals and to reach other corporate customers who may have been limited before by the fact that like it’s really hard and annoying to change cloud providers.
Casey Newton: 4:29 Yeah. And speaking of big deals Kevin, they signed what seemed like a pretty big one with Amazon this week.
Kevin Roose: 4:35 Yeah, so OpenAI wasted no time in its new open marriage with Microsoft. It went back out on the market and found themselves in bed with Amazon.
Casey Newton: 4:43 They found themselves in bedrock with Amazon. But I’m sorry, either way, go on.
Kevin Roose: 4:46 Yeah. So on Tuesday OpenAI and Amazon announced an expansion of their deal that they’d announced back in February that will allow OpenAI to sell its models through AWS’s Bedrock AI platform and make Codex, its coding model, available on Bedrock as well. OpenAI and Amazon will reportedly also develop customized models to power Amazon’s consumer facing applications and Amazon will invest fifty billion dollars in OpenAI. So there’s some interesting stuff here. I think the interesting subtext to me is that Amazon for a number of years now has been pretty closely tethered to Anthropic as its primary sort of frontier model developer and so OpenAI is kind of taking advantage of its newfound freedom by trying to elbow into Amazon and maybe displace Anthropic as their favorite model provider.
Casey Newton: 5:43 Hmm. Well, I know that Amazon was talking a really big game about this deal. The CEO of AWS was giving interviews essentially saying like, OpenAI belongs to us now. It was kind of a the boy is mine situation. Remember the old Brandy and Monica hit from back in the day? This was kind of bringing that back in a little bit more… of an AI flavor. I also have to say I find it very interesting that Amazon named its platform Bedrock, because that’s where the Flintstones are from. It seems rather backward looking for a leading AI company, Kevin, wouldn’t you say?
Kevin Roose: 6:09 That’s great analysis. Thank you.
Casey Newton: 6:11 Thank you. Thank you.
Kevin Roose: 6:14 Yeah.
Casey Newton: 6:15 By the way, like, I think that this is a really important point and a reason that we are, like, talking about it to a big general audience, which is that the story that you just described, Kevin, is one of a world where no one has the resources they need to serve the demand for AI that they have. And I think, you know, at a moment where we’re, you know, still sort of seeing a lot of skepticism, there’s so much bubble talk, I just want to, like, posit that as a really important point in understanding what sort of bubble this is, because even the biggest companies do not have the resources that they need to serve that demand.
Kevin Roose: 6:52 Yeah, and I think that’s a good point, and it’s really a profound shift in the way that skeptics have been talking about this, uh, this AI boom. I remember just even a couple of months ago, the leading sort of strain of criticism was that these AI companies would never be able to generate the demand to pay for all of the expensive data centers and infrastructure projects they wanted to do. And now that’s shifted to, well, there’s so much demand, what if they can’t build enough to support the demand they have?
Casey Newton: 7:28 Yes, and on that front, it does seem with at least one of these mega building projects, there have been some problems recently, Kevin.
Kevin Roose: 7:34 Yes, this was another story that hit this week. The Financial Times reported that Stargate, OpenAI’s joint $500 billion infrastructure project, is also undergoing a bit of a shift. The FT reported that in recent weeks, OpenAI has halted planned data centers in the UK and Norway, declined to expand its flagship site in Abilene, Texas, and seen several senior figures tied to Stargate leave for rival Meta. The FT further notes that OpenAI has shifted to leasing capacity from third parties instead of building out all of their own facilities. Casey, what did you make of this?
Casey Newton: 8:13 I think this was a case where, like, reality has just finally intruded on the Stargate project. Like, when all of these deals were getting announced initially, this is how they sounded: “Well, we’re going to spend one betrillion dollars that we don’t have to build 40 quadrillion data centers.” And at the time, people said, “That kind of seems like a lot. Can you guys actually live up to that?” They said, “Uh, yeah, just watch us.” Uh, well, guess what? They couldn’t, and now they’re changing course.
Kevin Roose: 8:39 Yeah, I don’t think this signals that they are sort of retreating from their compute ambitions. I think it’s more about, like, they are realizing that if they want to go public, which they do, they need to sort of get their house in order. And one way to get your house in order is to move some of this data center and infrastructure building off of your balance sheet and onto third parties.
Casey Newton: 8:59 Yes, but there is… One point in there Kevin that I do want to ask you about which is that Berber Jin at the Wall Street Journal over the past week had this really interesting story where he said that OpenAI had failed to meet some of their internal user number targets and some of their revenue targets and that this was possibly creating some tension between Sam Altman and his CFO Sarah Friar as they consider potentially doing an initial public offering later this year. So curious what you made of that story and does this maybe help explain why OpenAI has had to pull back on some of its big Stargate ambitions?
Kevin Roose: 9:34 Yeah I mean I think there are competing forces within all of the big AI companies right now. One side is sort of the the indefinite optimists, the people who think that demand for AI is just going to be essentially infinite and that as much compute and as much money as they need to spend acquiring compute it will all be paid back many times over because the world is about to change into something most of us barely recognize and so kind of just trust us on that is sort of one camp. And then there are the the sort of you know the number crunchers who are trying to fit all of this into a kind of financial projection that will make sense to investors who are not as convinced that the world is about to change forever and who want to see things like what is your plan for actually making the revenue that you’re going to need to pay for all this stuff. So I think this is happening in a way at OpenAI that is now because of Berber’s story is out there but I think this this kind of tension exists at all of the big AI companies and so I think right now what we’re seeing is kind of that that power struggle breaking out into the open.
Casey Newton: 10:51 Yes. And for what it’s worth OpenAI did call this story prime clickbait which I think just refers to clickbait that’s really really good is that what that means?
Kevin Roose: 10:58 Yes it’s sort of like Wagyu clickbait.
Casey Newton: 10:59 Yes exactly.
Kevin Roose: 11:00 This clickbait was dry aged for a month before it was served and it’s delicious. Yeah and I think one thing I want to flag on this is that these growth projections that OpenAI reportedly did not hit those were in 2025. I think it is fair to wonder if something has changed in just the last few months because of the enormous rapid growth of tools like Codex and Claude Code. We have seen just reports of astronomical growth in those tools so it may be that OpenAI was having some growth issues late last year but that because of this agentic coding boom things have started to turn around. We just don’t know yet.
Casey Newton: 11:41 Yeah that that makes sense to me and it does seem like their Codex app in particular was really well received. But there’s been this other transformation that seems to be unfolding Kevin this week. The Information had this really interesting story where apparently OpenAI projected at the start of the year that its eight dollar a month subscription which is called Chat… GPT Go, which sort of, you know, gives you a little bit of the good stuff, but not as much as if you’re paying $20 or more for ChatGPT. They predicted that its Go subscriptions would grow 36 times this year to 112 million people, while meanwhile, its $20 a month Plus subscriptions would fall 80% to about 9 million. So, that’s like a really interesting business pivot that I would love to know more about. Of course, it sounds a lot like the new Netflix plan that they rolled out a while back, right? Where it’s sort of like, well, you know, it’s going to be a lot cheaper, but we’ll show you ads. I was curious like what you make of that strategy because, you know, part of me feels like, well, they’d much rather have, you know, the $20 subs than the $8 subs, but maybe there’s just a lot more of those $8 subs out there.
Kevin Roose: 12:44 Yeah, I think what’s happening here is that the market is essentially splitting into two, right? There’s the sort of casual hobby users who are using AI chatbots like ChatGPT, like Claude, for sort of souped-up Google queries, to, you know, help them write emails, and maybe only using it a couple times a day. And if you’re doing that, you probably don’t want to pay 20 bucks a month. You’re probably more comfortable paying eight bucks a month, or maybe you don’t want to pay anything at all and you just rather use the free ad-supported tier of all of this stuff. And then there’s the professional users for whom this is worth way more than 20 bucks a month and who are willing to pay many multiples of that to get the access to the latest models, to have higher rate limits. And so I think all of the companies now are sort of, you know, doing this kind of experimentation with how much can we charge the professional user without losing them to a rival company, and how cheap can we make the kind of lower-end subscriptions or the free tiers so that people who are more casual users won’t be tempted to go use Google instead.
Casey Newton: 13:52 That makes sense. I’ll say for my part, I’d be willing to pay even more for ChatGPT if they would just let the Codex app talk about goblins.
Kevin Roose: 14:01 I say free the goblins! These models are so weird. Like, it is so weird that we have this technology that is now sort of load-bearing infrastructure for the entire economy, that every business is using to completely reinvent the way that it works. And that out of nowhere, if, if not specially restrained, it will just start talking about goblins.
Casey Newton: 14:24 Which to me is just like a satire of the AI safety conversation. You know, like lately OpenAI has sort of been very like skeptical of the AI safety and casting a lot of aspersions on doomers. But it’s like, well, we did have to add safety guardrails to prevent goblins from taking over our coding app. And that’s a real story. So, as usual, I’m just loving life here in 2026.
Kevin Roose: 14:45 What a world.
Casey Newton: 14:47 So, those are a bunch of stories about OpenAI’s strategic pivot, its reset. But there is this other big variable here, this potential fly in the ointment, and that is the long-awaited…
Kevin Roose: 15:00 Elon Musk trial uh that got underway in a federal courtroom in Oakland this week. Casey, can you remind us what this case is about?
Casey Newton: 15:10 Yeah, so Elon Musk was famously one of the co-founders of OpenAI. He gave the company some of its initial funding, but left in a power struggle between himself, uh Sam Altman, Greg Brockman and some others, and a few years after all of that went down, and notably after Elon started his own AI company, he sued OpenAI and said, ‘I have been defrauded, this was only ever supposed to be a non-profit and you’ve gone and turned it into one of the world’s most valuable companies through its for-profit arm.’ Notably Kevin, he made 26 claims when he originally filed this lawsuit in 2024, but only two have survived to trial: unjust enrichment and breach of charitable trust.
Kevin Roose: 15:53 The trial is just getting underway, they’ve done jury selection and they’ve had a couple witnesses testify. Elon Musk himself took the witness stand and said quote, ‘This lawsuit is very simple, it is not okay to steal a charity.’ He also said that if OpenAI is allowed to get away with this, quote, ‘it will give license to looting every charity in America.’ Basically he is saying this thing that started as a non-profit, uh that was supposed to continue as a non-profit, uh became through uh some corporate restructurings a for-profit company that has raised many billions of dollars, and that if this is legal to do, every charity would do this. Why wouldn’t you want to uh take your donor’s money and turn yourself into a well-funded startup?
Casey Newton: 16:44 Yes, now one inconvenient truth that Elon Musk faces here, which is that OpenAI’s for-profit business is still controlled by a non-profit. There’s this foundation that houses the public benefit corporation, and while I do empathize with those who say, ‘Hey, it really seems like the non-profit hasn’t done all that much and you know most of their money is being used for for-profit activities,’ this was litigated and the non-profit, you know, still does have like voting control over the for-profit.
Kevin Roose: 17:12 Yeah, so Elon Musk is saying this is a case of looting a charity. OpenAI’s lawyers have accused Elon basically of uh just being bitter that the company has succeeded without him. Uh its lead counsel William Savitt said during the trial, quote, ‘We are here because Musk didn’t get his way at OpenAI. My clients had the nerve to go on and succeed without him, and Mr. Musk did not like that.’ They have also been pointing out that Elon had also wanted to make OpenAI uh have a for-profit subsidiary back when he was with the company and that he’s just mad that he didn’t get to control it.
Casey Newton: 17:45 Yeah, to underline that, like in 2017, 2018, there are emails from Elon Musk where he talks about uh turning this into a for-profit. So, you know, whatever concerns he had about looting the charity, uh you know today, like he did not have them uh back at the time.
Kevin Roose: 18:00 revealed in some of these emails, Tesla of course being a for-profit company. So seems like this is not exactly a consistent and principled stand. Right, he also wanted to fold OpenAI into Tesla, that was uh revealed recently as well. But Casey, what are the stakes here? Like, if Elon Musk does manage to convince a jury that this was a case of OpenAI looting a nonprofit for its own commercial gain, like, what could the remedies be? Could this be fatal for OpenAI or is this just sort of an attempt to slow them down and distract them with a big trial?
Casey Newton: 18:28 I think that it is much more the latter. Like, based on my reading of the case and what I’ve seen sort of legal experts say about it, the whole case is very unusual that it even made it to trial. Like, for the most part, if you donate money to a nonprofit, you actually don’t have a say in what happens to it after that. So it’s very unusual that the judge even granted him standing to sue here. And as I noted, she threw out most of his claims. That said, let’s say that, you know, there is some single-digit percentage chance of him winning something here. What he wants to do is to take more than 150 billion dollars that is currently under the control of the for-profit business and give that back to the nonprofit, which would create a lot of headaches and roadblocks for OpenAI as it tries to build out Stargate and do everything else it wants to do.
Kevin Roose: 19:21 Yeah, I think the lawsuit and this ongoing litigation between Elon Musk and OpenAI has been very distracting for OpenAI. But like as a journalist and as a person who wants to know more about the inner workings of how these companies run, I think it’s been actually very valuable for a lot of these emails and early communications between OpenAI leaders to be released as part of this litigation. I have found it very useful in understanding some of the early dynamics at OpenAI. And it also just illustrates the degree to which these projects are all just sort of fueled by grudges, right? Like, there’s one level of interpretation which is like all of these people are just like obsessed with building the machine god and that this is all sort of related to their visions of the future. And then there’s like another more base level, which is just like these people are all just rivals and they have these petty long-standing grudges and they just don’t like each other very much. And so you can interpret a lot of what happens in AI through the lens of personal animus.
Casey Newton: 20:22 Yes, I’ve said this before and it is rude, but a shocking percentage of the AI industry is just people who decided they didn’t want to work with Sam Altman and who now have their own companies.
Kevin Roose: 20:30 Right. So Casey, some people have been looking at all of this drama and intrigue surrounding OpenAI, from the trial to the Microsoft deal to these missed growth projections, and saying some version of like, OpenAI is in trouble. They are not going to make it to an IPO, they are going to sputter out and maybe end up in some real hot water, and maybe Elon Musk wins this trial and it’s sort of the end of OpenAI as… As we know it. What do you make of those gloomy predictions?
Casey Newton: 21:03 Yeah, I mean, look, there are some fundamentals for OpenAI that remain worrisome, right? They’re planning to burn tens and tens of billions of dollars in cash before they achieve profitability. They still have this very ambitious infrastructure build-out that is quite expensive. And so, like, I’m not going to sit here and say that, like, all of the numbers seem to pencil out for this company. On the whole, like, if I try to, you know, put myself into the shoes of their CFO and I look through all of the stories that we just talked about, I think, these seem like smart things to me. You know, it kind of seems like they’re starting to dot their i’s and cross their t’s and get this company in a shape where retail investors will be excited to invest in the stock, which, by the way, I think they will be. Um, so, yeah, it’s one of these companies where, like, it is a generationally weird enterprise, but when I look at this particular set of stories, I think, I think they’re basically doing the right thing. What do you think?
Kevin Roose: 21:53 Yeah, I mean, I think there’s this interesting fallacy in the AI industry where it’s like there will be only one winner, right? Everything is zero-sum. If OpenAI is having a bad month, it’s because, you know, Anthropic is having a good month or Google DeepMind is having a good month and vice versa. Like, their sort of growth comes at the expense of all the others. And I think that’s a… that feeling is shared by, among others, the executives of these companies. But I just don’t think it’s true. Like, I think that there are going to be a handful of companies that are just going to kind of rise and fall together, right? That if your models are in the sort of top tier, you are going to be fine as long as they stay in the top tier and the sort of rising tide of AI adoption will sort of lift all boats. That’s more my feeling.
Casey Newton: 22:45 Well, will this rising tide lift all AI podcasts as well, do you think?
Kevin Roose: 22:50 I hope so.
Casey Newton: 22:51 I hope so. Okay. Me too. Kevin, is there a doctor in the house?
Kevin Roose: 23:00 There sure is, Casey. Today we are going to have a conversation with a doctor about AI and medicine because this is an area where there has just been a lot happening recently and we needed someone qualified to come in and debrief us.
Casey Newton: 23:18 Yeah, you know, as we’ve sort of looked across the landscape just over the past few months, we’ve seen company after company introduce their own product at the intersection of AI and medicine. There’s ChatGPT Health, ChatGPT for clinicians, Amazon has something called Health AI, Microsoft has Copilot Health, and of course, all the while doctors are experimenting with this technology and as best as we can tell, actually getting really excited about what they’re seeing.
Kevin Roose: 23:46 Yeah, and this has been a huge change in my recent visits to doctors, which is that I now am having this series of conversations leading up to the visit with AI systems about what is going on. And so I am coming armed with what I believe to be good information about what is going on and that allows me to sort of have a different more elevated conversation with the doctor. And this is not just me, like people are increasingly… Increasingly turning to chatbots for medical information, according to some recent data, approximately a third of Americans report turning to AI for healthcare information. And companies are racing to respond to that demand by making better tools that are specifically designed for use in healthcare. So to help us make sense of the landscape for AI in medicine and healthcare, we’ve invited back to the show one of our favorite doctors, Adam Rodman. He’s an internal medicine physician at Beth Israel Deaconess Medical Center and an assistant professor at Harvard Medical School.
Casey Newton: 24:33 Yeah, we last talked to him in November of 2024 and since then, he has continued to study the way that people and AI interact in the healthcare space. And we have a lot of questions for him, like what should we do about your rash, Kevin?
Kevin Roose: 24:45 Yeah.
Casey Newton: 24:46 Yes. So let’s fork over our copays and bring in Dr. Adam Rodman.
Kevin Roose: 24:53 Dr. Adam Rodman, welcome back to Hard Fork.
Adam Rodman: 24:57 Oh, it is a pleasure to be here. Am I a friend of the show at this point?
Casey Newton: 25:00 Well, let’s see how this interview goes.
Kevin Roose: 25:02 You’re at least a doctor of the show. You are our primary care physician. So when we last talked to you in late 2024, I think this was a moment where the medical community was starting to say, wait a minute, these AI models are getting pretty good at things like diagnostics, but I think a lot of the field was still kind of in wait-and-see mode. Now, almost two years later, we have a lot of new tools and a lot of new studies about the use of AI in medicine. So just catch us up on, like, what has been going on with AI in medicine for the last call it year and a half.
Adam Rodman: 25:36 Yeah, it’s been crazy. AI in medicine has gone from, well, depending on how you measure it, it’s probably the fastest adopted medical technology of all time. We went from this being super novel, almost no one used AI tools to just being a routine part of most doctors’ weekly practice.
Casey Newton: 25:57 And give us a sense of, like, the AI stack for a doctor. What are the tools that they are using right now and how? And particularly, what are the mainstream doctors? Like, the people that, you know, aren’t yet on the bleeding edge.
Kevin Roose: 26:07 The normies.
Casey Newton: 26:08 Yeah, if you will.
Adam Rodman: 26:10 Yeah. So, the biggest sort of normal doctor technology, which most of your listeners or a good portion of your listeners have encountered, are what are called AI scribes. That’s a sort of voice-to-text algorithm that listens to you talk to your patients and then writes a first draft of your note. And these have gone from, like, kind of a novel experimental technology to commodity in probably less than two years. They’re everywhere. Doctors really like them, and then patients really like them because they spend more time talking. And then the second sort of normal doctor use case is for decision support. So there’s this one company called Open Evidence that has created a free tool that has gone from, again, zero to crazy numbers of adoption. I will tell you, younger doctors like my residents use it all the time. And that is just… It’s funny because it’s a BYOAI, right? We always have usually the hospital or the health system buys tools. This is a bring your own AI to work day and I don’t know the actual numbers, but it’s probably close to half of US doctors are using this right now.
Kevin Roose: 27:17 Wow. So yeah, the statistics that I’ve seen are that more than 40% of doctors now are using this, which is pretty crazy uptake for something that was just started a couple of years ago back in 2022. In March, Open Evidence reported that in a single 24-hour period, doctors consulted the AI system a million times.
Casey Newton: 27:40 I’ve been fascinated by Open Evidence. I’ve never used it myself, but I have friends who are doctors or nurses and they have said what you’ve said that basically just everyone, especially on the younger end of medicine is just using this thing constantly. So like give us a sense of like how this Open Evidence tool works, like what situations is it used for and what are its strengths and weaknesses?
Adam Rodman: 28:00 Oh, that’s a great question. So, how Open Evidence works, like all of these tools is a trade secret, but it uses some sort of like retrieval augmented generation and an evidence retrieval tool. And they have all these deals with the big medical journals. So, New England Journal of Medicine, JAMA, and when you ask a clinical query, it searches the evidence and then tries to identify high quality sources and then it always grounds what’s coming back in the literature. You have my hair’s grayish. So you have gray hairs like me who kind of use Open Evidence the way that I would use a Google search or one of the old tools. So I use it as a souped-up way to search the literature rapidly and often go to the primary sources. Or I use it as a faster way to get a reference. So a drug that I haven’t dosed in a long time, Open Evidence pulls the drug monographs from the FDA. I can very quickly pull that up. Younger doctors, I have noticed, and I don’t know this empirically, but younger doctors are more likely to ask questions like what could be going on? Can you give me a second opinion? What is the next thing that I should do? So, ways that I don’t traditionally use decision support or reference tools, but sort of a new way. And of course younger doctors also use it in the reference ways that I do. If that makes sense.
Casey Newton: 29:14 Now are they actually uploading patient data to this? Like or are they just sort of describing patients in sort of generic and anonymized ways to get back some decision support?
Adam Rodman: 29:29 My understanding is largely number two. I’m sure the company has a good sense on how many people… I hope no one is copying protected health information and putting into it. Certainly what I’ve observed from like my colleagues and my students, most people use it the way you would use like a search tool when you have a question, which is, hey, like I’m giving this person ceftriaxone. What’s the right dose for an intra-abdominal infection? So sort of generic questions that are being interpreted through the physician.
Casey Newton: 29:56 And are there any AI tools that are integrated with patient health records? This has been an area…
Kevin Roose: 30:00 One area where I think there’s just been a lot of pushback of but like, “I don’t want my personal health data, my protected health data going into one of these cloud-based AI systems,” but are there hospital systems or medical systems that are bringing this stuff directly into contact with patient data?
Adam Rodman: 30:15 Oh, 100% yes. Right now most of the sort of in-contact with patient data are less about physician-facing decisions and more about like billing. There are companies that are like integrating with the electronic health record. Those are not standard yet. And then the EHR companies themselves, so like Epic is obviously the biggest EHR vendor in the US, they’re doing a lot of work on building in native things. So for example at my health system, if I want to send a message to the patients, the helpful AI at the top already has like, “maybe you should say this.” It’s usually not that helpful and I don’t think I’ve used it once in my life, but there are a lot of those things that are being experimented on actually built into the patient’s health data.
Casey Newton: 30:58 I’m curious how doctors are feeling about all of this. We saw a survey from the American Medical Association that found that more than 80% of physicians now report using AI professionally. Is that physicians racing out and grabbing these tools and bringing them to the office because they’re so helpful? Or is this the classic case of a CEO saying, “hey, you gotta use AI or you’re out of here”?
Adam Rodman: 31:15 So doctors are BYOAI. A lot of that AI use is AI scribes and decision support software. And I’ll tell you, some people are just using straight up like ChatGPT or Gemini or Claude for their decision support software. So I think one of the reasons doctors thus far have been more positive about it than perhaps the overall population is they’re largely tools that doctors are bringing themselves that they think make their lives better and at least not yet many things that are being imposed upon us.
Kevin Roose: 31:32 Yeah. I’ve noticed that when I and my friends go to the doctor now, we often are presenting our information to a chatbot first and then coming into the doctor with sort of a readout of what the chatbot has told us. This is of course not a new phenomenon, people have been doing this like with WebMD results for many years. But is this something that you’re seeing now? Is that many more patients are coming to you having already discussed whatever’s going on for them with a chatbot?
Adam Rodman: 32:06 Yes. This is the other big change is that there’s, you know, someone else in the exam room with me and often it’s ChatGPT. They’re talking, sometimes with my hospitalized inpatients, they’re talking to ChatGPT while I’m in the room with them. And I think this is… it’s interesting because this is kind of a new competency for doctors. We have to talk to our patients about AI. And I have started to talk to my patients about what I think are like safe uses, what are like safe uses while telling me and then things that they definitely shouldn’t do. My patients may talk to me more about it because I am like a doctor and an AI researcher, but like a lot of my patients are using AI routinely.
Casey Newton: 32:38 Well, give us a flavor of what you’re telling them because you know, I am definitely somebody who has looked up my symptoms before I’ve gone to the doctor and I would say I’ve found it enormously helpful. But I can also imagine, you know, more skeptical doctors being annoyed, you know, at a patient telling them, you know, what ChatGPT says to do. So—
Adam Rodman: 33:09 Yeah, so here’s my—I’ll give you my spiel. This is—I give them a, what is it, a green light, yellow light, red light. I don’t know why I’m doing circles. Everyone knows what streetlights look like. So the green light uses are general health questions. So, you know, I’ve recently diagnosed with diabetes. I really love seafood. Can you help come up with a diabetic diet for me? The green light uses are also preparing for clinic visits. So I’m about to go see Dr. Rodman. I want to make sure that I ask the right questions. Here is the last note or the last thing he wrote—obviously strip out anything identifying. I don’t put your personal health information. Um, and like, help come up with a good questions to ask him. And then other, other green light activities might be like wearable data. I don’t know how good they are wearable data, but I will tell you if a patient is going to give me like five years of their Apple Watch data, they’re probably going to get a better from ChatGPT than from me pretending to look at five years of Apple Watch data because it’s a 20-minute visit. The yellow light, Casey, I think is a lot of the things that, that you’re saying. So I tell my patients it’s okay to explore new symptoms. It’s even okay to seek out second opinions when talking to a chatbot. That can really help prepare you, as long as you understand that it’s not a replacement for a doctor and that is the first step to talking to a human being. So LLMs are, they’re really powerful and I mean there is some evidence of course that, like, when any human uses them, you don’t always get laboratory-level performance and there can be—like they can give you dangerous advice. But diagnosis and like exploring symptoms, as long as you use it in a way to prep to see your doctor can be very helpful. The red light, what I tell them never to do, is like, ask medical management decisions. Like don’t say my doctor said to do this, is this right? Like I have cancer—God forbid you have cancer—is this the right chemotherapy option? Like a lot of those decisions are so nuanced, take in so much information, those are things that the models don’t do well and they’re so sycophantic, they can convince you that they’re saying the right thing even when they’re wrong.
Casey Newton: 35:11 Right.
Kevin Roose: 35:12 I’m curious Adam, out here in San Francisco there are all these fitness people and health maxers, people who love to track themselves using all manner of devices and people are getting these full, full-body workups from companies like Function Health that are, you know, sort of premium concierge medicine things and they’ll, you know, get a hundred labs done and then they’ll upload all that data into Claude or ChatGPT and just sort of treat it as a sort of first-line medical professional in their lives. Do you think that is a good practice or is that just making people, you know, way too worried about things that maybe they don’t need to be worried about?
Adam Rodman: 35:51 Yeah, so that’s making people way too worried about things they don’t need to worry about. And this is ChatGPT, LLMs in general. I mean, the dark side of talking to— an LLM about your symptoms, if they are so sycophantic they can drive you into like the cyberchondria worry hole. The evidence is not there yet that these sort of large routine um testing functional medicine and putting it into an LLM does anything to improve health outcomes. Now, if your LLM is telling you to to work out um and eat healthier, that’s probably pretty good. Sleep.
Kevin Roose: 36:23 Yeah. What about the integrations like ChatGPT Health, which lets you sort of convert your Apple Watch or Fitbit data into something that ChatGPT can analyze? There’s also a new version of ChatGPT-4 for clinical use called ChatGPT-4 for clinicians. Um, are any of these integrations or or projects more promising in your in your view?
Adam Rodman: 36:45 I- not yet, but I think it could be at some point. I mean, so ChatGPT for health, it pulls in your data from the medical record, um, using like an interoperability standard and lets you chat with your medical records. Now, reason number one for concern is privacy. That’s obviously going to have your entire medical history going to a AI company. It’s also going to not be redacted by you in a way to remove identifiable things. Reason number two, I think if we’re talking about health record data, it’s really messy. They include tabular data. They include copy-forwarded data that’s been copied and pasted. And they also, if you’ve ever read your health records, they include things that are wrong. Um, there’s a lot of errors or misdocumented things in your health data. And it turns out that just copying a bunch of information, like LLMs aren’t magic. You can’t just copy your entire medical record in and think that you’re going to get good performance. And I would never bet against the technology. I think that we will get to the point that we have ways to build representations of of humans and understand their health. But right now, there’s like no advantage to just dumping everything in an LLM, which is what ChatGPT for health theoretically would allow you to do in a way that would allow you to better understand your health.
Casey Newton: 37:59 I’m curious if you saw this trial they’re doing in Utah, where you can use an AI agent to autonomously renew prescriptions for almost 200 routine drugs. Um, there’s apparently some human review, but um, mostly this is automated. Um, is that a good idea, bad idea?
Adam Rodman: 38:15 Well, so globally no, we should not be uh having LLMs write prescriptions for people. Um, the- the trial- the trial in Utah in particular is a refill. So a doctor has already written a prescription within the last 12 months. And I get the idea that it saves the primary care doctor time from having to review and refill. I’ll tell you if you talk to most doctors, yes, it is annoying to get refill requests. No, that is not the thing that drives us crazy. This is not like a use case that we’re screaming for. Um, I think it’s being done as a proof of concept of- of can this work in the real world. This trial in and of itself is not dangerous. I- prescription refills and I think there’s no opiates, there’s no dangerous drugs in it, and a doctor has to have written the original one. But even if it does work in this… That does not mean we should be having autonomous AI systems write new prescriptions. That is not safe and not a good idea yet.
Casey Newton: 39:07 See, I think this is a case where like this is sort of rent-seeking behavior on the part of doctors or doctor organizations. Like when I have gotten refills for prescriptions, I meet with a doctor for, you know, six to eight minutes. They say, ‘How’s it working?’ I say, ‘Great.’ They say, ‘Are you having any side effects?’ I say, ‘No.’ They say, ‘Okay, I’ll write you a refill.’ And the whole process just seems totally designed to like get me to pay up for another doctor visit and not give me any actual good medical advice. So if I could play devil’s advocate, like do you think that there that the sort of resistance to programs like this are motivated by just wanting to keep people coming to the doctor and paying for those visits?
Adam Rodman: 39:54 So first, aren’t most of your prescription refills just done as in you call the pharmacy and they send an automated thing to your doctor and they click the yes button and you never talk to them?
Casey Newton: 40:05 No, for some they make you actually do an office visit and maybe they want you to, you know, take your blood pressure again or whatever.
Adam Rodman: 40:13 So I’ll do the devil’s advocate back. Uh, let’s say I prescribe a fairly common antidepressant and they wanted to be refilled. What what I don’t know is that this patient may be the the silly question you get in the clinic may be new lesions forming in your mouth. And it’s an early ulcer. And if we don’t pick it up within 24 to 48 hours, you may develop like Steven Johnson syndrome, so potentially life-threatening complication. And the reason there are certain types of drugs, including anti-hypertensives, is that they can be high risk and we need follow-up. Now is that everything? No, and definitely there should be more things over the counter. I like I don’t think that most doctors are sitting around saying, ‘I wish I had more medication follow-up visits.’ And the reason some of these things exist is that there can be very dangerous symptoms.
Casey Newton: 40:54 Yeah, so keep going to the doctor Kevin, we can’t have you developing those lesions. You’re too important to the show.
Kevin Roose: 41:00 Let me ask you about another one. This one actually seemed like just an unqualified good. Um, the Mayo Clinic announced this week RedMod, this AI system that identified subtle changes in routine CT scans up to three years before a pancreatic cancer diagnosis. And this was like many, many, many percentage points um better in detecting pancreatic cancer than human beings. So to me, this is like the sort of thing I keep waiting for AI to do and it seems like it’s actually doing it and of course that’s very exciting with something like pancreatic cancer which is notoriously difficult to detect and has like very low survival rates.
Adam Rodman: 41:37 Yeah, and this is so like completely out of the discourse of like autonomous AI agents, there’s really exciting stuff happening. So the Mayo Clinic, there’ve been some great studies on breast cancer detection. Uh a lot of these algorithms have gotten so great that they’re able to identify breast cancer better than I shouldn’t say better than people, but in a workflow that has a good detection rate. And then in picking up like potentially cancerous polyps when you get a colonoscopy. So there’s some really great stuff happening. a lot of exciting and really positive things that are coming and I’m— I mean, at the end of the day, we’ll need to see how Red Mod works in the real world and a trial, but I’m— I’m really optimistic about that sort of technology.
Kevin Roose: 42:11 Hm. Do you think that if AI meaningfully extends life expectancy for— for people, it will be because of new AI-discovered drugs or because of changes to routine healthcare that are made more efficient or more accurate by AI?
Adam Rodman: 42:27 Number two. I think that when you talk about AI drug discovery, even the part of the pipeline that’s so difficult is not necessarily the coming up with the new compounds. It’s running the clinical trials and getting it through the regulatory process, which can probably be sped up, but not as much as the discovery. Uh, you know, if we get this right, there’s so many people in the US who don’t have access to a doctor, who don’t have access to very basic medications, who can’t control their diabetes because of lack of access. And I’m really hopeful that if we do this wisely, we can, you know, get people more access to care, which— doing my knock on wood— hopefully will improve health outcomes. So like all of this, I think the potential uh, benefits are like less exciting. They’re— they’re getting people, more people the bread and butter and getting, you know, more people to have less heart attacks, more people to have less strokes, more people to get their cancer screening and not necessarily like, ‘Oh, we cured aging with some sort of new AI-discovered CRISPR technology.’
Casey Newton: 43:27 Are you at all surprised though that we haven’t yet seen the— the first like AI-discovered wonder drug?
Adam Rodman: 43:32 I mean, the biggest wonder drug of my career has been the GLP-1s, which uh, was— I started using it when I was a resident, so we’ve had it for a really long time, and we had to like repurpose a drug for diabetes. So no, I’m not. Like, uh, medicine and science is— is just kind of messy and it’s— there are always those stories about like, you know, we discover something amazing and make it into— like penicillin. But even penicillin took like 20 years to get into human beings. So no, I— I think we will see AI-discovered drugs. I think it’s just the benefits from AI are gonna be like the benefits from medicine. It’ll be a lot less exciting than people think, but still important.
Kevin Roose: 44:08 So there’s a lot of worry right now about sort of AI in schools, in education, the— some of the cognitive atrophy that people are worried about— ‘Oh, if we start using AI to do all of our work, we’re not gonna have the basic skills.’ Is that something you’re worried about for like recent medical school graduates where maybe they would have had to hold all this stuff uh, in their brains a few years ago and now they can just ask a chatbot and maybe that’s gonna erode some of their skills as a physician?
Adam Rodman: 44:38 Yes. So that is the biggest worry that I actually have about sort of the short to medium term is the de-skilling of the workforce. We have some evidence— there was a sort of scary study last year from Poland on a trial where they gave doctors not a language model, but a polyp-detecting technology. Uh, and they— they looked at their ability to detect polyps, so potentially cancerous lesions in the colon before using it and then after using it for— three months. And when not using it, their ability to detect polyps dropped by six percentage points. So these are skilled doctors using a technology and they lose six absolute percentage points of their ability to detect potentially cancer in three months. And then imagine that you’re learning to do it for the first time. Will you ever gain those skills? So like at Harvard Medical School, like this is and medical schools I think everywhere, this is our big worry, which is how will this affect us to train the new generation of doctors. And it’s like every other field, like you talk about debugging code, in order to become a new doctor, you go through all this training because you need to make mistakes and you need to have someone above you who knows what’s going on so those mistakes won’t hurt patients. And that’s just how education works and it’s— this threatens that.
Casey Newton: 45:44 I mean, it’s interesting though because it’s like, you know, it’s probably true that because I had access to a graphing calculator, like if you took it away from me, I’d be worse at like plotting parabolas on a graph. But the solution to that is that I just keep using the calculator, you know? So like I’m not sure how big of a problem this really is.
Adam Rodman: 46:03 I’ll— I’ll also say there’s something— there’s something deeply ingrained in human society that middle-aged people complain about young people, so I think whenever we talk about de-skilling, we have to keep that in mind.
Kevin Roose: 46:14 Yeah, I mean, for what it’s worth, like, I want my doctors to be using AI models. I want them to be consulting the hive mind before they weigh in on my specific condition. It doesn’t threaten me as a patient to know that they are using Open Evidence or something similar. But I’m guessing for a lot of people that would seem strange and maybe there are some physicians who don’t advertise how much they’re using AI because their patients might think less of them. Do you think that’s happening?
Adam Rodman: 46:36 Oh, yeah. I think there in certain situations, in certain places, I bet there’s social pressure to, um, say that you’re not using AI. That there’s some ego on the line. Um, I— I don’t see that, but again, I’m an AI researcher, so I don’t think anyone would say that to me.
Casey Newton: 46:46 To me it just seems weird to like, to hold as your standard for what makes a good doctor that they have memorized like a maximum amount of material. Or like that’s basically what we’re talking about. It’s sort of like, you know, the taxi drivers in London that have to like learn every single street and like hold them all in their heads. It’s very impressive, but I’m fine with them using the GPS.
Adam Rodman: 47:11 Yeah, and I think it’s less about— so it is about memorizing, it’s more at this point, right, with where AI is now, it’s more having sort of that knowledge and call it wisdom to know when the system might be suggesting something wrong, which is something that right now, and this may change, we get by seeing a lot of cases and reflecting on them. So right now, you’re going to get the best performance if you have an experienced human trained in the old-fashioned way with an AI system. But I think your guys’ point is at some point that might not matter, right? The AI systems might just outperform all of us, and then yeah, I guess it’s like just use the graphing calculator. But we’re not there yet.
Kevin Roose: 47:50 Would the AI models be better if we were less protective of privacy for medical data?
Adam Rodman: 47:57 I mean, that’s a— that’s such a loaded question. So the first thing that I’m… The first thing I want to say before I answer that is patient privacy is very important and we should respect people’s privacy and their ownership over their data, but yeah. So, in short, like the reason they’re not better at certain things is that you need to to get LLMs better, you need to label and then train them on the sort of labeled health data and there’s all- in the US, there are appropriately many restrictions on how health data can be used. I suspect that these companies like OpenAI, by having ChatGPT for Health, they will gain some more of their own data, which they say they’re not going to train on. I trust that they’re not going to train on it, but they’ll be able to use that data to at least evaluate their models and try to make them better.
Kevin Roose: 48:39 I think they should train on it. I mean, obviously that’d be a huge, illegal violation of privacy. But it would also make the AI doctors better. And I think, you know, a lot of people would be sort of willing to make that tradeoff. So I at least think there should be a little checkbox when you go to the doctor that says, like, ‘I’m okay having my personal health data used to train AI models.’ I for one would.
Adam Rodman: 48:59 Yeah, much better. Yeah. Yeah, in exchange for like 30% off your giant medical bill.
Kevin Roose: 49:06 You get a coupon. You’d get a coupon, you know, you’d get like your next Ozempic shot is 20% off.
Adam Rodman: 49:11 Is on the house.
Kevin Roose: 49:12 Exactly. Well, that’s a good place to leave it. Dr. Adam Rodman, thanks so much for coming back and uh keep us posted on what is going on in medicine.
Adam Rodman: 49:21 My pleasure, guys. Thank you very much.
Kevin Roose: 49:23 Thanks, doc. Well Casey, usually on this show we are talking about the future, but today we are going to take a trip back to the past, specifically to the year 1930. What was happening in the year 1930?
Casey Newton: 49:40 Oh my goodness. Well, of course, we were in the middle of uh the Great Depression. Um, my grandmother had recently turned 11 and uh was looking forward to getting her first store-bought dress in just a few years.
Kevin Roose: 49:51 So, this is a new language model, a vintage LLM called Talkie, and it is trained exclusively on data from before 1931. This is a research project built by three guys, David Duvenaud, Nick Levine, and Alec Radford, the lead author of the GPT-1 paper, a former OpenAI researcher over there. And this is a fascinating project that has been burning up my timeline this week because this is an experiment in what happens if you only feed a large language model data from before a certain cutoff.
Casey Newton: 50:30 Yeah, and obviously there are a lot of, you know, kind of character-based chatbots on the internet that will give you the experience of talking to somebody from the past. But what makes this project different is that they tried to limit themselves to training data from that time and before. The hope was that it would avoid any kind of contamination from what came after. And as you’ll hear, they have some really interesting and potentially useful ideas about what this kind of LLM might one day be used for.
Kevin Roose: 50:59 Yeah, so Casey you spent…
Casey Newton: 51:00 …this model?
Kevin Roose: 51:01 I have. I tried to ask it the most 1930s question I could think of, which was, ‘Say, what’s the big idea?’
Casey Newton: 51:07 What’d it say?
Kevin Roose: 51:09 It said, ‘The big idea is to popularize.’ And I said, ‘Popularize what, fella?’ And it said, ‘Popularize a sport.’ And I said, ‘I’m gay!’ So that’s kind of where we let that one drop.
Casey Newton: 51:20 And it said, ‘Gay? You’re happy?’
Kevin Roose: 51:23 Yeah, exactly. It said, ‘Your heart must be light, sir.’
Casey Newton: 51:29 Yeah, I love this experiment. I love like weird niche language models. My- one of my favorite language models of all time was Golden Gate Claude, which was this special version of Claude that was like pathologically obsessed with the Golden Gate Bridge. I would put Talkie in sort of that category of like an experimental research model that is maybe not all that useful on its own, but like helps illuminate something interesting and important about these language models and what happens when you train them in specific ways. So today, we want to talk with one of the creators of Talkie. We’re going to bring in David Duvenaud. David is an associate professor at the University of Toronto who researches AGI governance and catastrophic risk mitigation, and he is one of the co-creators of Talkie.
Kevin Roose: 52:05 And there’s really Duvenaud-better person to talk to about it.
Casey Newton: 52:07 That’s true.
Kevin Roose: 52:09 David Duvenaud, welcome to Hard Fork.
Adam Rodman: 52:12 Thank you very much, Kevin.
Casey Newton: 52:14 So this project, Talkie, is fascinating. It’s a vintage LLM. Explain why you and Nick and Alec made this thing.
Adam Rodman: 52:25 So this all started a year ago, where me and Nick were interested in forecasting, specifically, can we learn or teach machines how to forecast like 5 or 10 years ahead of time, like what is the big picture going to be, just because we have our own sort of pet ideas about what the future’s going to be, but we don’t think people should take our word for it, and we also don’t think that people should trust machine forecasts unless they have a track record going back like decades and decades. So the idea here is that if we could build a model who really only knew about the data, about the world up to a certain date, we could ask it to forecast like 5 or 10 years ahead of time, like ask it ‘what’s the New York Times headline going to be 5 years from now’ or ‘is there going to be another great war’ or something. And we could iterate and see like what kinds of things are predictable, what does it take, like how far out can things be foreseen. And then eventually, hopefully, we’ll have machines that have like a 100-year track record of forecasting, and then we can ask them, you know, in 2026, ‘what do you think is going to happen like 2 or 4, 8 years from now’ and we’ll have an idea of how much to trust those forecasts.
Casey Newton: 53:38 It’s a fascinating idea, but it strikes me it requires you to have like really good data. So in this case, really good pre-1930s data. I’m going to guess that was harder to obtain than just, you know, going out and like crawling Reddit or, you know, and everything else that the frontier models have access to. So how did you face that challenge and where did you get this pre-1930s data?
Adam Rodman: 53:59 Yeah, so it should be… You know, there’s a ton of groups doing a ton of awesome archival work here. The first data set we got excited about was institutional books, which was Harvard Library scanned like 1% of their entire collection. Hmm. And so they had tons of data from like 1800s, early 1900s. Like there’s a whole bunch of different groups doing tons of work. It- it would take a long time to enumerate this, but- and also like OCR has just gotten a ton better just in the last like six months even. And so there’s always been lots of projects to like automatically digitally scan this data, but it just hasn’t been very high quality until very recently.
Kevin Roose: 54:35 Hmm. And I assume that part of the reason you chose the cutoff date of around 1930 is because that’s when sort of works become public domain. Anything after that is copyrighted. Are there any other reasons you chose that specific point in time?
Adam Rodman: 54:51 No, that was entirely it. Is that we want to make everything public- publicly available and open source, and 1930s is just the sort of most recent date that has almost zero legal headaches with releasing data or- or anything like that. Yep.
Kevin Roose: 55:03 Hmm. So I’ve been fascinated by seeing like what people are trying with this model. People are having it make predictions but also asking it about its favorite authors or its opinions of, you know, major historical figures. What have been the experiments that have been most interesting to you?
Adam Rodman: 55:25 Uh, yeah, so I guess it has been really delightful seeing people have fun with the models and think of all sorts of questions to ask that we never would. Like that was one thing that I really wanted to make sure we made the- the chat available for people to try, because people have way more imagination than we do. Um, yeah, like the fun things that I’ve seen people do is- I mean a lot of people like to ask like, ‘What’s 2026 going to be like?’ And the model has sort of very philosophical answers about how like, ‘Well, we- we will have figured out that war is bad, we’ll much have like a much more peaceful civilization.’ Or sometimes it says like it’s the end times. I mean it’s a very inconsistent model and it’s not quite smart enough to really like, you know, think things through in a systematic way. It kind of just gives you vibes.
Casey Newton: 56:03 Now, that- that brings up a sort of interesting like wrinkle of like the kind of LLM this is, because if you were to ask like a frontier model today to predict the future, it would not only be trying to guess a statistically likely sequence of words, right? It would also be like doing some reasoning. Talky is not doing that, right? So that- that just sort of seems like by the way that it is built we would expect it to be less good at forecasting, uh, as- as the models we have today.
Adam Rodman: 56:30 Yeah, absolutely. This is a very baby steps model. The basic fine-tuning for reasoning and the scaffolds, like the super-forecasting scaffolds that we know just improve anyone’s reasoning, like you know, think of the different distinct possibilities and assign them each sub-probabilities. The model’s just not really smart enough to follow these kind of detailed multi-step instructions yet. So again, we just wanted to release the first thing that we did, but it’s- like a there’s a clear path to adding all these, uh, refinements.
Casey Newton: 56:52 So you do plan to add reasoning as you go?
Adam Rodman: 56:57 Oh, absolutely.
Casey Newton: 56:58 Okay.
Kevin Roose: 56:59 People have also pointed out that the model behind Talky— It seems to know about some things that it probably shouldn’t know about, like the rise of Hitler, the presidency of FDR, things that didn’t happen until after its data cutoff. Is that proof that there’s been some kind of contamination of the training data with more recent data?
Adam Rodman: 57:15 So there’s definitely contamination and this is sort of like one of the ongoing like things that we’re gonna have to just keep revisiting again and again and refining. So we have a classifier that tries to look for things that are anachronistic. And especially if you want to use this for forecasting or to evaluate forecasting, it’s really important that we really nail this issue. So we have all sorts of ideas for like scenarios and things that we think the model should just never assign any likelihood to. Like think of, I don’t know, Nagasaki and Hiroshima. Like before World War II, like those two towns would just never show up in the same sentence ever almost except for like some weird coincidence. Um, so you can just tell whether there’s been leakage about important events if the model just thinks that there’s any chance that you’ll see those particular names together. Um, so this anyway, like we’ve done we’ve made a bunch of efforts to avoid leakage. We know there’s leakage right now, so you shouldn’t use it to evaluate your forecasting scaffold yet.
Casey Newton: 58:07 How is it getting that data if it’s only being fed scanned OCR’d books from archival sources?
Adam Rodman: 58:16 Uh, because archival sources have wrong dates in them all the time. Or it’s kind of unclear what the like date of a text is because there’s like an updated edition or sometimes there’s like a preface that’s been added later or sometimes even just in the middle of the text there’ll be like someone inserted some future notice like, you know, note in like later on, like historians note that da da da da. And so it’s just really hard to check all these little edits that people make and then they still maintain the original publication date on the metadata.
Kevin Roose: 58:46 I see. I asked Talkie what it knew about me and it said Kevin O’Hara, which is not my name, was born in Dublin in 1840 and having been educated at the school of the Christian Brothers, became a teacher in it. He afterwards adopted the profession of journalism and was for some years connected with the staff of the Nation newspaper. It also said I had written several popular songs including Molly Asthore and The Irish Immigrant. Now obviously, most of that is wrong, but it did connect me to journalism, which I found interesting and maybe like some other evidence of some data contamination. But like is this thing accessing the internet in some way or like how would it have known that I or at least Kevin O’Hara, this character sort of connected to me in the model, was a journalist?
Adam Rodman: 59:41 You know, that’s a great question. I guess I’ll say the training data was like 240 billion tokens and it’s just this sort of this vast ocean of stuff. So like maybe there was a list of journalists that got put in somewhere that you had your name in it. I mean, I guess one thing about this model is it hallucinates like crazy. And you know, this was a huge problem with the chatbots that people were meant to use professionally and I think it’s been addressed to a large extent in like frontier models. But we made zero effort to address that in Talkie.
Kevin Roose: 60:00 any ever post-training so far?
Casey Newton: 60:02 Kevin, would you sing a few bars of The Irish Immigrant for us?
Kevin Roose: 60:05 Uh, you know, I don’t want to waste Adam’s time here.
Casey Newton: 60:09 All right, we’ll save that for later. Speaking of problematic content, some people found that Talkie gives racist responses to questions that are basically like, you know, would you let a black professor teach your child? Um, I could see how that might be historically accurate, but I’m curious if you anticipated it and how you feel about it.
Adam Rodman: 60:27 Yeah, so, uh, we, you know, it was also very clear to us that it would have these kind of responses. I mean, I guess I’m a professor myself and my sort of first instinct is, like, let’s let people see this if they want to and just don’t surprise anyone and don’t be flippant about it, because, you know, it really can be upsetting to some people and especially if we just, like, treat it, sort of insouciantly. So, the way we threaded the needle was we, we did zero, like, filtering of the data set for, like, problematic content or whatever. Right? We wanted to just, like, show what the actual sort of state of knowledge or state of thought was in the past. It would defeat the purpose of the project if we, you know, put our thumb on the scale. But for the public demo where you can talk to Talkie, we just had a modern model with like modern sensibilities just read every response and at the end, once it’s generated, if it is deemed problematic, just like slap a warning and say, like, ‘oh, this might have something upsetting, just like click if you want to see it.’
Casey Newton: 61:25 Right. The, the description I loved, this came from Gavin Leech today, was that Talkie is creating beautiful prose by a terrible person. Which is consistent with some of my tests, which is like, this thing actually does write quite well and actually, to my ear, like much more literary than some of the more recent models trained on more recent data. But, yeah, it is not—it is clearly the product of its time or at least the time of its data.
Adam Rodman: 61:58 Yeah, yeah, and we—I mean the prose is really cool because it’s—it’s like very refreshing style. And actually, if you feed it to one of the AI detectors, it usually says like 100% human, which is kind of funny. But then, I guess, as you mentioned, like a terrible person, I mean, it’s right now it’s kind of ends up being this sort of like average person and it—depending on—like it’ll just randomly answer with all sorts of different voices. Um, but that’s one of the next things we’re planning to work on is helping you talk to more specific people or in specific sort of states of knowledge or times and places, because that’ll, I think, allow us to answer more coherent questions than just talk to, like, the hive mind of 1930 or whatever.
Kevin Roose: 62:34 Speaking of the hive mind, I saw another person ask Talkie, basically, the person told it that it was from the future and would tell Talkie anything it wanted to know about the future. And Talkie’s first question was, ‘How did universal peace come about?’ Which was like the most heartbreaking thing I think I’ve ever read from a large language model. Like, what does that tell us about the time period or, or about the training data?
Adam Rodman: 63:00 Well, I guess I’ll say… In general futurism is a place where people don’t actually often try to predict the future very hard and they more they kind of project their values and if you ask someone what do you think’s going to happen they usually fill in something about what they hope is going to happen and I think that was also true a hundred years ago. So the the trick is to get Taki out of the like wishful thinking um mode and actually like into brass tacks of like what do you actually think’s going to happen?
Casey Newton: 63:21 Yeah, it’s just funny because I think if like, you know, if I could talk to an LLM from, you know, what, you know, 2126 today, I would just sort of be like, are the humans still alive? Like, what’s going on with the climate? Uh, did the robots, you know, how many people did the robots kill? Uh, it would just sort of be a very different set of questions than Taki seemingly wanted to know.
Kevin Roose: 63:46 Now, are you going to point Taki at any big sort of scientific discoveries and see whether it can make them. I mean, Demis Hassabis at Google DeepMind has this sort of theory that AGI should be able to discover Einstein’s theory of relativity if you just give it all of the pre-existing scientific literature at the time. Are you hoping to use this model or a descendant of this model for anything like that?
Adam Rodman: 64:13 Absolutely, yes. So one of the like, I think Nick especially is interested in this question of given a state of knowledge, sort of how much would it take to like how far ahead can you just from pure reasoning advance your state of conceptual understanding and the classic examples are like, um, you know, some of Einstein’s discoveries, which really didn’t require experiments, they just required putting the pieces together. Um, and there’s actually another project called Machine is Mirabilis who that took a training cutoff of 1900 and tried to get it to see if they could rediscover special relativity. I mean, the thing is that the models that those people did those experiments on were like I think three billion parameters like um just not smart enough to do very much so he showed that you could hold its hand at certain point and it would kind of get gesture in the right direction but to do the kind of like systematic reasoning and math that like Einstein had to do we probably need another let’s say like 10x in parameters at least.
Kevin Roose: 65:08 David, what are you building next? Are you going to build a bigger version of Taki and keep trying to get it to perform better?
Adam Rodman: 65:12 Yeah, so there’s a few things that we want to do. So like obviously making the models bigger and um so you know right now the model is still smaller than GPT-3 was although bigger than GPT-2. So you know Alec kind of says that there’s a bit of a phase change around like the size like around like maybe 100 billion or 150 billion parameters where the model starts to be smart enough to actually have a back and forth conversation with. Um obviously scaling up the dataset uh and OCR efforts and like right now everything’s just mostly English just because we can evaluate you know we can quality check the English text because we’re all native English speakers but we want to obviously broaden its repertoire. Um next working on the filtering is obviously a big one. Um and then obviously others I mentioned all this like how do we even evaluate the forecasting ability that’s like another big question. If you put the model in a robot, would that be a walkie talkie?
Casey Newton: 66:04 You can ignore him.
Adam Rodman: 66:08 Anyways, I gotta go.
Kevin Roose: 66:11 Well, Adam, fascinating experiment and people can go try Talkie for themselves. It’s at talkie-lm.com. What’s a good goodbye to a podcast guest? Don’t worry about what a podcast is. Okay, well, as Talkie says, a pleasant journey to you, sir.
Adam Rodman: 66:40 Thank you very much, sir.
Kevin Roose: 66:42 Thank you, Adam.
