YouTubeFeed

Give People Back Their GPT-4o | 3 Wow and 1 Promise

Summary

This episode of “3 Wow and 1 Promise” covers a week where all major AI labs launched their biggest releases simultaneously. Rather than getting lost in the hype, the host focuses on understanding the underlying truths and making meaningful connections. The episode features three “wow” moments and one “promise” about AI developments.

The first wow explores Kaggle’s Game Arena where LLMs from Google, Anthropic, OpenAI, and xAI compete in chess - not for superhuman performance, but to test generalist AI reasoning on tasks they weren’t specifically trained for. The hilarious results show models making grandmaster-level moves followed by amateur blunders. The second wow examines the messy GPT-5 launch and user backlash - despite extensive influencer preparation, millions of everyday users were furious about losing their “companion” GPT-4o personalities due to the new routing system. The third wow demonstrates ElevenLabs Music, showing real-time AI music generation from text prompts.

The “promise” section critiques Genie 3 from Google DeepMind - an impressive demo-only release that can’t be evaluated because it’s not available to try. The host emphasizes that demos always look good by definition, and models that can’t be tested don’t really matter for real-world evaluation.

Highlights

”Models play chess like amateur humans”

Clip

Clip command
yt-dlp --download-sections "*2:22-3:15" "https://www.youtube.com/watch?v=3Xa7B6sHnSo" --force-keyframes-at-cuts --merge-output-format mp4 -o "3Xa7B6sHnSo-2m22s.mp4"

“I watched the match between Grok and Gemini Pro and it was pure theater. They act like us amateur chess players. Five moves like a grandmaster sometimes and then followed by 10 moves as a hangover amateur. It was hard to explain why this model would make this move. It was really hilarious.” — Ksenia Se, 2:22

”People were furious about losing their companion”

Clip

Clip command
yt-dlp --download-sections "*4:44-5:15" "https://www.youtube.com/watch?v=3Xa7B6sHnSo" --force-keyframes-at-cuts --merge-output-format mp4 -o "3Xa7B6sHnSo-4m44s.mp4"

“We saw from Reddit discussions after the launch of GPT5. People were furious about losing their companion. Yeah, exactly that. When GPT5 launched as a model, we all realized that we lost something that was very special.” — Ksenia Se, 4:44

”700 million users a week - complete ignorance”

Clip

Clip command
yt-dlp --download-sections "*5:47-6:35" "https://www.youtube.com/watch?v=3Xa7B6sHnSo" --force-keyframes-at-cuts --merge-output-format mp4 -o "3Xa7B6sHnSo-5m47s.mp4"

“This uneven results, this complete ignorance about the 700 million users a week and potentially more is a funny wow to notice. My wow is about how they encapsulated and bubbled in their own world thinking how to prepare influencers and developers to say good things about GPT5 - not thinking about all these millions and millions of users.” — Ksenia Se, 5:47

”It’s not about model quality anymore”

Clip

Clip command
yt-dlp --download-sections "*7:00-7:55" "https://www.youtube.com/watch?v=3Xa7B6sHnSo" --force-keyframes-at-cuts --merge-output-format mp4 -o "3Xa7B6sHnSo-7m00s.mp4"

“It’s not that much about the quality of a model. Now it’s much more about the quality of user experience and what we can do with these models. We crossed this invisible threshold about if the model is good or not. Now it’s about relationship between the model and the end user.” — Ksenia Se, 7:00

”Demos always look good - that’s the definition”

Clip

Clip command
yt-dlp --download-sections "*11:48-12:35" "https://www.youtube.com/watch?v=3Xa7B6sHnSo" --force-keyframes-at-cuts --merge-output-format mp4 -o "3Xa7B6sHnSo-11m48s.mp4"

“That’s my problem with models that launched only with demo. Demo always looks good. That’s the definition of demo - to demonstrate the best of the model. So when you cannot try it, when you cannot actually put your hands on it, it doesn’t really matter.” — Ksenia Se, 11:48

Key Points

  • Kaggle Game Arena (0:35) - LLMs from Google, Anthropic, OpenAI compete in chess to test generalist reasoning
  • Not about superhuman play (1:26) - Unlike AlphaGo, this tests reasoning through games models weren’t built for
  • Grok’s lack of reasoning (1:48) - Grok doesn’t write out its thoughts unlike other models, just makes moves
  • Dynamic reasoning evaluation (2:51) - Game Arena creates objective, entertaining way to measure strategic reasoning
  • GPT-5 messy launch (3:16) - Despite 2-3 weeks of influencer preparation, launch disappointed regular users
  • “Just do stuff” phrase (3:50) - Both Ethan Mollick and Simon Willison used same phrase - possibly from press release
  • Different models for workflows (4:25) - Users had GPT-4o, o1, o3 for different purposes including roleplays
  • Router problems (5:10) - Smart router that chooses models for you caused uneven quality experiences
  • Models as thinking partners (8:20) - Models are your mind space - relationship matters more than raw capability
  • ElevenLabs Music (8:35) - Real-time music generation from text descriptions
  • Copyright protection (10:05) - ElevenLabs blocks prompts referencing copyrighted music like Back to the Future
  • GPT-OSS released (7:20) - New open-weight model from OpenAI on Hugging Face - important milestone
  • Genie 3 demo-only (11:48) - Google DeepMind’s world model only available to select academics and creators
  • Building understanding (11:30) - Show’s philosophy: build your own understanding, don’t depend on AI hype

Mentions

Companies

  • OpenAI (3:16) - Launched GPT-5 with extensive influencer campaign
  • Google DeepMind (11:48) - Released Genie 3 demo but limited access
  • Kaggle (0:55) - Created Game Arena for LLM chess competition
  • ElevenLabs (8:35) - Launched AI music generation product
  • Anthropic (1:09) - Claude competes in Kaggle Game Arena
  • xAI (1:48) - Grok performs without explaining reasoning

Products & Technologies

  • GPT-5 (3:16) - OpenAI’s new model with controversial launch
  • GPT-4o (4:25) - Previous model users were attached to
  • Genie 3 (11:48) - Google DeepMind’s world model (demo-only)
  • GPT-OSS (7:20) - OpenAI’s open-weight model on Hugging Face
  • Claude Opus 3 (7:30) - New model from Google (mentioned)
  • AlphaGo (1:30) - Referenced as superhuman game player

People

  • Demis Hassabis (1:18) - DeepMind CEO noted games as AI proving ground
  • Ethan Mollick (3:50) - Influencer who wrote about GPT-5
  • Simon Willison (3:50) - Influencer who wrote about GPT-5
  • Sam Altman (6:00) - Reading Reddit threads to understand user feedback

Surprising Quotes

“Grok is just like ‘bishop C4’ because you know, it’s going to be legend. Wait for it, dairy. I don’t know. It just has like a Barney Simpson vibe to it.” — 2:08

“Whatever it means, you naughty humans” (about role-plays with AI models) — 4:40

“Did I just create all this internet garbage? The answer is yes, I did. I’ve built like 70 apps across the different app builders.” — (Similar sentiment to the Webflow interview)

“How are we supposed to follow all the AI news if almost every AI lab launches their best models in one week? We don’t. We can relax. We can chill.” — 11:10

Transcript

0:00 If you feel you’re missing out with everything happening in the AI world, chill - because for some reason all the AI labs decided to launch their biggest things last week hoping that all of us will stay on top of it. We don’t have to. What we need to do is understand what’s happening in the long run, to make the connections that make us feel better, to see the underlying truth of that. So let’s go.

0:35 My first AI wow is not related to any of the models launched last week. It’s about chess gladiators - because how do you truly test an AI’s intelligence? The benchmarks are good, especially in-house benchmarks, but it’s easy enough to train the model to be good with the benchmarks and tests. So how about just launching all these models in the wild world of chess, especially when chess is not their specialization?

1:06 This was the idea behind launching Kaggle Game Arena where all the main LLMs from Google and Anthropic and OpenAI can pit against each other in complex games starting with chess. As DeepMind’s CEO Demis Hassabis noted, games have been always a proving ground for AI. But this is different. This isn’t about superhuman performance like we saw with AlphaGo playing Go with a human champion and actually winning that game. No, this is about generalist AI trying to reason their way through a game they weren’t specifically built for.

1:48 Now, it doesn’t take very long for Grok to just straight up forget about it. What’s interesting also is Grok doesn’t write its thoughts out. That’s what I’ve noticed two days into the tournament. The other models, they explain all their reasoning - as ridiculous and dumb and absurd as it is. Grok is just like “bishop C4” because you know, it’s going to be legend. Wait for it, dairy. I don’t know. It just has like a Barney Simpson vibe to it. “Bishop C4.” And Gemini just simply takes the knight. It just leaves the knight in the middle of the board. “I never said these games would be good. I said this matchup would be good.”

2:22 I watched the match between Grok and Gemini Pro and it was pure theater. They act like us amateur chess players. Five moves like a grandmaster sometimes and then followed by 10 moves as a hangover amateur. It was hard to explain why this model would make this move. It was really hilarious.

2:43 So the wow here is not about finding the next Alpha Zero. We already have superhuman chess engines. The point of the Game Arena is to create a dynamic, objective, and endlessly entertaining way to measure the strategic reasoning of generalist AI. Through this process, we see their flaws. We see how stupid their reasoning is. And with all the models opening their reasoning, it’s an important part for researchers to understand how these models work.

3:16 Okay, my second wow is about GPT-5 launched last week by OpenAI, but not about the model itself and not even about the messy launch - considering how much work they put into working with influencers, developers, media and preparing the whole campaign for launching GPT-5 at least two, three, probably more weeks before that.

3:40 When GPT-5 was launched, we saw many articles coming from the main influencers with detailed diaries how they were using GPT-5 for the last two weeks. We saw it from Ethan Mollick. We saw it from Simon Willison. The funny thing is that both of them say that GPT-5 “just does stuff” - as if it was a phrase from a press release or something. But maybe they both came up with the same exact phrase. I don’t know.

4:10 What surprised me is that so much work put into the usage of the model that influencers reported to be their daily partner turned out to be so messy from the point of actual users who got used to using all the suggested models that OpenAI had. They had GPT-4o, they had o1, they had o3 and different people would use different models for different workflows and even different role plays. Whatever it means, you naughty humans.

4:44 We saw from Reddit discussions after the launch of GPT-5. People were furious about losing their companion. Yeah, exactly that. When GPT-5 launched as a model, we all realized that we lost something that was very special. It was not my experience to be honest. I have a lot of workflows with GPT-4, but when I tried it with GPT-5, it was basically the same.

5:10 But I think the problem was and maybe still is with the router - this new thing that OpenAI introduced. When you collaborate with a model, this smart router chooses for you which model is needed for your task. And some people reported that suddenly GPT became extremely stupid. My experience and the experience in my household was very positive. We have very good results with GPT-5. But it depends on the day actually. When I tried to use it yesterday, it was much dumber than when I tried to use it today - it was good quality.

5:47 So this uneven results, this complete ignorance about the 700 million users a week and potentially more is a funny wow to notice. Now Sam Altman and all the team carefully reads every Reddit thread, listens to Twitter, understanding what did they miss. My wow is about how they encapsulated and bubbled in their own world thinking how to prepare influencers and developers to say good things about GPT-5 - not thinking about all these millions and millions and millions of users who actually use GPT-5 for like enormous amounts of tasks that developers and influencers can’t even come up with.

6:37 We still need to see how GPT-5 performs. There’s still not enough evaluations. It’s still all over the place. I believe the team will fix it pretty soon. I like GPT-5. ChatGPT was always my first model to go to. But with this new model, I haven’t noticed big difference in quality. I believe they will fix it.

7:00 The another interesting notion here is that it’s not that much about the quality of a model. Now it’s much more about the quality of user experience and what we can do with these models. A lot to learn for every company. And last week was a crazy week because we also saw GPT-OSS the new open-weight model from OpenAI launched on Hugging Face. That was a very important milestone for OpenAI to become open. We saw new Claude Opus 3 from Google.

7:38 Anyway, launching big models in bulk is not what users need. So the thing and revelation for all the companies should be to think about the millions of users they actually serve. So yeah, we crossed this invisible threshold about if the model is good or not. Now it’s about relationship between the model and the end user. Not even about relationship of the model and a person who studies it - researcher, influencer, developer.

8:08 AI is in the world. Millions of people use it. Soon it will be billions, maybe it is a billion people already. What are their relationships with these models? What are they using it for? How do they get used to these models? What are their workflows? And again, relationship with their thinking partners. Because a lot of time the models are your mind space - getting into this mind space. That’s the biggest mess of the last week launch.

8:35 Okay, our final wow this week comes from a company that has already mastered the human voice and now they’re turning attention to music. We’re talking about ElevenLabs Music and I want to show you how it works hands-on. Let’s go. “What song do you want to create? Describe your song.” The song about summer on a lake with kids, laughter, adults talking softly on the background, the heat is tender, the water is lovely, the ice cream is waiting in the freezer after a swim.

9:20 [Music plays] “Their feet on the dark, stories told, echo after dancing on the waves.” Okay, that’s pretty cool. I need a song for “Three Wows and One Promise.” Let’s see how it does with a copyrighted part of it. “A song with a main theme from Back to the Future about the show with the name Three Wows and One Promise about AI and the real world. Make it sound as Back to the Future main music theme.”

10:05 “Oh, my prompt appears to have violated our terms of service. Please try again with a different prompt.” Huh? “A song that reminds the main theme from Back to the Future.” Let’s try with “reminds.” Oh, no. Still violated. Okay, I bet if you tweak the prompt and come up with something in your own words that describes it more precisely, you will come up with a good melody that will remind you of Back to the Future.

10:35 Anyway, that’s a next level of creating music. I’m not sure how musicians will feel about it. I as a user, of course for me it’s easier to use this tool to create music if I need it than to approach a musician and have all these interactions, time back and forth. So that’s definitely saving my time. I hope for the musicians it will save their time if they have a blank page moment, if they want to experiment with different styles. So I hope that will be a great instrument for them.

11:10 How are we supposed to follow all the AI news if almost every AI lab launches their best models in one week? We don’t. We can relax. We can chill and we can see what was the coolest part of it and what was the underlying truth that will help us think about AI and the world we’re building with it. This is the thing about Three Wow and One Promise. We are building our own understanding not depending on the hype of the AI world.

11:48 I mentioned Genie 3 from Google and that’s my problem with models that launched only with demo. Demo always looks good. That’s the definition of demo - to demonstrate the best of the model. So when you cannot try it, when you cannot actually put your hands on it, it doesn’t really matter. On this show, the promise is about something incredible on the horizon but something that isn’t quite here yet. And Genie 3 from Google DeepMind is exactly that.

12:18 Genie 3 is currently offered only as a limited research preview to select academics and creators. And it’s a bummer. How am I supposed to judge and how am I supposed to say if it’s a wow or not wow if I cannot try it? But we will try it at some point. Until then, I trust the promise.

12:40 Thank you for joining Three Wows and One Promise. We’ll be back next time to connect the dots, to save you from FOMO, to give you the ground and a few ideas how to think about AI and the real world. See you there. Until then, hit subscribe, follow us, and leave me a comment.