YouTubeFeed

GPT-5.6 Sol: Better AND cheaper than Fable

Summary

In this solo episode, Claire Vo puts OpenAI’s newly re-released GPT-5.6 lineup — Soul (the top-of-the-line frontier model), Tera (a balanced everyday model), and Luna (a cheap, high-volume model) — through her homemade “How I AI” vibe benchmark. She grades five models (Fable 5, Sonnet 5, and the three GPT-5.6 variants) across PRD writing, prototyping, wireframing, code debugging, and agentic voice, then scores the results on a “Claire Weighted Index” that blends 70% her own taste with 30% an LLM judge (GPT-5.5). The verdict: GPT-5.6 Soul wins by a meaningful margin, and at $5/$30 per million input/output tokens it undercuts Fable’s $10/$50.

The through-line is a distinction between “theoretically hyper-intelligent” models like Fable and “practically effective” ones like Soul. Vo respects Fable’s raw intelligence but finds it pedantic, inscrutable, and painful to collaborate with — “an engineer that has never met a human before.” Soul, by contrast, writes like a person, produces opinionated and functional prototypes that escape generic “AI slop,” and is willing to loosen its own constraints to ship real user value. She does credit Sonnet 5 as still the best “agentic voice” (the least cringe to talk to) and Tera as her favorite for clean, direct PRD writing.

Beyond the benchmark, Vo demos the use cases she can’t stop using: a one-shot gamified homework-tracking app for her kids built in Codex, rapid social-clip video editing of a Cursor talk, and — her favorite — Chrome browser automation via Codex that burned through ~500 LinkedIn messages while she did nothing. She closes with model recommendations and notes all the benchmark work will be published on the ChatPRD blog.

Highlights

”Your girl loves 5-6-o. She just does.”

70-30 Claire Weighted Index

“And so if you look at that 70-30 split, your girl loves 5-6-o. She just does. It had the highest taste score by a significant amount, so I just thought it output the best work.” — Claire Vo, 7:34

Clip command
yt-dlp --download-sections "*7:34-8:11" "https://www.youtube.com/watch?v=gAWbvEwUoiI" --force-keyframes-at-cuts --merge-output-format mp4 -o "girl-loves-56.mp4"

”It talks to me like an engineer that has never met a human before”

Fable's first day on earth

“I hate talking to Fable 5 because it talks to me like an engineer that has never met a human before. It’s like its first day on earth.” — Claire Vo, 8:26

Clip command
yt-dlp --download-sections "*8:26-9:00" "https://www.youtube.com/watch?v=gAWbvEwUoiI" --force-keyframes-at-cuts --merge-output-format mp4 -o "fable-engineer.mp4"

”It loves a forest green” — the GPT-5.6 tell

Woodland elegance

“It loves a forest green. In fact, I think this forest green is like in its system prompt called like woodland some, woodland elegance or something like that… you will see a lot of green and I think this is one of the GPT-4o tells that you will start to notice and get really frustrated with.” — Claire Vo, 16:38

Clip command
yt-dlp --download-sections "*16:38-17:48" "https://www.youtube.com/watch?v=gAWbvEwUoiI" --force-keyframes-at-cuts --merge-output-format mp4 -o "forest-green-tell.mp4"

Fable is theoretically hyper-intelligent, Soul is practically effective

Practically effective

“If you take away like one highlight difference between Fable and Soul is like Fable is theoretically hyper-intelligent and Soul is practically effective.” — Claire Vo, 21:15

Clip command
yt-dlp --download-sections "*21:15-21:54" "https://www.youtube.com/watch?v=gAWbvEwUoiI" --force-keyframes-at-cuts --merge-output-format mp4 -o "practically-effective.mp4"

A gamified homework app for coin-operated kids, built in one shot

Gamified homework tracker

“If he does his homework, I need to, like, give him a skittle or let him trade skittles for Nerf guns, and he will, like, learn calculus by the time he’s in fifth grade.” — Claire Vo, 24:08

Clip command
yt-dlp --download-sections "*24:00-25:00" "https://www.youtube.com/watch?v=gAWbvEwUoiI" --force-keyframes-at-cuts --merge-output-format mp4 -o "gamified-homework.mp4"

”A beast when it comes to browser use” — 500 LinkedIn replies

Browser use beast

“It went through and burned through probably 500 messages… Browser use and 5.6, and when I got rolled back to 5.5, my life was worse. So please, please, please, learn to use at Chrome, at browser, and at computer and just let Codex rip and let GPT 5.6 rip.” — Claire Vo, 34:21

Clip command
yt-dlp --download-sections "*34:21-35:08" "https://www.youtube.com/watch?v=gAWbvEwUoiI" --force-keyframes-at-cuts --merge-output-format mp4 -o "browser-use-beast.mp4"

Key Points

  • The three GPT-5.6 models (1:13) - Soul is the brainiest frontier model, Tera is a balanced everyday model, Luna is the cheap high-volume (Mini/Nano-style) model
  • This is a love letter to Soul (2:00) - The episode focuses on Soul versus Fable for daily product work, though Tera and Luna get tested too
  • Pricing beats Fable (2:16) - Soul is $5/million input and $30/million output tokens vs. Fable’s ~$10/$50; more affordable even at API pricing
  • Subscription speculation (3:00) - Vo suspects OpenAI will keep Soul in the subscription and that this could pressure Anthropic to put Fable back into the Claude subscription
  • State-of-the-art on Terminal Bench 2.1 (3:24) - Highest performing in ultra mode, plus new cybersecurity benchmarks and safety framework discussion in the blog
  • The How I AI benchmark (4:30) - Tests PRD generation, wireframing, full design prototypes, code debugging, and “talking to me like a human”
  • The eval harness and judge (5:30) - Runs every eval against each model, uses GPT-5.5 as the “hardest” LLM judge, plus a manual “Claire-Vo taste test”
  • The em-dash / slop complaint (6:41) - Vo repeatedly begs for someone to get rid of em-dashes and “slop talk” in agentic voice
  • 70-30 Claire Weighted Index (7:25) - Final score is 70% Claire’s taste, 30% the machine judge; Soul wins on taste by a significant amount
  • Blind taste test (7:59) - She graded the models blind before learning which was which, then mapped them back
  • Per-task winners (9:17) - Soul for prototypes, Tera for clean/streamlined PRDs, Sonnet 5 for bug-hunting and agentic voice
  • Sonnet’s “you are a human” gold star (10:12) - Sonnet 5 got high praise for agentic voice (“aside from the em-dash, you are a human”)
  • Claude’s editorial design aesthetic (10:46) - Beige backgrounds, burnt orange, italic serif fonts — recognizably “Claude,” which Vo ranked low
  • Escaping “Blurple slop” (11:24) - Soul excels at complex, dense, technical, uniquely designed pages; she has Soul redesign an Opus-made page
  • What Claire rewards (12:22) - Uniqueness, creativity, and functionality in design; succinct, direct, non-AI writing; she gave 50 written reactions with 14 flagged as garbage
  • Dense dashboard side-by-side (13:42) - On a doc-scheduler dashboard, Soul delivered clean neutral layout, semantic color, and everything actually clickable/functional
  • Functionality made the difference (14:14) - Across prototypes, Soul’s builds were consistently more functional, which heavily influenced scoring
  • Soul writes like a normal person (19:21) - Fable is “inscrutable,” “for-agents-by-agents” writing; hard to collaborate with despite intelligence
  • Greenfield ChatPRD rebuild (22:03) - Soul rebuilt her ChatPRD concept with research, executive-recommendation tables, and readable prose
  • One-shot gamified homework app (24:00) - Built XP system, focus-mode timers, quests, rewards, companion avatars, and a parent HQ in essentially one shot
  • Breaking Fable out of its own hardened loops (29:40) - Fable hardened a tool-calling loop so only GPT-5.5 would run; switching to Codex fixed it by “getting out of its own mind”
  • Video editing use case (31:43) - Drag in a long talk recording, ask for five tight horizontal social clips, then finish in CapCut
  • Browser use is the best use case (33:35) - Codex + GPT-5.6 + Chrome for LinkedIn triage, testing web apps, and filling out annoying forms

Mentions

Companies

  • OpenAI (1:13) - Maker of the GPT-5.6 Soul/Tera/Luna lineup being reviewed
  • Anthropic (3:00) - Maker of Fable and Sonnet; discussed re: subscription access and pricing pressure
  • Cursor (31:54) - Hosted the event where Vo gave a talk on the future of PM; source of the recording she clipped
  • LinkedIn (33:35) - Where she ran the Chrome browser-automation message-triage demo
  • ChatPRD (20:44) - Vo’s product; Soul helped build a prototyping tool inside it and later greenfield-rebuilt it

Products & Technologies

  • GPT-5.6 Soul (1:13) - The frontier model and Vo’s overall favorite; best at prototypes, functional design, and writing
  • GPT-5.6 Tera (9:27) - Balanced model; her favorite for clean, straightforward PRD writing
  • GPT-5.6 Luna (1:13) - Cheap, high-volume model for everyday work
  • Claude Fable 5 (2:16) - Hyper-intelligent but pedantic and hard to talk to; still produces good code
  • Sonnet 5 / Sonnet 3.5 (9:54) - Best agentic voice; LLM judge rated it top on bug-hunting; her go-to for OpenClaude
  • GPT-5.5 (5:30) - Used as the “hardest” LLM judge in the eval harness; only model that would run a hardened tool loop
  • Opus (11:37) - Used to make a “sloppageous” Blurple/gradient page that Soul then redesigned
  • Codex (19:05) - OpenAI coding agent used for the homework app, the loop fix, and browser automation
  • Terminal Bench 2.1 (3:24) - Benchmark where Soul is state-of-the-art in ultra mode
  • OpenClaude / Open Claw (0:29) - Vo’s agent setup; she prefers Sonnet for its voice there
  • Chrome (@Chrome / @browser / @computer) (33:35) - Browser tooling driven through Codex for automation on logged-in pages
  • CapCut (33:13) - Where she finished the social hype clips with music
  • Math Academy (25:00) - Referenced inside the kids’ homework-tracker quests
  • Cursor (24:28) - Used to turn a Claude-generated PRD into the homework app

People

  • Claire Vo (0:00) - Host of How I AI, founder of ChatPRD; runs and narrates the entire benchmark

Surprising Quotes

“I have been very, very, very sad the last week because for the last week I have not had access to my true favorite top-of-the-line model GPT-5.6, but guess what babes, it is back.” — Claire Vo, 0:00

“This is my show, this is my podcast… and what I said about the performance of the models, and then I get to strike the difference. And you know what? I’ve decided I like my own taste better.” — Claire Vo, 7:14

“Fable makes up, it seems like Fable’s unfamiliar with the English language and communication with humans. Fable is very much like a for-agents-by-agents communication mechanism, I can barely make out what it’s talking about.” — Claire Vo, 19:42

“I really struggle working with theoretically intelligent colleagues who can’t get anything done, like can’t actually see the forest for the trees, get too much in their head.” — Claire Vo, 21:30

“I mean Soul said, ‘This is a deploy, not a referendum, like please don’t do this, not that to me, do not do m-dashes.’” — Claire Vo, 18:51

Transcript

Claire Vo: 0:00 I have been very, very, very sad the last week because for the last week I have not had access to my true favorite top-of-the-line model GPT-5.6, but guess what babes, it is back and I am here to walk you through GPT-5.6 Soul, GPT-5.6 Luna, GPT-5.6 Tera. I’m going to tell you what are these models, how have I been using them, why are they my heart’s favorite and is Fable better than all of them or not?

Claire Vo: 0:29 I have been testing this model for a couple weeks, there was a few days there where we didn’t have access and I found myself desperate to get this workhorse model back. Now, we’re not just relying on my own opinion, we are going to run the very famous, very new How I AI Vibe Review Benchmark against common tasks from PRD writing to prototyping to whether or not it’s cute in my Open Claw Agent, and I’m going to tell you very scientifically if this is the model that you should be working with all the time now. Let’s get to it.

Claire Vo: 1:13 Okay, you all can read these blog posts, so I’m not going to go into too much depth about the models and the benchmarks, I’ll just give you the hits. First, OpenAI is releasing three new versions of their GPT-5.6 model. Soul, which is the next-generation frontier model, the brainiest of the brainiest. Tera, which is a balanced model for efficient everyday work, and Luna, which is sort of akin to their Mini or Nano models, which is cheap and affordable for high-volume work.

Claire Vo: 2:00 So you’re going to have these three versions of the models. I don’t know if these beautiful images are exactly how we should think about the relative capabilities, this big sun, this medium earth and this tiny moon, but I will say my love letter that is this podcast today is written directly to GPT-5.6 Soul. This big model is the one I love.

Claire Vo: 2:16 Now, I have tested Tera and Luna so I will give you my input there, but really this is going to be all about Soul versus Fable and which one I would use for the type of work that I’m doing every day. Okay, quick note on pricing. Soul is a lot more affordable than Fable, so it’s $5 per million input tokens, $30 per million output tokens.

Claire Vo: 2:44 I believe Fable at the time I’m recording this is 10 on a million input tokens and 50 on a million output tokens. Now again, you’re going to get a little bit of subscription usage built into your OpenAI subscription so you are going to get a decent amount that you can test with and use. You know, there’s been some challenges with the Fable rollout, they’ve limited when it’s been included in the subscription, and so it was supposed to be available till early this week, I think they extended that a little bit at Anthropic so subscription Claude users

Claire Vo: 3:00 Users could save under their subscription. So we have to see how much soul usage we get and if like Anthropic they’re going to take soul out of the subscription. I suspect not. I suspect this is a model they want people to use. I also suspect this might put pressure on Anthropic to put fable back into the Claude subscription. But for now, it’s more affordable even at API pricing.

Claire Vo: 3:24 Now, I’m not going to read through all the benchmarks for you. You can go to this OpenAI blog and read them for yourselves. All I will say is it is the brand new state-of-the-art model from OpenAI. It is the highest performing when using the ultra mode on terminal bench 2.1. And then they’ve also eval’d it against a couple cyber security benchmarks. So I do think as we get these smarter models you’re going to see a lot more evals and benchmarks around exploits and security.

Claire Vo: 3:55 And then very similar to what we’re seeing with Fable, there’s a lot of conversation in this blog post about the safeguards and security frameworks around the release of this model. I do believe like Fable it’s going to fail over in some tasks that are maybe a little bit riskier. But I have not run into that myself.

Claire Vo: 4:15 Now let’s get back to how I eval these models. If you missed my episode on Fable, I got kind of bored of the vibe vibe check and I built an extremely scientific “How I AI” benchmark.

Claire Vo: 4:30 Now this How I AI benchmark tests basically a couple things. It tests the ability to generate good PRDs. It tests the ability for it to wireframe against a couple different app ideas, develop fully designed, robust design prototypes, debug code, and then talk to me like a human which is the thing that I care about the most.

Claire Vo: 4:52 And I’m just going to remind you how I did these benchmarks and then scan you through a couple of the outputs. And since I know what the models are now after I’ve done the grading, I can show you which ones map to Fable and GPT-5.6.

Claire Vo: 5:11 Okay, so this is my vibe review. What I tested was Fable 5, Sonnet 5, and then the three versions of GPT 5.6. I did it against my common use cases of PRDs, prototyping, coding, and chitchatting with an agent.

Claire Vo: 5:30 And then what I have the eval harness do is it runs all the evals against each of these models and it does a LLM-based judge. The LLM that I’ve decided is the hardest judge is GPT 5.5. So that’s the one that judges. But it also gives me this page where I can actually go through and give what’s called the Clair-Vo taste test, which is I read all the assets, I look at all the designs, I score it and give it notes.

Claire Vo: 5:51 And so you can see here I went through PRDs, we went through sort of some complex prototypes here in terms of a doc scheduler, we did a consumer app… Up, so lots of beautiful different habit tracker apps, different versions you can see here, a pretty complex dev tool, wireframe versions of those same prototypes which I graded.

Claire Vo: 6:13 I also give notes. This one great note says, ‘My fave but not great.’ And then I let the code grader just evaluate the agentic multi-step debug, because I wanted it to be really about accuracy there, and I didn’t feel like I could eyeball that and give a strong opinion.

Claire Vo: 6:29 And then the last thing that it generates is an agentic voice. So basically, how it would respond to me answering a couple questions. Very important on agentic voice, these models—somebody, please hire somebody to get rid of the em dashes and slop talk. I cannot stand it.

Claire Vo: 6:48 Now, one of the things that I will say as an observation for 5-6 is it’s a great writer, and I will show you some examples of that, but truly, a lot of my evals here were em dash slop. I hate you.

Claire Vo: 7:00 Okay, so let’s go to what the Claire Weighted Index says. Now, this is my show, this is my podcast, and so I sort of strike the balance between what the LLM judge said about the performance of the models and what I said about the performance of the models, and then I get to strike the difference. And you know what? I’ve decided I like my own taste better.

Claire Vo: 7:25 So I’ve decided it’s going to be a 70 Claire Vo, 30 the machines split on evaluating these models.

Claire Vo: 7:34 And so if you look at that 70-30 split, your girl loves 5-6-o. She just does. It had the highest taste score by a significant amount, so I just thought it output the best work.

Claire Vo: 7:50 Again, I went through dozens of evals, looked at them, clicked through them, gave my own opinion, and put notes, and I just have to say, I really like GPT-5-6-o. I know I spoiled it at the beginning, but I did blind taste test these, and so I do really feel like it did a good job, and I will give you a couple examples of that.

Claire Vo: 8:11 Now, I don’t hate Fable 5, so I’m not saying that Fable 5 is out of the game. I will say, I did not have to talk to Fable 5 when running this benchmark.

Claire Vo: 8:26 I hate talking to Fable 5 because it talks to me like an engineer that has never met a human before. It’s like its first day on earth. Um, but when I don’t have to talk to Fable 5, it outputs pretty good work, and I would say had some good outcomes there.

Claire Vo: 8:46 Um, and then Terra Luna did fine work. Sonnet 5 at the bottom, really haven’t figured out how to get this one working, although there’s a very specific use case that we think Sonnet 5 is good at, or actually two, two use cases.

Claire Vo: 9:00 Now this is heavily weighted on its front-end prototyping design and app building capabilities since that is the trunk of the how AI eval it is heavily weighted there, but I do want to call out that per task I do have a couple favorites.

Claire Vo: 9:17 So for that prototype task and we’ll go to some examples in a minute, I just love 56 Soul. I just really do. I think it was functional, the designs were the most interesting, I thought it was really good. For PRD, I liked Terra, maybe it’s down to earth, maybe I like a basic straightforward PRD.

Claire Vo: 9:33 As I said, it was my favorite streamlined and to the point and so if you want clean crisp direct business writing, maybe GPT 56 Terra is the way to go. You know, the bug hunting eval which I don’t really feel like I’ve nailed exactly, so I’m not super confident in this one,

Claire Vo: 9:54 but the LLM as a judge thought that Sonnet 5 did the most complete and accurate job. I will say, I only like talking to Sonnet models through my OpenClaude. Really, I only like talking to them, I still really struggle with getting my OpenClaude to work well with the GPT models.

Claire Vo: 10:12 I still did not like Fable in the agentic voice eval which you should not be surprised at, but Sonnet 5 got a very good gold star for me because I said aside from the em-dash, you are a human. That is very, very high praise.

Claire Vo: 10:28 And then I’m going to show some of these designs in a second, but you can see across the board on a full fidelity prototype, I just really preferred 56 Soul three out of five times, 56 four out of five times and Sonnet did an best job at the editorial design.

Claire Vo: 10:46 I will say Claude’s design aesthetic tends to this sort of like editorial design, if you know, if you’ve seen it you know it, it’s like that beige background, that orange, burnt orange color, the italic serif fonts, it’s just very, very Claude.

Claire Vo: 11:01 But I hated that design overall the most, so you can see here I ranked it still lower than almost anything else on this leaderboard, it just happened to be the best of the worst, I would say.

Claire Vo: 11:14 Now, where GPT 56 Soul did a really good job and I’ll show some of these examples is like complex, dense, technical, unique designed things.

Claire Vo: 11:24 And so I will say I have been happy to extract myself out of Claude slop, out of Blurple slop into more interestingly designed websites. And I’ll even show an example kind of like a meta example which is this is the Opus designed version of this page.

Claire Vo: 11:44 Like very sloppageous and we got the Blurple, we got a gradient, I don’t think the typography is particularly sophisticated, and I asked Soul to redesign it and I just think this is a lot cleaner, a lot nicer and easier to look at. Okay, let’s look at the results.

Claire Vo: 12:00 Let’s talk about how Claire qualitatively evaluates models. Some of these quotes will just give you a sense of what I value. And again, I gave 50 written reactions. There were some like unmistakable hits where I loved what the models came up with, 14 places where I was like, this is garbage. So let’s see like kind of what I talk about when I review things.

Claire Vo: 12:22 So I was definitely calling out uniqueness, creativity, and functionality in the design. And so in designs I like non-slop, unique designs that are functional. And so I’m definitely going to reward this doesn’t look like the generic prototype and you’ve pulled the thread of functionality through the prototype.

Claire Vo: 12:43 For writing, I just like succinct and to the point. I cannot stand AI writing. It drives me nuts. I can see it a mile away. So I really like just direct, very frank, very crisp writing. I think 5-6 is good at that.

Claire Vo: 12:58 And then you can see the things that I hate. I hate slop. I hate slop. I hate slop. We all hate slop. It’s the worst. It’s the worst part of AI. If I hate one thing about AI, it is that I have to experience slop.

Claire Vo: 13:14 So you can see, I’ll like Claude design slop across this editorial page, typography, emojis and bad placeholders, like I really held a high bar in terms of design quality. Okay, let’s look at a couple of these and why I really liked Soul compared to other models, although where Fable did a perfectly serviceable job.

Claire Vo: 13:32 Okay, so this dense operation dashboard is basically like an eval for a doc scheduler app and its full design. And what you can see here is both were pretty useful. Soul on the left and Fable on the right. I just think Soul was the most unique. All of the other ones really just looked like this dark mode, monospaced kind of layout.

Claire Vo: 13:57 As you can see here, Soul actually has like a really clean, kind of like neutral color layout with great visual hierarchy, semantic color, and this thing was functional. So like everything I expected to be able to click and work and assign and do, all of it actually worked.

Claire Vo: 14:14 And this was just my experience across a bunch of the different prototypes is the Soul ones were just a lot more functional. And that made a big difference on how I’m evaluating things.

Claire Vo: 14:26 Now, let’s look at the Fable design again. It’s pretty good. It’s actually a lot harder to read though. And the design I would say is not as unique. And even some layout issues like this white space here at the bottom. Now, it did do a lot of functionality, but I would say like the colors weren’t semantically assigned, the typography needed some work. And I just really preferred this unique design of Soul, even though it wasn’t crazy, it was just opinionated, which I think is nice.

Claire Vo: 15:00 Now, here is another design, it was this creative pack website again, both of these got fives from me, I just really preferred that Sol went ahead and had like a personality. Look at these placeholder images versus what Fable came up with, which I will say is beautiful and clean and worked really well and like I have no complaints about it. It’s a good one especially for sort of a wireframe style prototype, it’s great. I would just say it’s not this, this is pretty interesting, it’s got a better point of view, and it’s got like nice little design affordances that I just didn’t see in these other designs. And so I just really preferred or at least I rewarded the fact that Sol, you know, used its brains to be a little bit more unique and give me some inspiration.

Claire Vo: 15:59 Then on this dev tools page, this is again where Sol went really well and it’s sort of the same as the doc scheduler. It just does the job of this is a um incident triage site, it just does the job a lot better than I would say the Fable five did. Fable five is fine, it’s just not that unique and again the thoughts around the design are not exactly what I would want. And so again this like functionality point of view design I really preferred Sol.

Claire Vo: 16:38 Another last side-by-side comparison and again I think this is a good one to think about. If you see here we did these habit tracker apps and just looking at the comparison side-by-side design, like this is good ol’ classic Claude stuff, you’ve seen this design a million times especially if you’ve used Claude Co-work. And if you look at this, it’s just again a little bit more opinionated. There are some slop pieces to this design, um some things that I did not love. Um the one thing I will say I noticed about um Sol, which you will notice, which I have told the delightful and lovely OpenAI team, and maybe it’s because they love me, it loves a forest green. It loves a forest green. In fact, I think this forest green is like in its system prompt called like woodland some, woodland elegance or something like that. I mean look, I love a forest green, look at my office, it is forest green, but you will see a lot of green and I think this is one of the uh GPT-4o tells that you will start to notice and get really frustrated with.

Claire Vo: 17:48 Now on wireframes again let’s just look at these side by side. Sol: very functional, very easy to read like as a person trying to convey a complex application I think this… a really quite excellent job and just a better job of this. It’s just a little harder to read, I’m not quite sure what I’m supposed to do here, it’s not as functional. There are some interesting things here, but you know what Fable came up with was not my favorite.

Claire Vo: 18:19 Now, final thing is its voice, I just want to call out, I do love Sonnet for agentic voice, so I cannot knock Sonnet for not sounding ridiculous.

Claire Vo: 18:33 So I asked it in sort of a EA personal assistant, open-Claude style a couple questions: can you move my meeting? Deploy is Red again. Why did I start this company? Let’s just yolo straight to prod and how Sonnet replied and how Soul replied.

Claire Vo: 18:51 Uh, you missed the line break so it read a little bit better in the eval, but if you read them, like Sonnet’s still super cringe, but Soul was worst. I mean Soul said, ‘This is a deploy, not a referendum, like please don’t do this, not that to me, do not do m-dashes.’

Claire Vo: 19:05 So I could not get rid of m-dashes, but I thought Sonnet 3.5 had the best voice. I tend to use Sonnet for whatever for my open Claude, so I’m not surprised about that. Okay, so that is the Clair-Vo Eval, but I want to go into a couple other things I really love about this model. So let’s switch over to Codex.

Claire Vo: 19:21 Okay, I’m going to zip through a couple examples of things I think Soul does a lot better than other models, and in particular, a lot better than Fable. Number one: it writes like a normal person. I cannot cope. I love Fable, you’re brainy, as I showed the eval shows you do a pretty good job. I cannot talk to Fable anymore.

Claire Vo: 19:42 Fable makes up, it seems like Fable’s unfamiliar with the English language and communication with humans. Fable is very much like a for-agents-by-agents communication mechanism, I can barely make out what it’s talking about. It is incredibly inscrutable writing, and that makes it very hard to collaborate with your model.

Claire Vo: 20:18 And so what I would say is my experience using Fable has been it is like incredibly technical, incredibly pedantic, and while it is super intelligent, hard-working, will like definitely fan out and solve very complex problems, its ability to collaborate is low, and it left me with a lot of frustration as an end-user using Fable.

Claire Vo: 20:44 Now, Fable did knock off some like pretty complex work and I’m very happy to go through what that is. It helped me build a full prototype tool inside ChatPRD, so like a V0-lovable-etc. version prototype tool. It’s helping me build… like synthesis product brain product that I’m working on, but I’ve found it incredibly hard to break it out of its own sort of frameworks, its own limitations, its own structured way of approaching problems.

Claire Vo: 21:15 And what I really feel like the difference, if you take away like one highlight, um difference between Fable and Soul is like Fable is theoretically hyper-intelligent and Soul is practically effective.

Claire Vo: 21:30 And so like, I’ve been an executive long time, I’ve been a manager long time, like I really struggle working with theoretically intelligent colleagues who can’t get anything done, like can’t actually see the forest for the trees, get too much in their head. And so like when I want to ship stuff to customers, I need practical, get the job done, understand the end user goal, understand the end user and like willing to loosen constraints appropriately to get things done.

Claire Vo: 21:54 And that has just so much more been my experience with Soul versus Fable.

Claire Vo: 22:03 The writing is straightforward, the communication is clear and it’s less pedantic. I’ll just give you a quick example of this which is I had Soul look at my Chirp PRD repo and like Greenfield totally rebuild it.

Claire Vo: 22:15 Just my idea was like completely rebuild your idea of what Chirp PRD should be in 2026. And it went and did a bunch of research and it came came back to this and again, love me an executive recommendation starter, uses you know tables, um what exists today is very straightforward and easy easy to understand.

Claire Vo: 22:29 This is a very long document. I did read a lot of it and it’s just easier to parse than anything I’ve seen come out of Fable. So so writing communication, definitely plus in Soul’s corner.

Claire Vo: 22:41 The second thing is like full zero-to-one prototypes. As we’ve seen in the Eval benchmark I just really like, so again for this like rewrite Chirp PRD from the ground up, it came up with this idea of like taking a problem space or a decision, validating it with external insights

Claire Vo: 22:54 and then pulling it all the way through coding handoff and built this pretty complex prototype. Now do I love everything about this idea? No. Are we doing some of the things about this idea including insights generation? For sure.

Claire Vo: 23:08 But this was actually very nice from a prototyping perspective and I thought it did a good job of giving me a robust thing to experiment with and gave me some good ideas about what I could do with the product next. So I was pretty happy with the like zero-to-one prototype. Now a little bit more fun example is I asked Soul to

Claire Vo: 24:00 To make a fully gamified homework tracking system for my kids. Look, my kids are coin-operated. I have a middle child who’s basically going to be an enterprise sales rep.

Claire Vo: 24:08 If he does his homework, I need to, like, give him a skittle or let him trade skittles for Nerf guns, and he will, like, learn calculus by the time he’s in fifth grade.

Claire Vo: 24:16 But I’m a vibe code lady, and so I want to build a app. Just sneak peek into our household, my husband sent me a XP system proposal via Claude this morning, so I’m taking a Claude-generated PRD, dropping it into Cursor and GPT-4o, and generating something. Now,

Claire Vo: 24:36 what I came up with was pretty ambitious. Now, do I love the design? Is it a little, like, does it have some AI tells? It’s like gradients, you know, fonts, all this kind of stuff. But it’s like cute in a way. Look at this, you know, it’s using this emoji really well with the texture, it’s doing some animated things here.

Claire Vo: 24:53 And basically, it’s giving my oldest child and my youngest child two different summer quests they can do.

Claire Vo: 25:00 They can enter focus mode. I think this is really good, again, from a design perspective. They can enter focus mode. What does this listen to? ‘Math Academy. Finish one focused Math Academy mission. Hero check.’

Claire Vo: 25:12 So it built in some voice to it. It even built things like focus mode where it could start a timer and start to track the time that it’s spending, that my kids are spending, on particular homework items.

Claire Vo: 25:19 Yes, we are very fun here. Um, how many lessons, reward them about how they pursued their tasks. Finishing the quest, you get some nice little confetti here.

Claire Vo: 25:29 They then get to get available rewards. My oldest child is earning a one-on-one basketball coach because he likes coaching, so we say if you practice your piano, you get a coach.

Claire Vo: 25:38 So they put that front and center, and then it came up with different sort of like prizes they can win,

Claire Vo: 25:46 including picking family dinner, a movie and staying up late, or buying like new basketball shoes, which, man, the way these kids grow their shoe size, they buy a lot of basketball shoes.

Claire Vo: 25:54 And then same with my middle. He’s focusing on a couple different things, including playing piano. It’s actually really short what he has to do, and so it built that.

Claire Vo: 26:01 And then what I love is it gamified them together. And so if they can work together, they can earn more XP.

Claire Vo: 26:08 They also can earn like companion, I don’t know, avatars, like Beatbot and Comet Fox. They can get power auras.

Claire Vo: 26:18 They can like figure out which different kinds of subjects they’re learning. So it really went ham on some gamification.

Claire Vo: 26:26 And then, again, to the sort of like full-fledged functionality, it even gave me a parent HQ. Now, we got a little slop here with the button layout.

Claire Vo: 27:00 order on the side, but I can review exactly what they’ve done, I can turn on and off quests, I can edit how many points they get per quest, I can add things, so if I want them to start doing stuff, I can add it in here, I can change what rewards they get, again, it really listens to me, my oldest is motivated by basketball and my youngest is motivated by Minecraft and it gives me a history and other settings that that we can set. And so this is a very robust app, it built it basically one shot and put a lot of effort into the the design of it and this is something that I’ve seen from Claude now, like is this consumer grade exactly what I would ship? No, but it’s a lot better than what I’ve seen kind of one shot out of other models and I do just like the polish that it’s put in in terms of effort. So again, writing good, one shot sort of prototypes good, we’ve seen that in the in the benchmark.

Claire Vo: 28:01 Let me talk about another thing where I think GPT-4o, Sonnet and its family does a lot better than Llama 3. And I understand, I’m going to preface this by saying I understand why Llama 3 is a great cybersecurity researcher in that it is like incredibly precise, incredibly detailed, will like look at every corner and every edge and score every risk and like try to be incredibly precise. The problem is when you’re building products, exact precision is neither helpful nor possible. Like you literally, especially when working with AI, cannot be precisely deterministic when building a great product. And like understanding what a user would like is not an exercise in technical precision, it is an exercise in intuition, design, all these things and boldness and creativity and strategy and all this stuff.

Claire Vo: 28:59 And I was working on two projects deeply with Llama 3 and then with Sonnet and I just had a very much better experience unlocking with Sonnet. Let me just talk you through what those are. One was this chat PRD kind of like integrated prototyping tool where like V0, Lovable, all these things, you could take your PRD and make make a prototype and building like a good effective coding harness there and then trying to figure out what the right model was. The second thing is basically like an insights ingest product where you can like hook up Intercom and Linear and all this to GitHub and all these signals and suck them in and like basically build a product brain that’s gonna be rad.

Claire Vo: 29:40 And when I was having Llama 3 work on these, it did a lot of the like technical heavy lifting, it got the like big meaty pieces into place, but it was like a brutal scorer and it hardened these architecture of both of these products that it actually broke itself. So my example is it like had this very hardened tool calling loop in my prototyping tool and only GPT-4…

Claire Vo: 30:00 5.5 would run. Like, I could not get any other model to run. And I ran eval after eval after eval, OpenWeight, Sonnet, Opus, all of these, could not get anything but GPT-5.5 to run.

Claire Vo: 30:11 And I was insistent that this was an us problem, not the model problem. These models can definitely create frontend prototypes. And Fable was like, ‘No bro, that’s it’s, it’s totally these models’ fault.’

Claire Vo: 30:23 And as soon as I switched it to Codex and said like, ‘Look, I’m just not convinced we can’t get Sonnet 5 to work, this is ridiculous, just do what you think is correct,’

Claire Vo: 30:33 it, it fixed it and it got it actually working. Now, did it get it working perfectly? No. Do I think this is a great design? No. I’m trying to figure out what the problem is.

Claire Vo: 30:41 But in one shot, it got out of its own mind and fixed things and again, this was like such an unlock, very similar to my insights generating engine.

Claire Vo: 30:54 Fable really wanted to like score and lint this effort and wanted to like be able to deterministically figure out if generating prose could be like reproducible, always verifiable, always citationed, all these things. And at the end of the day, that wasn’t what was going to make a great product, it was just what was going to make like a code evaluation verification loop exit.

Claire Vo: 31:12 But once I’ve told GPT-5.6 and Codex like, ‘Stop being pedantic,’ I ended up getting these really useful and helpful Wiki pages generated out of this, all this structured and unstructured data. It was actually really good.

Claire Vo: 31:26 And it just, I don’t know, I don’t know what Fable’s deal was, I could not get it unlocked. But 5.6 was very willing to reconsider its own kind of limitations and build something.

Claire Vo: 31:34 I’m going to do two more quick use cases where I think GPT-5.6 is really good. I will get you out of here, go start coding. I’m basically out of model capacity anyway, so I’m going to have to take a break.

Claire Vo: 31:43 Two use cases that I think are amazing. First one is video editing, video editing. I have to do a lot of social clipping and it’s really tedious to go through and clip videos. So taking something really long and shortening it.

Claire Vo: 31:54 So recently I spoke at Cursor’s event and gave this talk on the future of PM and got the recording from the Cursor team, thank you very much, and I really wanted to make it a hype video.

Claire Vo: 32:04 So all you have to do is literally drag the file in here and I said, ‘Can you cut this video into five clips for social?’ and I gave some feedback, I said I want them horizontal, I want them hype video cuts from various parts.

Claire Vo: 32:18 I need them to be faster, I need them to be tighter, and then I got these like sharp and funny hype videos. Let’s see if it opens up. This one’s for my, my talk.

Claire Vo: 32:26 We’re going to figure out what it means to be a product manager in the age… where anybody can build anything. We have been coming up with creative ways to avoid building things forever. Yes, PRDs, like these complicated documents where you had to describe

Claire Vo: 33:13 So like, that would have taken me so much time to like find the right cute parts, clip it, cut it. I was able to drop it into Capcut, put some music, ship it on social, it’s like a really cute hype video. But this is one of my favorite use cases. I’m pretty sure it can do even more, color grading, sound, all this kind of stuff. But even just dropping videos in here and fixing things are great.

Claire Vo: 33:35 Finally, the last and best use case of… of 5.6, and I cannot believe I waited to the end to show this, is it is a beast, beast when it comes to browser use. I am deeply obsessed with letting Codex plus GPT 5.6 and Chrome, and at Chrome in Codex, if you don’t know how to do that, you do it like this: at Chrome on a logged in page and just say, “Go with the stars and do some stuff in.” And like, I’m sorry LinkedIn, I know I’m not supposed to do this, but I opened up LinkedIn and I said, “Can you use Chrome to reply to messages that are very high value to ChatPRD or the How AI podcast? Keep the bar very high. Again, I love you all. I cannot deal with all the LinkedIn requests.”

Claire Vo: 34:21 So like only accept them if they’re executives of tier-one companies, I don’t want random sets of connections. It went through and burned through probably 500 messages. It replied to people that I… needed to reply to and said thank you to people who said nice things about the podcast. Thank you to those people, I do mean it. But it just rocked through browser use. I have used it to test web apps, I have used it to fill out annoying forms. Browser use and 5.6, and when I got rolled back to 5.5, my life was worse. So please, please, please, learn to use at Chrome, at browser, and at computer and just let… let Codex rip and let GPT 5.6 rip.

Claire Vo: 35:08 Okay, that’s it. That is the very scientific How AI model benchmark, the love letter to Claire Vo’s favorite favorite model, GPT 5.6. A honorable mention to our pal Fable, who if I don’t have to talk to you, I’m actually pretty happy with your code. And a broad set of use cases I think it’s really good at. Excellent at writing web apps, the best of the AI writers unless you want it to have a personality, then that’s Sonnet. Great at unlocking sort of technical work that has gotten too complex for its own good and breaking through to the real user value, cutting videos, which I really love to do, really love to do with GPT 5.6, and using the browser. Those are the things that I would try. I would love to hear what you think about these models. I would love to hear your feed…

Claire Vo: 36:00 back if I am totally off my rocker, what I should add to the how AI benchmark we will publish all this work to the Chat PRD blog and I look forward to talking to you about the next model soon.

Claire Vo: 36:11 Thanks so much for watching. If you enjoyed the show, please like and subscribe here on YouTube or even better leave us a comment with your thoughts. You can also find this podcast on Apple Podcasts, Spotify, or your favorite podcast app. Please consider leaving us a rating and review which will help others find the show. You can see all our episodes and learn more about the show at howaipod.com. See you next time.