YouTubeFeed

Resemble AI and Flux.Kontext Rock, Odyssey Intrigues, Veo 3… Failed Again | 3 Wow and 1 Promise

Summary

This episode of “3 Wow and 1 Promise” dives into practical AI tool testing and reveals which ones deliver on their promises and which fall short. The first wow showcases Resemble AI’s newly open-sourced voice cloning model, which Ksenia tests live and finds surprisingly good quality - even noting that the cloned voice speaks better English than she does. The second wow highlights Apple’s underappreciated Voice Memos AI transcription feature, which dramatically saves time compared to traditional workflows using services like Otter.

The episode takes a critical turn when testing Google’s Veo 3 video generation. Despite being a highly anticipated model with access granted to Pro subscribers, the experience is frustrating - unintuitive interface, inconsistent results, and lack of proper documentation make it feel “not there yet.” In contrast, Odyssey’s interactive AI-generated video preview emerges as a surprise third wow, offering an immersive, explorable virtual world that hints at a new form of storytelling and game creation.

The bonus highlight covers Black Forest Labs’ Flux Kontext release, which enables powerful image editing capabilities with consistent results through iterative prompting.

Highlights

”My cloned voice speaks better English than I do”

Clip

Clip command
yt-dlp --download-sections "*1:10-1:38" "https://www.youtube.com/watch?v=TBZg86b9KdU" --force-keyframes-at-cuts --merge-output-format mp4 -o "TBZg86b9KdU-1m10s.mp4"

“It sounds like my sister became a robot. The best part about it is that the English of this voice is better than mine.” — Ksenia testing Resemble AI, 1:10

”Doctors cutting documentation time from 90 to 30 minutes”

Clip

Clip command
yt-dlp --download-sections "*1:42-2:25" "https://www.youtube.com/watch?v=TBZg86b9KdU" --force-keyframes-at-cuts --merge-output-format mp4 -o "TBZg86b9KdU-1m42s.mp4"

“I just read that a lot of doctors now use AI to record the conversations with the patients and then transcribe it and that it saves them. Usually it took them 90 minutes to prepare the documentation for patients and now it reduced to 30 minutes.” — Ksenia, 1:42

”Veo 3 - So unintuitive that I don’t understand how to use it”

Clip

Clip command
yt-dlp --download-sections "*7:03-7:30" "https://www.youtube.com/watch?v=TBZg86b9KdU" --force-keyframes-at-cuts --merge-output-format mp4 -o "TBZg86b9KdU-7m03s.mp4"

“It’s so not intuitive and not clear that a user like me who sees the computer not for the first time does not understand how to make a video with Veo 3. And it took me 11 minutes to just work with this little thing and I haven’t even received two scenes. So again, not there yet.” — Ksenia, 7:03

”Odyssey interactive video - A new form of storytelling”

Clip

Clip command
yt-dlp --download-sections "*8:22-9:25" "https://www.youtube.com/watch?v=TBZg86b9KdU" --force-keyframes-at-cuts --merge-output-format mp4 -o "TBZg86b9KdU-8m22s.mp4"

“So, not only you can move in this world, which is exciting, but you can also switch the world channels. That’s so awesome. Now imagine creating games with this outstanding format. This is truly a new form of storytelling.” — Ksenia, 8:22

”Flux Kontext - Magic image editing”

Clip

Clip command
yt-dlp --download-sections "*10:13-10:57" "https://www.youtube.com/watch?v=TBZg86b9KdU" --force-keyframes-at-cuts --merge-output-format mp4 -o "TBZg86b9KdU-10m13s.mp4"

“The coolest part about new Flux Kontext is that you can choose a picture and then you can edit it with your prompt. Magic, right? The image generation is pretty consistent.” — Ksenia, 10:13

Key Points

  • Resemble AI Open Source Release (0:16) - New open source voice cloning model that produces surprisingly good quality
  • Voice Clone Quality Assessment (1:10) - The cloned voice speaks better English than the original speaker
  • AI in Healthcare (1:42) - Doctors using AI transcription cut documentation time from 90 to 30 minutes
  • Apple Voice Memos AI (2:12) - Underappreciated feature that transcribes recordings directly on device
  • Microsoft DAX Copilot (3:11) - Medical transcription tool with safety concerns being addressed
  • Veo 3 Access (3:52) - Google gave Pro subscribers access to their newest video model
  • Veo 3 UI Problems (4:32) - Difficult to switch between Veo 2 and Veo 3, unintuitive workflow
  • Video Generation Inconsistency (6:45) - Sound and character consistency problems when extending scenes
  • Odyssey Interactive Preview (7:51) - AI-generated interactive video you can explore in real-time
  • World Channels Feature (9:10) - Odyssey allows switching between different virtual worlds
  • Flux Kontext Release (9:28) - Black Forest Labs’ new open source image editing model
  • Iterative Image Editing (10:24) - Flux Kontext enables consistent edits across multiple prompts

Mentions

Companies

  • Resemble AI (0:19) - Released new open source voice cloning model
  • ElevenLabs (0:27) - Mentioned as closed model competitor for voice synthesis
  • Apple (2:12) - Implemented AI transcription in Voice Memos
  • Microsoft (3:11) - DAX Copilot for medical transcription; safety focus at Microsoft Build
  • Google (3:52) - Veo 3 video generation model for Pro subscribers
  • Odyssey (7:51) - Interactive AI-generated video preview
  • Black Forest Labs (9:28) - Creators of Flux open source image model
  • Otter (2:34) - Transcription service mentioned as previous workflow

Products & Technologies

  • Voice Memos (2:17) - Apple’s recording app with new AI transcription
  • Veo 2 (5:25) - Previous Google video generation model
  • Veo 3 (3:52) - Google’s newest video generation model
  • DAX Copilot (3:11) - Microsoft’s medical transcription AI
  • Flux (9:33) - Open source image generation model
  • Flux Kontext (9:51) - New image editing capability from Black Forest Labs

Surprising Quotes

“The best part about it is that the English of this voice is better than mine.” — Ksenia on her Resemble AI voice clone, 1:20

“It’s so not intuitive and not clear that a user like me who sees the computer not for the first time does not understand how to make a video with Veo 3.” — Ksenia, 7:10

“Not only you can move in this world, which is exciting, but you can also switch the world channels. That’s so awesome. Now imagine creating games with this outstanding format.” — Ksenia on Odyssey, 9:06

“Your computer will probably going to melt but still you can use this model and create very very good pictures not depending on any paid versions of other services.” — Ksenia on Flux, 9:36

Transcript

0:00 Welcome back to three wow and one promise. Who needs more AI news, right? Well, you need apparently. So that’s why you’re here. Let’s dive in.

0:17 This week, a new open source model was announced from Resemble AI that creates your voice. There is a bunch of different open models. There is a closed models like ElevenLabs that allow to synthesize your voice and use it, but open source is always better. So, let’s try it. Get started with Resemble AI. Okay, clone your voice. Let’s do that.

0:40 Good morning, Resemble AI. My name is Ksenia. I’m the founder of Turing Post and I’m testing you for my news that I present every week. It’s called three wow and one promise. Let’s see what you can do. Good morning, Resemble AI. My name is blah blah blah. Cloning my voice. Ooh, that’s exciting. Oh, no. Of course, I will unlock and pay if it’s a good quality, but I need to check first.

1:10 Hey there, it’s me. Well, technically it’s you. What would you like me to say next? It sounds like my sister became a robot. The best part about it is that the English of this voice is better than mine. Hey there, it’s me. Well, technically it’s you. Hi there, it’s me. Well, technically it’s you. It’s good. It’s wow, it’s good. I think for the just released model, it’s pretty good.

1:38 To the next news, I just read that a lot of doctors now use AI to record the conversations with the patients and then transcribe it and that it saves them. Usually it took them 90 minutes to you know prepare the documentation for patients and now it reduced to 30 minutes. So again it’s a huge cut and save for human time.

2:06 And in that sense I wanted to show you which I think is highly underappreciated and I found it by a chance that Apple actually implemented some AI into their voice memos and if you use voice memos a lot like I do now you can use transcript. It’s super easy because before the flow was that you record an interview and a conversation. Then we upload it to a computer. Then you use such tools like Otter or other transcription services. It takes a lot of time.

2:38 Now you go to your old recordings in voice memos. You choose a recording. Then you choose these three dots over here and it says copy transcript or view transcript. And you can just copy to your notes and then they already in your computer. That’s a wow in my opinion because this AI feature saves my time and going back to the doctors it saves tons of their times to record the conversations with the patients.

3:06 Of course the question as you know I have questions about the safety. The doctors mostly use Microsoft DAX copilot. As we know Microsoft is on PC and PC is not that good with preserving safety as for example Apple. But I’m sure and what we heard at Microsoft Build they put a lot of attention into the safety measures so with AI tools and features rolling out to more and more people safety becomes a tremendously important topic and from what I heard Microsoft would make an effort to make it safer. So, we’ll see.

3:47 If you remember, I promised you to show how Veo 3 from Google AI works. They heard me, they listened to me, they gave access to pro subscribers to their newest model, Veo 3. So, if you remember, let’s take a look at the screen. And if you remember last time, we made this video with a peculiar bird and somehow the video updated itself.

4:13 Let’s take the same prompt that we had and create a new project, a tense moment inside a pilot aeroplane cockpit during flight. It was funny that I tried to create another video before showcasing Veo 3 to you. And don’t try using words. It’s not there yet. But I’m very excited to see what is going to show us in terms of video.

4:36 I don’t understand why they don’t talk. I don’t understand. Maybe I still don’t have access to Veo 3. I’ll try again. But today I read an article that pro subscribers of Google have access to Veo 3.

4:53 Why… it was actually my mistake. I didn’t figure out how to switch Veo 2 to Veo 3. We are going to use Veo 3 for our favorite prompt. Let’s check. It’s here. We added our prompt. We need to add that they talking. They actively talking about how to pilot the plane. I’m excited. I am looking forward to see what Veo 3 is capable of.

5:26 Okay, we have something. And it’s interesting that the pilots look almost the same as in our first prompt with Veo 2. Okay, so let’s listen how they talk. “I think we should try to turn to the left.” “Are you sure?” “Yeah.” “Are you sure that’s the right course of action?” I mean, that’s pretty good. At least they’re sitting. They’re facing the window, which I think is very important for the pilots.

5:46 “We’re losing altitude. We need to push the throttle.” “What do you mean? What do you mean?” Oh, the second one is much better. Now, let’s add something. Let’s do the bird again. My favorite bird. Extend. Suddenly, with loud noise, a bird rushes into the window. Go.

6:08 “We’re losing altitude. We need to push the throttle.” “What do you mean? What do you…” For some reason, the sound in the second video, it does not work. Maybe it was a very complicated prompt. Let’s delete it. Let’s extend it. And then the female pilot says, “Let me turn away from the window.” That’s what pilots do. I hope it has sound.

6:51 It’s not consistent. Also the characters are very different. Okay. So let’s try and then we’ll make sure that this one is highest quality. I don’t know why it doesn’t work. It’s either working but I don’t understand how.

7:03 But that only means that it’s so not intuitive and not clear that a user like me who sees the computer not for the first time does not understand how to make a video with Veo 3. And it took me 11 minutes to just work with this little thing and I haven’t even received two scenes. So again, not there yet.

7:30 And now to the promising thing. I mean, you know, Veo 3 was still a big promise that we showcased, but there is another one. There is another thing that I wanted to show you. I tried it. Look at the screen. I tried it once. It didn’t work. So I hope with you it will work.

7:51 What is it about? Odyssey just launched a research preview of interactive video generated by AI in real time. That means as I understand it that you can interact with this video and it will be like a new form of storytelling. The video that they have here is not really informative. I didn’t understand from that video how it works. But there is a button experience interactive video. So that.

8:22 Ooh. Oh, that’s much better than I used before. Before I saw just a brown blob. Now we can see a village. It’s with sound. So maybe we can walk. We can walk. Wow. Maybe that’s a third wow and not a promise.

8:40 So maybe we put Google Veo to the promise section again. I’m sorry Google. And this Odyssey interactive video. We put it to wow. You know, it’s still in the very beginning, but just imagine I just opened this interactive video. I’m in a some sort of a village. There is a music. It is very loud. And we’re walking and we can look around. It’s glitchy, but it’s just the beginning.

9:06 Oh my god. So, not only you can move in this world, which is exciting, but you can also switch the world channels. That’s so awesome. That’s really cool. So, now imagine creating games with this outstanding format. This is truly a new form of storytelling. Super exciting.

9:28 One more thing that I wanted to show you. So, this new release is coming from Black Forest Labs. They’re famous for their model Flux because it’s open source. You can download it on your computer and though your computer will probably going to melt but still you can use this model and create very very good pictures not depending on any paid versions of other services.

9:51 So this week they introduced Flux Kontext and let me show you what it can do now. It’s pretty cool. One of my favorites little prompts about a little toddler girl petting a big black dog. The dog licks the girl’s face which makes her smile.

10:13 The coolest part about new Flux Kontext is that you can choose a picture and then you can edit it with your prompt. So now I want to say move girl’s face far from the dog. Let’s see. Magic, right? Let’s choose this one. Now let’s edit again. The girl now looks straight in the camera.

10:39 The image generation is pretty consistent. And look at this. It’s just amazing what it does. Now the girl is looking into the camera. And you can download it. You can use it on the playground or you can download through API and use it on your computer. You will need a powerful computer. It’s a very cool thing.

10:59 Thank you for your attention. Don’t forget to subscribe, share, like, comment. I truly appreciate it. Let’s find more wow about AI and keep promising things rolling. Thank you.