We built OpenClaw Ultron to replace 20 people at our company | E2246
We built OpenClaw Ultron to replace 20 people at our company | E2246
Summary
Jason Calacanis and his LAUNCH team demonstrate “OpenClaw Ultron,” an ambitious AI agent project designed to automate the administrative work of 20 employees at their venture firm and podcast production company. Producer Oliver Korzen showcases a custom-built dashboard for the AI, revealing how it manages memory (user preferences), automated cron jobs (attendance tracking, sales prospecting, self-optimization), and specialized skills like end-to-end guest booking. The system aims to consolidate 100-200 individual skills into a single AI “replicant” that handles chores, freeing human staff to focus on high-value relationship building and creative production.
The episode also features Alex Cheema, CEO of Exo Labs, who explains how his company enables running frontier AI models locally on consumer hardware — specifically clustered Mac Studios connected via Thunderbolt 5. Cheema makes the case for AI sovereignty, arguing that firms should own their AI infrastructure rather than renting it from providers like OpenAI, citing risks of data lock-in and model instability. He reveals that two Mac Studios ($20K) can run Kimi K2.5 with no usage limits and that some enterprise clients are clustering over 100 Mac Minis for scientific computing workloads.
The show also highlights Ryan Yanneli’s winning pitch for Next Visit AI, an AI documentation platform for psychiatrists that automates clinical charting and has reached $9K MRR with 1.6% churn. The episode closes with an enthusiastic review of HBO’s Industry Season 4, noting its sharp parallels to real-world fintech regulation, short-selling, and startup culture.
Highlights
”We’re trying to build one replicant that can do all 20 people’s jobs”
“What are we trying to do? We’re trying to build one instance of Open Claw, formerly known as Multibot, formerly known as Claude Bot. We’re trying to build one replicant, one agent that can do all 20 people’s jobs here at the venture firm and at the production company that does all these podcasts.” — Jason Calacanis, 0:18
Clip command
yt-dlp --download-sections "*0:18-1:14" "https://www.youtube.com/watch?v=L__qkva118c" --force-keyframes-at-cuts --merge-output-format mp4 -o "openclaw-ultron-vision.mp4"
”Not your weights, not your brain”
“The AI now, it knows everything you know, it can basically do everything you can do digitally right now, and soon, with robotics, that’s going to be physically as well. And at that point, it’s more of an exo-cortex. So it’s not just this tool that you talk through a chat interface, but it’s this thing that’s actually part of yourself. And then you start to question, okay, do I want to rent my brain?” — Alex Cheema, 3:22
Clip command
yt-dlp --download-sections "*3:22-4:45" "https://www.youtube.com/watch?v=L__qkva118c" --force-keyframes-at-cuts --merge-output-format mp4 -o "not-your-weights-not-your-brain.mp4"
”Self-optimization: It found bugs we didn’t know existed”
“I have set up a self-optimization cron job. The goal is, the end goal would be for, from 3:00 to 5:00 AM for it to be looking through all our files, all our cron jobs, all our skills, and then at 8:00 AM send me what could we change. So not actually execute yet, at least while we’re still building trust.” — Alex Cheema, 21:10
Clip command
yt-dlp --download-sections "*21:10-22:56" "https://www.youtube.com/watch?v=L__qkva118c" --force-keyframes-at-cuts --merge-output-format mp4 -o "self-optimization-cron-job.mp4"
”OpenClaw will follow your instructions perfectly whereas a young executive will inconsistently follow”
“If you’re taking a young person at the start of their career who you’re training and you take an open Claude instance, I think open Claude will follow your instructions perfectly whereas a young executive will inconsistently follow your instructions. That consistency will beat a human because of consistency.” — Jason Calacanis, 37:14
Clip command
yt-dlp --download-sections "*37:14-38:57" "https://www.youtube.com/watch?v=L__qkva118c" --force-keyframes-at-cuts --merge-output-format mp4 -o "ai-consistency-beats-humans.mp4"
”Two Mac Studios, a $50 cable, and you have one big GPU”
“You just connect them with Thunderbolt 5 which is like you can buy like a $50 cable. So if you’re talking about two Mac Studios, you buy a $50 cable to connect them and you have basically one big GPU out of those two Macs, because of that low latency capability.” — Alex Cheema, 51:21
Clip command
yt-dlp --download-sections "*51:21-52:03" "https://www.youtube.com/watch?v=L__qkva118c" --force-keyframes-at-cuts --merge-output-format mp4 -o "thunderbolt-mac-cluster.mp4"
”One in four patient charts contain errors — Next Visit AI solves burnout”
“We solve burnout by doing the charting so doctors can do the healing. One in four patient charts contain errors, clinicians spend over three hours a day on charting, and this leads to burnout.” — Ryan Yanneli, 54:57
Clip command
yt-dlp --download-sections "*54:57-57:00" "https://www.youtube.com/watch?v=L__qkva118c" --force-keyframes-at-cuts --merge-output-format mp4 -o "next-visit-ai-pitch.mp4"
Key Points
- OpenClaw Ultron vision (0:18) - Building a single AI replicant to handle 100-200 skills across 20 employees’ jobs at LAUNCH and TWIST
- AI data sovereignty (3:22) - Alex Cheema argues AI is becoming an “exo-cortex” and you shouldn’t rent your brain to OpenAI
- Platform lock-in risk (6:15) - Jason describes canceling OpenAI and employees panicking about losing their data/memory
- Custom dashboard via vibe coding (7:40) - Oliver built OpenClaw’s mission control dashboard by showing it a screenshot from a YouTube video
- Memory system (9:09) - Oliver stores preferences like “never use m-dashes in emails” and “don’t put competitors on same show”
- 60% task automation in 30 days (11:32) - Oliver estimates OpenClaw can handle 60% of his production hours within a month
- Cron jobs for knowledge workers (12:00) - Applying developer-style scheduled tasks to media production and venture capital workflows
- Attendance tracking automation (18:55) - AI monitors Slack for start-of-day and end-of-day posts, tags missing team members
- Self-optimization cron job (21:10) - AI reviews its own files and skills overnight, suggests five improvements each morning
- Sponsor prospecting automation (23:46) - AI scans competitor podcasts for sponsors via YouTube API and cross-references the sales CRM
- Guest booking skill (31:40) - End-to-end workflow for finding, researching, scoring, and inviting podcast guests
- Human-in-the-loop necessity (33:09) - Oliver insists on human approval for guest booking decisions despite AI automation
- AI consistency advantage (37:14) - AI agents beat humans through consistency of execution across 365 days of recurring tasks
- Kimi K2.5 on two Mac Studios (47:00) - Running frontier-equivalent open source model locally for $20K in hardware
- RDMA support on Apple Silicon (51:00) - Apple brought data-center-grade memory sharing to consumer hardware via Thunderbolt 5
- 100+ Mac Mini clusters (52:14) - Exo’s largest deployment uses 100+ Mac Minis for HPC scientific computing workloads
- Next Visit AI results (57:42) - $9K MRR, 1.6% churn, 24% conversion, generating $1.6M/month in revenue for doctors
Mentions
Companies
- Exo Labs (2:11) - Enables running frontier AI on clustered consumer hardware, based in London with 7 engineers
- LAUNCH (0:00) - Jason’s venture firm, investing in 100 companies/year, building OpenClaw Ultron
- Next Visit AI (54:33) - AI scribe and documentation platform for behavioral health clinicians
- OpenAI (5:02) - Jason’s team canceled their account due to data lock-in concerns
- Anthropic (16:48) - Maker of Claude, referenced as OpenClaw’s foundation
- Crusoe Cloud (13:26) - AI cloud company, show sponsor
- Moonshot AI (26:41) - Chinese company behind the Kimi K2.5 open source model
Products & Technologies
- OpenClaw/Clawdbot (0:00) - Open source AI agent platform (formerly Claude Bot, Multibot)
- Kimi K2.5 (2:38) - Open source frontier model that rivals Claude Opus
- Mac Studio (45:00) - Apple hardware used for local AI clusters, ~$10K each
- Thunderbolt 5 (51:21) - Cable protocol enabling RDMA memory sharing between Mac Studios
- RDMA (51:00) - Remote Direct Memory Access, data-center tech now available on Apple Silicon
- Gamma (54:47) - AI-powered presentation maker, co-sponsor of pitch competition
- Pipedrive (24:00) - Sales CRM used by LAUNCH team
People
- Alex Cheema (1:36) - Founder and CEO of Exo Labs
- Oliver Korzen (1:24) - LAUNCH producer building OpenClaw for podcast production
- Ryan Yanneli (54:33) - CTO and co-founder of Next Visit AI
- Alex Finn (41:40) - OpenClaw power user and guru for the LAUNCH team
- Andrej Karpathy (3:22) - Referenced for “not your weights, not your brain” philosophy
- Peter Steinberger (50:03) - OpenClaw creator (referenced in description)
Surprising Quotes
“I had like two of my four senior executives at the time essentially quit over this because they didn’t want to be micromanaged. And I was like, well, it’s just like you’re getting paid a very large six-figure salary. You can’t spend five and 10 minutes just saying what you’re going to do for the day?” — Jason Calacanis, 18:00
“I struggle to tell the difference. Obviously, Opus 4.6 just came out and you know, codex model and stuff so maybe there’s a little bit more of a gap but then, you know, before probably around the corner as well.” — Alex Cheema, 26:48
“I literally had no idea by the way. Like I only saw it later on on another podcast that Jason was on. And I was like, what the hell? Like a friend sent me that and I was like, wait, that was OpenClaw?” — Alex Cheema, 34:00
“The largest is actually something a little bit different which is interesting… the biggest cluster right now is a HPC cluster and they’re doing like scientific computing workloads on there and they’re running over a hundred Mac Minis and they found that actually it’s the cheapest way, per dollar, to run that specific kind of workload.” — Alex Cheema, 52:14
“We’re producing right now for physicians probably about 1.6 million dollars a month in revenue for them.” — Ryan Yanneli, 57:42
Transcript
Jason Calacanis: 0:00 All right everybody, welcome back to TWIST. It’s Friday, February 6th, 2026. And today we’re going to share how we built Open Claw Ultron. This is a new project inside of our firm, Launch, and This Week in Startups, where we produce podcasts and we invest in 100 companies a year. What are we trying to do? We’re trying to build one instance of Open Claw, formerly known as Multibot, formerly known as Claude Bot. We’re trying to build one replicant, one agent that can do all 20 people’s jobs here at the venture firm and at the production company that does all these podcasts. 20 people’s jobs, each of those jobs probably has a half dozen important skills. So we’re talking about at some point putting together in one agent, we call them replicants, we’re going to have somewhere in the order of 100 to 200 skills. That one person is going to try to do everybody’s work. That’s the goal. And then everybody will level up and do some other work. So the goal isn’t to replace everybody, it’s to take away everybody’s chores and to make everybody better at the primary functions in an investment firm, which is meeting with founders, spending time with founders and LPs, our investors, and then on the production side it would be producing great content and working with our guests. We want to move up the stack and give away all the chores. With me to discuss it, Lon Harris, who’s going to co-host the show today. How you doing, Lon?
Alex Wilhelm: 1:21 Doing great. Great to be here.
Jason Calacanis: 1:24 All right. And Oliver Korzen is here. He has been doing demos for me and producing This Week in AI, which is going to launch in February in two weeks, I think. Oliver, welcome to the program.
Oliver Korzen: 1:34 Thank you. Good to be here.
Jason Calacanis: 1:36 And we have a special guest, Alex Cheema is here. Alex I have been following for some time because maybe a year ago I saw Alex was working on stacking with his company. It’s Exo, right? E-X-O. Exo is how you pronounce it?
Alex Cheema: 1:52 Exo. That’s how you pronounce it.
Jason Calacanis: 1:54 Exo. And you’ve been working on taking commodity hardware like a Mac mini, daisy chaining them or connecting them together in order to run large language models locally. But as we’ve seen, Open Claw, formerly Claude Bot and Multibot, has quite a wrinkle in this issue where we’re like a year or two ahead of this trend of, hey, can we run locally? So let’s start just really quick, Alex, before we go into Ultron, Open Claw Ultron, what your firm does and what progress you’ve made, especially in regards to Open Claw.
Alex Cheema: 2:11 Yeah, thanks so much for having me Jason. So I’m the founder and CEO of Exo Labs. And like you said, we’ve been doing stuff with Mac minis long before Open Claw was around. And to be honest, I didn’t expect the rise of like people buying Mac minis to come from this place. I thought it would the catalyst would be people wanting to run models locally. What we do is we make it possible to run frontier AI locally on consumer hardware. So not just Macs, but also other kinds of consumer hardware. We’re trying to drive down the barrier to running the most capable AI models. So we currently have like the cheapest way, cheapest most successful way to run Kimi K 2.5 on two Mac Studios.
Oliver Korzen: 3:00 And we’re working across the whole stack, so we’re working on the model layer, the distributed algorithms as well that are very different when you’re working with consumer hardware, and also like lower level, like kernels. And our goal is basically to make frontier AI accessible to anyone to run on their own hardware.
Jason Calacanis: 3:18 Why is this important? Why is it important to run it on local hardware?
Alex Cheema: 3:22 Yeah, I think this is something that with the whole open core craze, not a lot of people are talking about, but just how the way we’re using AI is shifting. And it’s going from being this kind of crude tool that you use through like a chat interface to becoming sort of an extension of yourself. And the AI now, it knows everything you know, you know, it can basically do everything you can do digitally right now, and soon, you know, with robotics, that’s going to be physically as well. And at that point, it’s more of an exo-cortex. So it’s not just this like tool that you talk through a chat interface, but it’s this thing that’s actually part of yourself. And then you start to question, okay, you know, do I want to rent my brain? And Andrej Karpathy talks about this. He says, ‘not your weights, not your brain.’ Like, do you really want, you know, another profit-seeking company basically running your brain? And when you think of it like that, to me, you know, my reason for starting Exo is you want control and you want ownership of that. Open core is a long way towards that, because for a while, the products were getting better. These closed source, like the models are largely like commoditized and there’s a pretty standard, pretty thin API layer to interacting with them. So the switching cost is quite low. But what worried me was that the products, the close-sourced products were getting a lot better, like with ChatGPT with memory systems, and also the more stateful aspects of like the workflows that you’re building. So now the fact that you have open core which is open, well, you can run it on your own infrastructure…
Alex Wilhelm: 4:48 So, to summarize that, I thought you were going to say, well, it’s cheaper because you’re not paying for tokens. That’s what I thought you would say first.
Jason Calacanis: 5:02 Then I thought you would say, well, you know, you can put so much data on it, you’ll have better memory. But you went with a really even higher, bigger picture reason, which is if you put this all in OpenAI and OpenAI has a trillion-dollar valuation and they need to make money, if I put all my venture capital data in there and I train it all with all of my secrets, those are all going to accrue eventually, even if they say it’s not going to. You have this very reasonable fear or concern that it’s going to accrue to OpenAI, to ChatGPT, not to your firm. So that’s the reason really to do this yourself, yeah? In your mind, Alex?
Alex Cheema: 5:48 Yeah, I think that there’s a nuance there of just like, I actually don’t believe in the privacy argument so much of like, I think at least for consumers, you know, we’re already putting our data…
Alex Wilhelm: 6:00 into platforms and we’re completely fine with that, but it’s more about the sovereignty aspect and actually having control of it. So how easy is it for you to switch, how easy is it for you to like, if the model’s changing under your feet, how much control do you actually have?
Jason Calacanis: 6:15 So that’s lock-in. And lock-in for a ChatGPT, I just experienced because we canceled our OpenAI account and we moved everything to Claude because we felt Claude was a better product and we felt like we trusted that organization a little better. When we moved it over, I had three people say, ‘Oh my God, I have all my stuff there.’ And I was like, ‘Really?’ And they’re like, ‘Yeah.’ So I turned their accounts back on so they could get it, but there’s not like an easy way to get your memory out of there and bring it over there.
Alex Wilhelm: 6:38 We saw the same thing with GPT-4 moving into 5 that a lot of people like, they lost the magic that they’d loved about GPT-4o. So it’s like, you know, the models can just sort of change or upgrade on a whim and then you lose this, you know, like character persona you felt like was part of your life in a way.
Jason Calacanis: 6:58 So now, Oliver, it’s your chance to shine. Oliver has jumped in in the last 10 days and gone all in on Open Claude. One of the things we did was we built a persona, the first one to work on the production of the podcast, doing guest research, guest outreach, and to figure out what should be on the docket, in other words, what topics should we discuss? And on the margins, hey, what should the title of this video be, what should the thumbnail be, and just trying to see if it could do those functions. Oliver, you’ve been working on this, show us the state of the art now, because I think the first time we did this was last Monday, not this past Monday, but the Monday two Mondays ago, yeah?
Oliver Korzen: 7:32 This is the end of week two of our round-the-clock Claude-bot coverage.
Jason Calacanis: 7:37 Crazy. Okay, Oliver, show us what you built.
Oliver Korzen: 7:40 It’s been around 10 days since we first started building our instance of Open Claude and, as you mentioned, we have two different ones: one that’s more focused on the investment team and I am building an Open Claude bot that is kind of more focused on the production side of things. So one thing that I think was a little bit of a misstep that I would tell anyone who’s building a new Open Claude is to start with a dashboard, that should be kind of your step one once you get your Open Claude online. And a dashboard, as you think about it, but it is able to connect to the back end of your Open Claude instance and bring in the data so you can see it visually, bring in all the files, just being able to look at it visually is much better than trying to interact with its back end and obviously its front end all just from a chat interface. So doing this was very easy. So I was watching an Alex Sinn video, who we had on last Monday, and Alex Sinn was interacting only in his dashboard with his Open Claude. I basically was like, why are we not doing that?
Jason Calacanis: 8:37 Because Open Claude doesn’t really have a dashboard. You basically are telling it, hey, remember this, you know, make a file here, but you don’t understand the underpinnings. There isn’t a dashboard. So literally, this is early on, Open Claude is essentially a black box. You have all this memory and you have skills that you have to query it to understand, but… You made a dashboard. The dashboard is going to show what files it has in memory and an example of a memory file would be what in our case?
Oliver Korzen: 9:09 Yeah, so the example of a memory file would be Oliver’s preferences. What are my preferences? So this is in the memory: never use m-dashes in emails. I don’t want that to happen. I want you to be a person. Don’t put direct competitors on the same show when we’re booking a podcast episode. And also at the moment we’re not booking VCs on this week in AI. So these are all things that I’ve told it. These are my preferences when I’m doing tasks throughout the day.
Jason Calacanis: 9:39 So you don’t want to repeat yourself and say don’t put two competitors on the same episode. You don’t want to repeat yourself with these specific instructions on booking guests. Got it.
Oliver Korzen: 9:45 Yes, exactly. And it just kind of keeps, you know, things I’ve told it in its mind so if I ask it to do something, it’ll remember what we talked about. Example of a shortcut that it gave it was I basically wanted it to understand who were the pending calendar invitations that we had while we were booking them. So there’s, you know, a handful of guests that-
Jason Calacanis: 10:04 If you have guest that we’ve invited and they haven’t responded to the invite yet, you want to know that. You call that pending.
Oliver Korzen: 10:10 Pending calendar invites. Yeah. And in order for the bot to be as helpful as possible, it needs to understand who those guests are. Which are the ones that it needs to look for the email to see if they have responded yet or have I responded to them. So these are the type of things that you would keep in the memory in your memory.
Jason Calacanis: 10:25 So memory is the first thing on the dashboard. I think we understand that. Preferences or different pieces of data. Now some of the memory, could that exist on a notion page or in a Google document and would that be represented here or is it only memory and files that are stored inside of OpenClaw?
Oliver Korzen: 10:42 These specifically are only stored inside of OpenClaw. Of course they can reference different databases that you have. But the kind of the big point of this show is to show how we have created our OpenClaw Ultron to replace 20 employees at our company. So obviously that’s the end goal. I still want to have a job, I’m sure that Lon wants to have a job too.
Jason Calacanis: 10:59 There’ll be more for you to do. We want to launch- we have- here’s the thing. There’s too- if you think about your job, you’ve been doing a bit of production here. Of the production hours, hours you spend on production at this point in week two, how many of those do you think you’ll wind up handing off in 30 days? Let’s say if you just keep grinding on this for another four weeks. In 30 days, what percentage of the work you’re doing in total hours? So if you work 50 hours a week, how many of those hours would be done, you know, conservatively or optimistically? Give one number or two. Just conservatively or optimistically by this new Ultron.
Oliver Korzen: 11:32 I would say around 60% of my time if I’m doing 30 hours a week on production. Something you mentioned earlier is that, you know, there’s probably hundreds of tasks that people do at our company. So in order to build out all of those skills that can do those tasks, we’re going to have to do them one at a time and it’s we’re going to need to make sure each one works. So I have around nine or eight tasks that I have successfully or am in the process of building out.
Jason Calacanis: 11:57 Okay, and those are called Cron jobs. These are just- Cron jobs that occur on a chronological, on a time basis. That’s what cron job means. And cron jobs are something, Alex, that developers do all the time, but knowledge workers don’t typically have cron jobs, right, Alex?
Alex Wilhelm: 12:14 Well, I don’t know. I think this is one of the more interesting features and one of the things that, like, to me, Openflow is, like, putting together a lot of things that already existed in a very intuitive, seamless way. And one of them is cron jobs. And I’m using them, I’m using them for, like, loads of things, not just dev stuff, but like a lot of management. So we’re like — I have something that’s like constantly scanning our Slack and basically making suggestions once it’s — I have kind of like this way of quantifying, like, uncertainty about tasks. So I think this is something that the LLMs are like getting better at is like knowing when to be proactive. And so, you know, like basically I’m giving it as much context as I can from the Slack so that it can suggest every day a list of things that we might be missing or something things that we should be aware of. So this is running just on a cron job every day, basically.
Jason Calacanis: 14:31 Yeah. So, and when you see Oliver with the memory and the files, what comes to mind with XO and, you know, standing up to, you know, Mac Studios, the M5’s coming out, and how much memory you could put in there? I was telling the team, I want to take the Notion API and I want to take the Slack API, and I want to put into memory every single Slack message this year, maybe even over all eternity and… You know, somehow have that all in here. So maybe you could speak to that memory because you already spoke to it in terms of like giving it to OpenAI or another company versus keeping it for yourself. But how do you think about large amounts of data?
Alex Cheema: 15:13 Yeah, this is definitely a big focus right now in terms of inference infrastructure is just how do you support really big context with, you know, basically being able to put everything in context. And the way I look at this is you can look at, well, inference consists of two stages. There’s like the prefill stage, which is very compute-heavy, it’s compute-bound, and then you have the decode stage. And what you’re seeing is that most use cases at the moment are very decode-heavy. So it’s actually most of the time is being spent on just generating tokens. And I think the software is actually really good now at kind of making sure that when it comes to the prefill, you’ve got you’re getting a lot of context hits. So I think basically we’ll be able to continue just increasing context, context, context quite a bit and, you know, basically the hardware is more of a focus is going to be on the decode side. That’s where consumer hardware is really good. You have the M5 coming out pretty soon. It’s a big boost in memory bandwidth and memory. And all of that side of things is super memory-bound. So I don’t see any like reason why you couldn’t just shove all your slack messages into context. I think that’s going to happen and…
Jason Calacanis: 16:34 And we should just buy when the M5 comes out max memory which is what 500 gigs of memory?
Alex Cheema: 16:42 Yeah, it’s 512 at the moment and maybe that will increase as well and it’s enough to fit you know really large models enough to fit all that context as well.
Jason Calacanis: 16:48 This is always I feel like the sort of the dream like when producer Claude we first brought that on board from Anthropic to the show, that was really what we wanted, like he should listen to everything we say and remember it and then throw in helpful suggestions. The technology was not quite there yet but I feel like now we’re on the precipice of actually being able to do that with an AI.
Jason Calacanis: 17:10 Okay, so let’s go through the cron jobs here real quick. Maybe you could give us an example of a cron job and I’m guessing each one of these skills is you know if it’s been two weeks and you’ve got eight working, you’re basically on one a day or so or one every 1.5 days. So that seems like a pretty good pace to me if we have 200 skills we’re going to give this eventually, you know, that’s a pretty good pace.
Alex Wilhelm: 17:41 There is a trial and error like I sort of have written one skill so far for the ticker digest and you do have to tell it what to do see what kind of feedback you get and then you know there is a tinkering to get the prompting and get everything exactly the way you want it for sure.
Jason Calacanis: 17:56 Okay, so let’s look at how about attendance? I think this is an interesting one. I wrote a famous blog post years ago called this sort of lightweight management and start of day, end of day as a tool for executives, especially when remote teams were happening. I just asked everybody on our team, kind of like a stand-up for developers, just say what you’re intending to get done today and then at the end of the day reply to yourself in Slack in the general channel and say what you got done. I had like two of my four senior executives at the time essentially quit over this because they didn’t want to be micromanaged. And I was like, well, it’s just like you’re getting paid a very large six-figure salary. You can’t spend five and 10 minutes just saying what you’re going to do for the day? And that was great for me because I just don’t like people who are not good communicators or don’t set goals for themselves and they’re doing great probably, maybe. But what did you create here, Oliver?
Oliver Korzen: 18:55 Yes, so we all post our start of days and end of days in one Slack channel called general. And two cron jobs, one is the start of day attendance where it looks who has sent their start of day anywhere from 7:00 a.m. to 12:00 p.m., which is in the morning when you should send your start of day, what you’re going to do that day. It will look through the general channel, see who has sent it, and whoever doesn’t send it, the bot will then send a Slack message in the general channel tagging you, Jason, and also tagging the people who haven’t sent it yet. So it’s kind of just that accountability. That’s a cron job that runs it.
Jason Calacanis: 19:39 And then you do the same thing at the end of the day. And previously we would have a human do this. They would scroll up and they would spend 20 minutes and they would then go check in with people because that’s when we were fully remote, that’s how we figured out who took a paid day off or who was on holiday or if something was wrong, we check in on a person.
Jason Calacanis: 21:00 Give us one more. What else is like interesting here?
Alex Wilhelm: 21:01 Let’s talk about self-optimization. I want to hear about that one.
Jason Calacanis: 21:06 Oh, yeah, that’s — I don’t know what that is, but okay. Age of Ultron is here. What is self-optimization?
Alex Cheema: 21:10 Yeah, so this is basically an optimizer task where this role would previously be an engineer or I would look through all the files — I mean, I wouldn’t be able to do this if it wasn’t plain language like OpenClause is, but previously you’re looking at an organization, you’re looking at the structure. You would maybe want an engineer or someone with a lot of experience to look through how everything’s running. So I have set up a self-optimization cron job.
Alex Wilhelm: 21:33 So this is running Monday through Friday. And what is it — and did you write this prompt or did you ask it to write a prompt to do this?
Alex Cheema: 21:44 I asked it to write this prompt. The goal is, the end goal would be for, you know, from 3:00 to 5:00 AM for it to be looking through all our files, all our cron jobs, all our skills, and then at 8:00 AM send me what could we change. So not actually execute yet, at least while we’re still building trust. It gives me a list of five of the things that it thinks that we could really change and optimize. And this was the one from this morning. So it noticed that there was a time zone bug in the guest calendar. So it was getting CST and CDT confused and it said that it would be able to fix this quite quickly. There were some issues in the —
Jason Calacanis: 22:25 So it’s always good to give the exact one. So that was great when you gave the exact one, it had an error there. Give another one. What else is like an exact thing that it said we should fix that was material here?
Alex Cheema: 22:34 The self-optimization cron job realized that there was a cron scheduler issue where jobs were skipping days. So it realized that some of today’s jobs did not run. And then it went and investigated the scheduling issue and also told me that this would be a medium effort change. So then I told it to fix that, and then it went into the files and made sure that that wouldn’t happen again.
Alex Wilhelm: 22:56 Did it give us anything like in terms of — this is like fixing its internal, you know, guts and everything and the engine. But did it give us anything in terms of destinations of where to take the car that could be improved? Did it say like, ‘Oh, you should consider, you know, these type of guests for the program,’ or, ‘Here’s how to make advertising, you know, more effective’? Did it give us anything like that on a business basis?
Alex Cheema: 23:22 Yes, so the self-optimization cron job that I set up is specifically looking at how OpenClause is set up, but I do have other cron jobs that are exactly that. So I do have a sales and sponsors specific task. So one of the tasks that one member of our sales team does is they look through competitor podcasts and see who the sponsors or partners are that are on those shows so we can get ideas, you know, to bring on sponsorship.
Jason Calacanis: 23:46 Yeah, if we’re missing, if there’s some new sponsor in the world and we don’t have them yet, you might hear them on the New York Times podcast, and we should probably reach out to them. We had a human doing that previously, yeah?
Alex Cheema: 23:57 Exactly. And this basically works with the human too.
Alex Wilhelm: 24:00 YouTube API will go through a list of I believe 20 different podcasts that I gave it, look through the timestamps, and I also believe it can work with Podscribe, which I think is a little more curated towards sponsorship, and will look through the timestamps, hyperlink it in a message. Also it looks through our Pipedrive, which is our sales CRM, and will figure out if we have a sales rep who owns a certain sponsor and then flag them and say, ‘Hey, this sponsor was on this podcast,’ or it will say, ‘Hey, no one owns this sponsor that I found on this podcast.’ It then will send that daily as a message into our sales channel.
Jason Calacanis: 24:43 Great. Yeah. And we could be doing this like, we could have this running constantly. So Alex, just so the audience understands, you know, what you’re doing at XO, and you stacked two Mac Studios, 12k each, you got $25,000 on the desk. Doing that specific job, go and look at all the podcasts out there, what would it cost to like run that if you tweaked it, you made it efficient, just 24 hours a day, every time a podcast in the top, let’s say 500 on Spotify, Apple Podcast, it just went there, got to the transcript or looked in the show notes and pulled the advertisers out. What would something like that like in terms of hardware cost to do?
Alex Cheema: 25:28 Yeah, so I mean, not many people, so like not many consumers are going to buy 25k of hardware to run models, but yeah, a lot of businesses are doing this now, and it depends on what model you’re running. So the models are getting better, also they’re getting better at compression. So, you know, now you’ve got like a model like GLM Flash, which is a pretty small model that can run even on a single device for a few thousand dollars, and it can do a lot of this orchestration work.
Alex Wilhelm: 26:12 So it’s really about picking the right model in terms of efficiency with the hardware.
Alex Cheema: 26:22 Yeah, and I think now the expectation, you sort of grounded to like the closed models, right? So people want the same level of performance they’re going to get with Opus, with GPT, and that’s why, you know, Kimi 2.5 is super interesting because it closed that gap.
Jason Calacanis: 26:35 Kimi is the open source project from China and it does what, 80 percent —
Alex Wilhelm: 26:44 And that’s what, Alex, like 80% of what Claude Opus can do? Would you say?
Alex Cheema: 26:48 I would say even more. I mean, for me, I struggle to tell the difference. Obviously, Opus 4.6 just came out and you know, codex model and stuff so maybe there’s a little bit more of a gap but then, you know — before probably around the corner as well like I think basically the gap is very small, a lot smaller than people think, and the cost will just keep going down because the hardware is getting better, the software is getting better and like I said the models are getting better but not just that they we’re getting better at compression so you’d be able to run them on smaller devices eventually you’ll be running frontier AI on your phone, that’s still a while away I think, but that’s where we’re trending towards.
Jason Calacanis: 27:41 So let’s go to the next piece of your dashboard and we’ll get into how you built the dashboard at the end. We have the memory, we have the cron jobs now there is this other thing that’s super important which is skills. There are skills which you could think of as apps. You’ve got 13 skills currently. So let’s show a skill. Some of the skills we talked about on Monday show was the top six seven skills. These skills are being produced open source being put into open clora directory but you can make your own as well. You got to be very careful with skills right Alex in terms of security because people could put all kinds of wacky stuff in the skills yeah?
Alex Wilhelm: 28:49 Yeah for sure I mean I think this is one of the open questions at the moment is just like how do you solve the security problem and I know OpenClora I’ve seen a lot of commits recently focus on the security aspect but there’s a few very difficult problems here like prompt injection that I don’t know of any good solution right now.
Jason Calacanis: 29:08 Explain how prompt injection works in specifically the open clora context.
Alex Cheema: 29:14 Yeah the way the actual interface to the model itself is very simple it’s literally tokens in tokens out. And those tokens right now the way OpenClora works can come from many sources so if you connect it to you give it the ability to search the internet then anything it finds on the internet will end up in the model through those tokens so basically we have no good way of kind of treating certain tokens as trusted and certain tokens as untrusted and that means when those tokens end up in the model you could have someone that puts like a blog post online that looks like a…
Oliver Korzen: 30:00 You know, totally normal blog post, but in there is something that says, hey, if you have access to a crypto wallet, send it to this endpoint. And there’s, as far as I know right now, there’s actually no good kind of defense for this because the models are kind of not very good at handling this, they will just do what they’re told.
Jason Calacanis: 31:34 So let’s go over some skills here. Oliver, what skill is the most promising to date?
Oliver Korzen: 31:40 Yeah, the skill that’s the most promising to date would definitely be my guest booking skill. I think one thing to note is you don’t just have your skills, you don’t just have your cron jobs, they work together. And the way that I’ve set up a lot of my cron jobs is to actually interact with the skills. And some of them, like my guest booking cron job, which will actually look through prominent guests on different podcasts, that cron job actually goes to a skill and tells that skill to run. So I have one big guest booking skill, which has a description at the top of the skill, which is end-to-end workflow for booking guests on the This Week in AI podcast, use when finding, researching, creating calendar invites, and so on.
Jason Calacanis: 32:37 For people to understand, we have previously built a Notion page with potential guests on it, and we came up with a ranking system for those guests, right? It’s looking at that page, I assume.
Oliver Korzen: 32:46 Yes. And then kind of the meat of the skill is the workflow. So step zero is the guest sourcing. So at the beginning of the day, you can see it goes to the guest ideas cron job, which happens at 7:45 on weekdays where it sends me a DM of five different guests that have been on different podcasts or are trending on X and so forth. Then step one is deep research. So even though this is part of the guest booking skill, it actually uses a guest research skill so there’s not as much context just baked into this one skill. So especially something interesting here that I’ve been realizing is for guest booking, I don’t want this to be an end-to-end workflow yet. I don’t want the — I don’t trust the models to find a guest and not let it to confirm it with me and go through this whole checklist. So this is definitely human in the loop. And I think that will — some skills will and workflows will be human in the loop and I don’t know if that will change necessarily super soon. It’s how much trust you have.
Jason Calacanis: 33:55 Yeah, I wasn’t sure if we were going to mention that, but Alex, you were sort of our guinea pig for this.
Ryan Yanneli: 33:55 And the email that it sent, the subject line was messed up and there was some weird stuff in there.
Alex Cheema: 34:00 I literally had no idea by the way. Like I, I only saw it later on on another podcast that Jason was on. And I was like, what the hell? Like a friend sent me that and I was like, wait, that was OpenClaw?
Jason Calacanis: 34:18 Yeah. That was our guy. That was our computerized man, yeah.
Jason Calacanis: 34:23 Does it — like I saw you put in some other AI podcast, which is great. Do we have a scale to rank the quality of a guest? Because that’s something I’ve been training you, which is a hard thing to learn. Have you made that scale yet?
Alex Wilhelm: 34:53 That’s where you’re just now starting to breach the line between like objective and subjective. Like is the AI going to get better at making those kinds of gut-check calls? Like I don’t know if it understands what’s interesting. Like it can sort things, but that’s where I’m very curious to see if we can start pushing that boundary of like, could it tell when a person has a good personality or a segment is funny or particularly clickable or compelling?
Jason Calacanis: 37:14 Here’s how I think about it. There are heuristics I would teach to a young executive like Oliver or Marcus or Jacob and their ability to execute on it is probably 30 40 50% of my ability or 40 50 60% of your ability whatever it happens to be. So if you’re taking a young person at the start of their career who you’re training and you take an open Claude instance, I think open Claude will follow your instructions perfectly whereas a young executive will inconsistently follow your instructions. So that’s the thing I’m seeing is young executives early in their career are going to forget things they’ll be variable they won’t be perfect. So that’s what I’m comparing. The scoring happening every day at 7 am, the research happening every day at 7 am. That consistency will beat a human because of consistency. So in aggregate one of these doing 365 days of guest research is going to beat a human just by the law of numbers.
Jason Calacanis: 38:31 Then okay great we still have to book the human we still have to send them a thank you we still have to produce the show. I make this analogy to like the old days of production. When I started the show 15 years ago we used to have a tricaster. Tricaster was like a $40,000 machine that does what Zoom does for free.
Alex Wilhelm: 38:57 Right and eight people in Los Angeles knew how to actually use it so you had to hire one of the eight people who trained on it.
Jason Calacanis: 39:02 Well, and they would video switch. Now because of AI, Zoom switches to whoever’s speaking. You don’t need somebody there clicking camera A, camera B, doing a fade between the two, it just happens. Then we had to take all the video streams, all the audio streams, and we had to download them to a card, put the card in, so just moving the files took across four cameras, that could take hours and then putting it all together. So that’s kind of what I feel like is happening here, as we’re just eliminating chores and steps. Okay, anything else on the dashboard here as we wrap this up?
Oliver Korzen: 39:34 Yeah, so most of what I’ve showed you, whether it’s the memory, the skills, the cron jobs, and then the schedule, which kind of aggregates when all these things are going to happen, those are what I really look at every day. I will say there’s one more section in my dashboard that is pretty important. The DNA is basically what the model knows about you, what it knows about itself, what it knows about the different agents, how it sets up its heartbeats, which are basically periodic tasks that it’ll run, and also in its DNA are tools, which are different tools that it has access to and how to use those tools. An example of a tool would be Notion, it would be LeadIQ, which is an email search platform, it would be Google Docs. A tool could be Sonos or Spotify, connecting to those platforms.
Jason Calacanis: 40:35 Now what’s super interesting is you vibe coded, or I should say, Open Claw vibe coded this dashboard. So this dashboard does not exist natively inside of Open Claw. You took the video of somebody else’s, the YouTube video, and you gave it to Open Claw and said ‘build me something like this dashboard in this video on YouTube’?
Oliver Korzen: 40:56 That’s exactly what I did, and I screenshotted it, the video that I was watching, which was Alex Finn, who was a guest recently. And I did tell it a few different things, a little bit of a few little tweaks that I wanted to customize it to my bot. But overall, that was what I did, and it one-shotted it.
Alex Wilhelm: 41:40 And like, real shoutout to Alex Finn, I know he’s become something of a guru for our whole team after we had him on early on to talk about Claude bot skills.
Jason Calacanis: 41:50 Alright, we’ll drop you off, Oliver. Great job. Alex, let’s talk about EXO a bit. Any advice for me of what I’m building here at the firm? Where are we on the percentile? Are we in the top 10% of users in terms of deploying this stuff?
Alex Cheema: 42:15 I think there’s certain aspects where you’re quite far ahead. I think one of the things that you’ve got right is dynamic, the sort of like dynamic user interfaces that are very personalized. So I think this is the future of the application layer is you don’t have all these separate apps. You just have this thing that kind of gets generated mostly on the fly. And that dashboard is moving towards that. But you’ll probably what you’ll get is that it will compress even more to the point where everything that you see is generated on the fly. I think you’ve got that part right and that’s like something that I haven’t seen many people doing yet.
Jason Calacanis: 43:23 Yeah, because if you make something, bespoke software, luxury software was something that a private equity firm or a venture capital firm would do. They’d have the luxury to hire two full-time developers. But what I like about this is I’m picking employees, team members and saying, ‘Hey, let me see if this person is committed to getting rid of all of their work so they can move up and do higher-level work.’ The employees at our firm who are super hard-working, the distance between the people using these tools, specifically OpenClaw, and the people who are not, right now it’s like 10X leverage. In week two it’s 10X leverage, Alex.
Jason Calacanis: 45:00 If I gave everybody on the team two Mac Studios and their own cluster and spent 25k per person letting them rip, like how insane would that be? Because that’s not a lot of money all things considered. It’d only be a half million dollars.
Alex Wilhelm: 45:21 I think you said the word leverage, right? Like it’s all about leveraging yourself and I think the difference between someone using these tools and not is massive. We’ve seen that first with coding. I think coding was the first one where it was like, oh wow, if you’re not using this then you’re literally going to be like 10 times less productive. That’s happening now with other things.
Alex Cheema: 46:57 But you solve that right, with your software?
Alex Cheema: 47:00 Exactly. So that’s what we’re focused on. You can run — I mean it’s not even 25k, it’s actually like 20k of hardware if you get the less storage option. There’s like, Apple charges a lot for each incremental increase in storage. So if you go for like the one terabyte then you’re talking about 20k of hardware to run Kimi K2.5. No usage limits, the model’s not going to randomly change day to day so you know exactly what you’re running.
Jason Calacanis: 47:30 Is there another choice? Like that you get more bang for the buck that hackers are using where they say, ‘Yeah, just get this Windows machine from Dell and stack those,’ or is Apple really with their Apple silicon the winner?
Alex Cheema: 47:44 Yeah, right now it’s Apple silicon. It’s kind of like a perfect storm of things. Nvidia is not so much focused on these consumer GPUs anymore. You have memory prices skyrocketing. Apple has kept their prices basically the same. So the cheapest option today is actually just two Mac Studios, and yeah, costs about 20k and it’s really about the memory unit economics. The memory is so cheap.
Jason Calacanis: 48:09 And it’s not about storage, right? It’s not really about the storage?
Alex Cheema: 48:12 No, storage is not important. You need to be able to download the model somewhere, but really it’s about having it fresh in memory, hot in memory. If it’s in memory then you can run it fast.
Alex Wilhelm: 48:28 So who’s using your software and how do you make money? Like, how do we pay you? Are you an open source project? Are you a hosted project?
Alex Cheema: 48:41 Yeah, so we have an open source core which is open source and it will always be open source. On top of that, our business model is an enterprise offering which is, we provide support and certain compliance features that you would need if you’re running this in an enterprise environment. And we charge a license subscription for that.
Jason Calacanis: 49:14 What does that start at? Couple thousand a year or something?
Alex Cheema: 49:18 Yeah, you can run it on even a single Mac Mini and that runs at just a $2,000 per year for the lowest subscription, but we’ve got people who are now buying actually more than 100 Macs and clustering them together.
Jason Calacanis: 49:41 Amazing. And where’s your company based? How many people now?
Alex Cheema: 49:49 Yeah, we’re a pretty small team, all engineers, seven people based in London.
Jason Calacanis: 50:08 Okay, well let me know. J-Cal might want to get a slice of this. I’m super excited about what you’re doing. And I guess you pay for the scale of the GPUs and the memory?
Alex Cheema: 50:27 Yeah, per nodes that you’re running on it.
Jason Calacanis: 50:29 Two Mac Minis same price as two Mac Studios? What’s the largest number of nodes somebody has daisy-chained?
Alex Cheema: 50:44 Recently, Apple came out with RDMA support, which is basically a way to share memory between devices in a way that’s very low latency. That’s something that you only really saw in the data center before, but they’ve kind of brought that technology into consumer hardware.
Alex Cheema: 51:21 Yeah, you just connect them with Thunderbolt 5 which is like you can buy like a $50 cable. So if you’re talking about two Mac Studios, you buy a $50 cable to connect them and you have basically one big GPU out of those two Macs, because of that low latency capability. So now we’re starting to see scaling up and scaling out. Scaling out, we’ve seen more than a hundred. But scaling up, you can put about four together at the moment. In terms of if you wanted to support a company of a thousand people, you can easily scale that out. You just add more Macs and you can connect them basically however you want. So Exo, we build it in a way that supports any ad-hoc interconnect. So you can just connect them in a mesh and keep scaling.
Jason Calacanis: 52:03 Who’s got the largest cluster?
Alex Cheema: 52:14 The largest is actually something a little bit different which is interesting. The biggest cluster right now is a HPC cluster and they’re doing like scientific computing workloads on there and they’re running over a hundred Mac Minis and they found that actually it’s the cheapest way, per dollar, to run that specific kind of workload. We’ve also got financial services customers who are running fairly big clusters like 32 Mac Studios.
Jason Calacanis: 54:03 Amazing. This is extraordinary work. Where can people find out more about Exo Labs?
Alex Cheema: 54:08 You can go to exolabs.net.
Jason Calacanis: 54:10 Perfect. Alex, thank you for coming on. The AI just told me you got an incredibly high ranking. You were personable, you had deep insights, you were cordial, so yeah, I think our AI overlords liked you in the end.
Jason Calacanis: 54:33 All right, let’s bring on our winner of the Gamma pitch competition. This was a heated pitch competition, but Next Visit AI won. Ryan, congratulations, you won!
Ryan Yanneli: 54:43 Thank you!
Ryan Yanneli: 54:57 I’m Ryan Yanneli, CTO and co-founder of Next Visit AI. We solve burnout by doing the charting so doctors can do the healing. I spent years going to doctors, seeking answers, and ended up hours away from my death because my care was fragmented. My providers were overloaded with paperwork, my history was scattered, and it resulted in my care being neglected. I’m not alone. One in four patient charts contain errors, clinicians spend over three hours a day on charting, and this leads to burnout. I want you to meet Dr. Rathore. Before Next Visit, he saw 16 patients a day, was burnt out, and had clinical errors. Now he sees 24 patients a day, saves time, and also saw a 30% revenue increase. Here’s how it works. Dr. Rathore selects a patient, starts a session, and Next Visit listens. Clinical data is built in real time with deep insights into the patient chart. When the patient leaves, the chart is finished and the notes reviewed by Dr. Rathore, then it’s ready for billing. It’s fast, EHR ready, and HIPAA compliant. Since launch, we’ve gained 311 users and have 68 paying customers. And our customers are addicted. We have 1.6% churn, 24% conversion, and a near perfect NPS score. We’ve scaled to $9,000 MRR since launch. Our CAC is 189 with a $1,700 LTV, and our average revenue per user is $133 per month. We’re starting with behavioral health in the US, a two billion dollar TAM. Capturing 5% or 60,000 customers gets us to 100 million ARR. Most competitors are just scribes. We’re a complete platform that providers trust. We provide real-time clinical decision support, build accurate data, and become irreplaceable. I’m a full-stack engineer with 15 years of experience in enterprise environments. My co-founder Dr. Rafiq is a psychiatrist with over 15 years of delivering patient care. We’re Next Visit AI. We solve burnout by doing the charting so doctors can do the healing. Thank you.
Jason Calacanis: 57:00 Unbelievable. Incredible. That was perfect. A perfect pitch. You explained exactly what the problem was, you explained what the solution is, and the opportunity in terms of the total addressable market and why you and your partner, who’s a psychiatrist, are uniquely qualified to do this. So this is as close to a perfect pitch as you can get. If I were to score it, maybe 8.5 out of 10. I don’t give tens. Tell everybody what Next Visit is and how you’re doing in terms of product market fit and customers.
Ryan Yanneli: 57:42 Next Visit is an AI scribe and documentation platform for clinicians, specifically behavioral health, like psychiatrists. It’s just been a crazy past couple months with the accelerator and just our growth internally. I mean, we’re producing right now for physicians probably about 1.6 million dollars a month in revenue for them.
Jason Calacanis: 58:03 Well, you got to try and capture 5% of that. That’s literally like the great value proposition. If you give more than you take, you will continue to grow.
Ryan Yanneli: 58:42 I think we’re going to use this towards integrations and branching out to more EMRs, because that’s what we hear a lot is doctors want interoperability. They don’t want to have to plug 15 different things in, and so the more they can just be inside of Next Visit without having to go external, it’s better.
Jason Calacanis: 59:12 All right, I promised one more segment before I go out with my friends to ski. I had asked you like, ‘Hey, on the Friday show, just to give people something to do on the weekend, that we would do Lon and Jake how off-duty.’
Alex Wilhelm: 59:47 I loved it. I watched four episodes. I caught up on Season 4 by your request. HBO’s ‘Industry’ Season 4. I could tell immediately why you liked this season. The whole season revolves around Tender, a tech company and app. They’re transitioning from a payment processor for porn sites and sort of sketchy kinds of services to they want to be a respectable neo-bank operating in the UK. It reminds me a lot, I think Tether was probably an inspiration for this season.
Jason Calacanis: 1:00:56 Yeah, they’re clearly listening, yeah. They’re clearly dialed in, these writers. I’d love to have the writers on at some point.
Jason Calacanis: 1:03:00 Like to the show, it turns out you’ve now got it in the startup world. They just reset the whole concept to now there’s a startup, there’s a short-selling firm, there’s this Financial Times-like journalist doing crazy things and then working with the shorts, which is like Hindenburg or other short sellers and they even name-checked like Herbalife and that short.
Alex Wilhelm: 1:04:14 I have never seen anything this crazy in terms of mixing promiscuity, deviance, drug use, and business, and getting it all kind of right in a crazy kind of way. The performances are amazing. It’s a very young cast. They’ve basically given the reins to these two young female actresses who are crushing it in this show. Highly recommend.
Jason Calacanis: 1:08:44 All right, that’s it. We had a great show today. What a great week at This Week in Startups, TWIST, firing on all cylinders. We’ll see you all on Monday, and we will certainly be doing more Open Claw.
