Will Everyone Become an AI Builder? Clem Delangue on Hugging Face, Agents, Local AI & Robotics
Will Everyone Become an AI Builder? Clem Delangue on Hugging Face, Agents, Local AI & Robotics
Summary
Clem Delangue, co-founder and CEO of Hugging Face, sits down with Ksenia Se to discuss the rapid democratization of AI. He predicts the number of AI builders will explode from a few hundred thousand today to tens or even hundreds of millions, fueled by coding agents like Hugging Face’s “ML-intern” that lower the barrier to entry for fine-tuning models, creating datasets, and converting formats. He argues this expansion will shift AI’s focus away from “video slop” and toward meaningful problems in biology, chemistry, medicine, and climate.
A core thread of the conversation is Clem’s passionate defense of open source AI against renewed lobbying in Washington. He argues that American AI leadership was built on open source contributions like the Transformer architecture, and that restricting open source would concentrate power among a few large companies, kill competition, and ironically increase cybersecurity risk by leaving defenders without the tools attackers can already access through leaks. He pushes back on fear-based marketing, noting that GPT-2 was once deemed “too dangerous to release” and we now laugh about it.
Looking ahead, Clem is most excited about local AI — running models on your own hardware for privacy, cost, and control — and open robotics through LeRobot and the Reachy Mini, which Hugging Face has shipped nearly 10,000 units of. He sees Hugging Face becoming “agent-native,” expecting agent users to surpass human users by year’s end. He closes with a meditation on Camus’ Sisyphus as a metaphor for founder life in AI: enjoying the task of building itself rather than obsessing over the outcome.
Highlights
”The numbers of people who are going to be able to become AI builders is going to explode”
“I think the numbers of people who are going to be able to become AI builders is going to explode, right? It’s going to go from maybe a few hundred thousands or or low millions of people who have the skills to do this kind of work, to maybe tens of millions, fifties of millions, maybe 100 million at some point.” — Clem Delangue, 2:29
Clip command
yt-dlp --download-sections "*2:29-3:30" "https://www.youtube.com/watch?v=DfJV722V1WY" --force-keyframes-at-cuts --merge-output-format mp4 -o "ai-builders-explode.mp4"
”More biology, chemistry, medicine, climate change AI — less video AI slop”
“If more people could build AI, maybe we would have a little bit less, you know, video AI slop and maybe a little bit more of like, you know, biology, chemistry, medicine, you know, climate change AI, like some things that a couple of like Silicon Valley guys don’t care so much about but that a lot of other people care care about.” — Clem Delangue, 3:44
Clip command
yt-dlp --download-sections "*3:44-5:01" "https://www.youtube.com/watch?v=DfJV722V1WY" --force-keyframes-at-cuts --merge-output-format mp4 -o "less-slop-more-medicine.mp4"
”Open source is actually a solution much more than a problem”
“Famously, I think it was GPT-2 that was like too dangerous to release, right? And now we’re laughing about it, that was like three years ago. I mean, it’s just like, it didn’t create any problem for it to be open source.” — Clem Delangue, 9:44
Clip command
yt-dlp --download-sections "*9:44-11:51" "https://www.youtube.com/watch?v=DfJV722V1WY" --force-keyframes-at-cuts --merge-output-format mp4 -o "gpt2-too-dangerous.mp4"
”Without open source, only OpenAI, Anthropic, and big tech can do AI”
“Without open source models, without open source datasets, without open source libraries, it’s impossible for anyone to do AI except OpenAI, Anthropic, and the big tech companies. And imagine a world where only a few companies can do AI, just like for example if you would have a world where only a few companies could do software, it would be quite quite scary.” — Clem Delangue, 15:35
Clip command
yt-dlp --download-sections "*15:35-17:25" "https://www.youtube.com/watch?v=DfJV722V1WY" --force-keyframes-at-cuts --merge-output-format mp4 -o "only-few-companies.mp4"
”By end of this year, we may have more agent users than human users”
“We’ve seen that also with with Hugging Face because we’re seeing a bigger and bigger part of our usage coming from agents. I wouldn’t be surprised if by the end of this year, we had more agent users than human users of Hugging Face.” — Clem Delangue, 18:18
Clip command
yt-dlp --download-sections "*18:18-19:13" "https://www.youtube.com/watch?v=DfJV722V1WY" --force-keyframes-at-cuts --merge-output-format mp4 -o "agent-users-surpass.mp4"
”You have to imagine Sisyphus happy — enjoy the task itself”
“Especially for us as parents sometimes to see like 20 year old in Silicon Valley like working 24/7 and you’re like fuck how how am I going to stay relevant in a world like that but adopting more of a mindset of just enjoying the task, enjoying the journey, the work is useful.” — Clem Delangue, 42:00
Clip command
yt-dlp --download-sections "*42:00-43:09" "https://www.youtube.com/watch?v=DfJV722V1WY" --force-keyframes-at-cuts --merge-output-format mp4 -o "sisyphus-happy.mp4"
Key Points
- Default coding agents are bad at building AI (0:54) - Clem cites Karpathy’s comment that he barely used agents for AutoResearch because it’s out of distribution. But with tweaks to harnesses and tools like the Hugging Face Hub, ML-intern is fine-tuning small models and creating datasets.
- ML-intern passed the Hugging Face researcher interview in half an hour (0:54) - The agent aced the test that human researcher candidates take.
- From hundreds of thousands to 100 million AI builders (2:29) - Eventually every software engineer should be able to fine-tune and optimize models themselves, breaking dependence on closed APIs whose vendors can deprecate models or raise prices arbitrarily.
- AI’s text-driven nature makes it more accessible than software engineering (3:44) - You don’t need to learn a programming language to contribute to AI through datasets.
- Building AI yourself changes your perception of AI (5:14) - Public perception is terrible because people are scared or hostile; hands-on building reveals AI as empowering rather than threatening.
- Reachy robot example: people don’t say they want AI robots, but they love them once they build one (6:41) - Hugging Face has sold nearly 10,000 Reachy units; users assemble them in three hours and become advocates.
- Fear-based marketing sells API access (8:10) - After Project Glasswing was announced, member companies immediately started pitching commercial agreements.
- AI is software 2.0 or 3.0, not Robocop (9:00) - It’s not a self-conscious entity; it’s a technology builders push in the right direction.
- Open source improves cybersecurity (9:44) - Open repositories are patched faster; closed-source attacks let attackers operate undetected for weeks. LLaMA’s leak demonstrated the risk asymmetry.
- GPT-2 was “too dangerous to release” three years ago (9:44) - We laugh about it now; the doom predictions never materialized.
- It’s fine for companies not to open source — just don’t lie about it (12:52) - Closed companies often invoke “safety” when the real reason is business strategy. Releasing small artifacts (a paper, dataset, or small model) builds credibility and helps hiring, as Mistral and Cohere demonstrated.
- US AI leadership came from US open source leadership (14:46) - Google open-sourcing Transformers (“Attention Is All You Need”) enabled GPT. Lobbying against open source would reverse American AI dominance within a few years.
- Open source creates competition, jobs, and growth (15:35) - Without it, only OpenAI, Anthropic, and big tech can build AI, concentrating all the captured value and destroying more jobs than it creates.
- Coding agents are the biggest change since paternity leave (18:00) - Agent users may surpass human users on Hugging Face by end of year.
- Hugging Face is going “agent-native” (19:19) - Focus on CLIs, APIs, headless interfaces, agents.md files, and token efficiency for agent consumption.
- Reachy Mini is the first agent-native robot (20:27) - Users talk to their agents to build robotics apps within hours of unboxing.
- Local AI is the biggest cybersecurity gain possible (23:13) - Free, private, fast, controllable. Llama.cpp is the most-used runtime. 99% of workloads today are proprietary API calls; ultimately ~95% should be local or specialized open source.
- Comparing open weights to APIs is apples to oranges (25:21) - APIs include tooling, harnesses, multiple models behind them. Open models trade some benchmark accuracy for cost, speed, privacy, and learning.
- Training your own models is the real differentiator (27:00) - As Cursor and Lovable make building features trivial, the moat shifts to fine-tuning and post-training on your own data — impossible with closed APIs.
- Hugging Face Buckets for robotics datasets (31:14) - New S3-like product using deduplication tech from acquired company XetHub to host large datasets cheaply.
- 15 million AI builders, 3 million public models, ~1 million datasets, with only 200 employees (33:14) - One new repository every 8 seconds.
- Open agent platforms have historically “sucked” for local AI (37:53) - The harnesses are technically open source but optimized for proprietary APIs like Claude, which is why people wrongly blame the local models.
- Camus’ Sisyphus as founder philosophy (41:02) - Find happiness in the task itself, not the outcome — especially in a field as overwhelming as AI.
Mentions
Companies
- Hugging Face (0:22) - The open AI platform Clem co-founded; 15 million builders, ~3M public models, ~1M datasets, 200 employees.
- Google (15:00) - Open-sourced the Transformer architecture in “Attention Is All You Need.”
- OpenAI (15:35) - Cited as one of few companies that could dominate AI without open source competition.
- Anthropic (15:35) - Same context as OpenAI; Claude later mentioned as well-supported by current agent harnesses.
- Eleven Labs (12:33) - Example of a company keeping its model closed as its moat.
- Mistral (12:52) - Example of a company that built a big business partially through open source releases.
- Cohere (12:52) - Same context — benefited from open source contributions.
- Apple (23:13) - “The champion of local” computing.
- NVIDIA (35:11) - Ksenia asks if Hugging Face turned down a big NVIDIA investment; Clem declines to confirm but says they could have raised 10x more.
Products & Technologies
- ML-intern (0:34) - Hugging Face’s coding agent for ML tasks; fine-tunes models, creates datasets, converts formats; passed the researcher interview test.
- Hugging Face Hub (0:54) - Central platform agents connect to for models, datasets, and tools.
- AutoResearch (0:54) - Karpathy project mentioned as example of work agents struggle to do.
- Reachy / Reachy Mini (6:41) - Hugging Face’s open robot; ~10,000 shipped; first “agent-native” robot.
- LeRobot (21:11) - Hugging Face’s open robotics library, most-used for open robotics.
- Project Glasswing (8:10) - Referenced as example of program whose member companies aggressively pitched commercial deals after announcement.
- GPT-2 (9:44) - Once labeled “too dangerous to release”; cited as a now-laughable example.
- LLaMA (11:51) - Famous Meta model that leaked, demonstrating fake safety of closed releases.
- ChatGPT / GPT (15:35) - The “T” stands for Transformer — Google’s open-sourced architecture.
- Llama.cpp (24:00) - Most-used runtime for local AI; part of the Hugging Face team now.
- Cursor (27:00) - Example of product making feature-building trivial.
- Lovable (27:00) - Same context — anyone can build apps now.
- Hugging Face Buckets (31:14) - New S3-like storage product for large AI datasets.
- XetHub (D-Z-I-T tech) (31:14) - Deduplication technology from a company Hugging Face acquired, powering Buckets.
- Transformers / “Attention Is All You Need” (15:00) - Google’s open-sourced architecture; foundation of modern LLMs.
- agents.md (19:36) - Documentation standard for making sites agent-friendly.
- Claude (37:53) - Cited as a model that current open agent harnesses are optimized for.
People
- Clem Delangue (0:22) - Co-founder and CEO of Hugging Face; the guest.
- Ksenia Se (0:22) - Publisher of Turing Post; the host.
- Andrej Karpathy (0:54) - Cited for saying he barely used agents to build AutoResearch because it’s out of distribution.
- Steve Yegge (3:30) - Quoted by Ksenia as predicting non-technical people will enter coding.
- Nathan Lambert (25:09) - Recently told Ksenia he doesn’t believe open models will catch up with closed models.
- Albert Camus (41:02) - Author of The Myth of Sisyphus; Clem’s favorite book and founder metaphor.
Surprising Quotes
“Today the team uh got it to pass the interview test that they had for for researchers in half hour it aced the test.” — Clem Delangue, 0:54
“If just a few companies do AI, they’re going to capture all the value and they’re not going to create enough jobs to compensate all the jobs that are destroyed.” — Clem Delangue, 15:35
“It’s not kind of like Robocop, it’s not kind of like a self-conscious entity that governs itself, it’s really kind of like a technology that we built and that we are going to push in the right direction progressively.” — Clem Delangue, 9:00
“When you run a model locally, it’s almost free, right? You already have the hardware most of the time. It’s private. When you’re talking about cyber security and then local is kind of like the biggest cyber security gain that you can ever have because you don’t send your data anywhere.” — Clem Delangue, 23:13
“We’ve 15 million AI builders using the platform now, one new repository created every, every 8 seconds on the platform. Like almost 3 million models, public models on the platform, almost a million datasets.” — Clem Delangue, 33:14
“Especially for us as parents sometimes to see like 20 year old in Silicon Valley like working 24/7 and you’re like fuck how how am I going to stay relevant in a world like that but adopting more of a mindset of just enjoying the task, enjoying the journey, the work is useful.” — Clem Delangue, 42:00
Transcript
Clem Delangue: 0:00 The numbers of people who are going to be able to become AI builders is going to explode. Life as if you would have a world where only a few companies could do software? It’d be quite, quite scary. You want open source to create more jobs. But how, how am I going to stay relevant?
Ksenia Se: 0:22 Thank you Clem for agreeing for this, I’m a big fan of Hugging Face and what you guys doing for open source community. It’s been amazing to know you for many years but meet you for the first time in person.
Clem Delangue: 0:31 Yes. Thanks for having me.
Ksenia Se: 0:34 Um, let’s start with, with your recent post about ML-intern, how you’re playing with it on Hugging Face. Um, what is the most surprising and funny lesson that you learned how agents work with real machine learning tasks now?
Clem Delangue: 0:54 Yeah, so, I mean what’s interesting is that the, I think the default coding agents are pretty bad at building AI. And you saw that I think it’s when uh Andrej Karpathy released I think AutoResearch, or I don’t remember if it was AutoResearch or something before, he said ‘oh I, I barely used any agents to build it just because it’s either like out of distribution or like is, it just doesn’t work yet really to build AI’. But with a couple of like tweaks to the harnesses, to the models, connection to tools like the Hugging Face Hub, you can actually uh make a lot of progress. Uh so we were surprised that uh now uh ML-intern is managing to uh to fine-tune some small models, to create some datasets, to convert models uh to different formats. Today the team uh got it to pass the interview test that they had for for researchers in half hour it aced the test. So we’ve been excited about it. If agents can lower the barrier to entry to build AI, uh it’s going to be very valuable for the world because it’s going to enable more people to do open source models, to do open data sets, to maybe play with local models, which historically have been a little bit hard to do, but I think now it’s getting easier.
Ksenia Se: 2:24 How do you see this development in the coming months? What where the exploration is?
Clem Delangue: 2:29 I think the numbers of people who are going to be able to become AI builders is going to explode, right? It’s going to go from maybe a few hundred thousands or or low millions of people who have the skills to do this kind of work, to maybe tens of millions, fifties of millions, maybe 100 million at some point. Maybe at some point every software engineer will be able to optimize models themselves, train models themselves, fine-tune models themselves. Uh and that obviously would, would be amazing, right? Because it obviously would be… mean that they’re not going to only rely on closed source APIs and on kind of like third-party vendors who can kind of like dictate their conditions in a way, right? Increase their prices whenever they want, deprecate models whenever they want, change them behind the scene that you’re not even sure why, why the quality is going down on your workloads and so to give back some some control to the to the builders which is which is nice.
Ksenia Se: 3:30 a couple of months ago I was doing a little interview with Steve Yegge and he said oh and non-technical people will definitely come to this coding world. How do you feel about that? Are we ready?
Clem Delangue: 3:44 The beauty of AI is that a lot of is of it is driven by data sets and text in general, right? So compared to software engineering where you had to kind of like learn programming language, I think AI has the potential to have much wider base of of users, of people who can contribute to it. So I hope I hope it happens. I think it would be good too because the more diversity of builders that you have, I think the wider the perspectives and I think for the field what is good is that it’s going to drive it towards more actual challenges and things that are important for people, right? Like I feel like if if more people could build AI, maybe we would have a little bit less, you know, video AI slop and maybe a little bit more of like, you know, biology, chemistry, medicine, you know, climate change AI, like some things that a couple of like Silicon Valley guys don’t care so much about but that a lot of other people care care about. It brings more perspective and so hopefully more more problem solved.
Ksenia Se: 5:01 You think that more creation with AI will essentially get to some like quality point when people stop creating slop? Because there’s a lot of slop right now. And when they stop creating slop and start actually like solving problems?
Clem Delangue: 5:14 I think there’s a lot of other things than slop, slop to to build and the more you enable empowered people to to build, the more they’re going to build other stuff than than slop. And also I mean I think like empowering more people to become AI builders will change the perspective of the public on AI, right? Like obviously right now we have terrible perception of AI by the public, you know? Like if you look at the studies it’s crazy. Either people are are crazy scared or they hate AI or they don’t want to hear about it whereas I think if you if you help them understand how to build these systems, if you let them build this this systems, they they’re going to kind of like see that it’s actually something that’s empowering, that they can use to do whatever they want, that they can use to solve new problems. And so I think this is kind of like one of the ways to, uh, to change public perception of AI versus if it stays kind of like just in the hands of a few companies, of a few builders. I mean, they can do marketing, they’re going to do marketing, but still I don’t think, I don’t think they’ll manage to convince a lot of people that AI is good.
Ksenia Se: 6:22 So to your point of opening the world and changing the perception of AI, who is the main player who can do that? Because now we are five main companies, we have government, we have, uh, attempts like Hugging Face to open the space. Who can actually change the opinion?
Clem Delangue: 6:41 A bit of everyone. I think it’s, uh, every part of the ecosystem has, has a role to play, right? I think obviously, uh, policy makers have, have a big role to play, companies have a big role to play, research, academia has a big role to play, and each kind of like at their level. Like, uh, good practical example is like we have, we have the Reachy robot here. I think if you, if you ask people in an abstract way, are you excited by AI robots? I don’t think a lot of people would say they, they are. A little bit in our bubble some people are, but most, most people are not. But what we’re seeing is when we, when we ship one of these, like we, we sold, uh, almost, almost 10,000 of them. We ship them to people who maybe wouldn’t be particularly excited about AI robots in general, but maybe bought this because it’s cute, because they think they can play with it with their kids. Then they assemble it, right? Like they take three hours to assemble it, build it themselves. They start to play with it, they build apps and automatically they start, yeah, they start to love AI robots and they’re like, AI robots are the best things. Uh, so that, that’s kind of like an example of, of just kind of like mechanics of like enabling people to take part in the process, build it themselves and how it changes their, their perception. Obviously, the biggest thing is that there’s a lot of fear-based marketing in, in AI right now, which serves some, some purpose.
Ksenia Se: 8:08 What’s the purpose of that?
Clem Delangue: 8:10 Oh, um, I mean, it sells obviously, um, you know, like just, I think just a few days after Project Glasswing was, was announced, we started, I’m sure it’s the same thing for a lot of people, we started to receive emails from some of the companies part of the program, telling them that they were part of the program and trying to sell us some commercial agreements. And so obviously, obviously it serves these kinds of, of purpose. I also do believe that a lot of people kind of like marketing like that do themselves believe in kind of like the fact that, you know, it’s, it’s important to restrict access and things like that. But I think it’s, uh, it’s a mistake and it’s, uh, it’s misleading. Again, what you want to enable is kind of like giving access to people so that they can build. Yo, realizing that it’s a tool, right? It’s just a new, it’s software 2.0 or software 3.0. It’s not kind of like Robocop, it’s not kind of like a self-conscious entity that governs itself, it’s really kind of like a technology that we built and that we are going to push in the right direction progressively. And so there’s no really need in my opinion to kind of like play too much on this fear based marketing approach, but—
Ksenia Se: 9:29 Let’s try to counter some of those views that try to make open source specifically kind of a tool to create, I don’t know, bio-weapons, deep fakes. What can you tell to these people? Why it’s true, or it’s not true?
Clem Delangue: 9:44 Cybersecurity is a good example, right? Because it’s being weaponized. It’s kind of fun, because they struggle that it can create a weapon, but then they weaponized the cybersecurity against other domains. But as kind of like an example of how you shouldn’t release some of these models, but if you look at really how cybersecurity works, it’s always by kind of like empowering more defenders than attackers, making it more expensive to attack than to defend and creating the resiliency in systems so that when you have an attack and when you have something that succeeds, that the system can be patched really quickly, resolved, and that the problem that it causes are not systemic, right? And in all these things, when you take them individually, open source is actually a solution much more than a problem. For example, we all know that open source repositories are patched much, much faster than any proprietary system. Once you have an attack on a proprietary system behind closed doors, like the attackers basically take advantage of getting all the user data, all the important information for weeks, weeks, weeks without it being patched. And when it’s patched, it’s too late. And so by keeping it closed source to a small number of people, then you increase the risk, because you increase kind of like the asymmetry of power and capabilities, right? You have some people who can get big capabilities, but the defenders don’t have this kind of capabilities. So that’s when you create kind of like the bigger risk, versus when you open them up, you keep kind of like the balance between the two. And so usually the defenders have the way to counter the attackers. We’ve had the same kind of like thinking, many steps, obviously in the past in AI, right? Famously, I think it was GPT-2 that was like too dangerous to release, right? And now we’re laughing about it, that was like three years ago. I mean, it’s just like, it didn’t create any problem for it to be open source.
Ksenia Se: 11:51 And LLaMA has leaked.
Clem Delangue: 11:53 Yeah. Yeah, it’s like, but it’s good example of the risks, if like a model is a bit more powerful and get leaked, or for an entity— there is doesn’t have the right intentions have access to this kind of things and not the rest of the world, that’s actually when you when you create a more risks. And I think the, you know, APIs or the kind of like small released, they give a fake impression of of control and safety, whereas it’s not. Like the leaks is good example of that. So if you look systematically at how cybersecurity works, open source is actually a solution to most of the problems rather than the issue.
Ksenia Se: 12:33 You mentioned API and for many companies, like for example, Eleven Labs, they don’t open source because this is their moat, they’ve been working on the closed model for a long time and, you know, with the big players, they just can’t. So when you advocate for open source so passionately, how is it considered with the business part?
Clem Delangue: 12:52 Yeah. So first, it’s totally fine for businesses, for companies not to do open source. There’s no nothing wrong about that, right? Sometimes what I get frustrated is when when companies don’t say that and instead say like, oh, I’m not open sourcing because of safety, right? Because I don’t think deep down that’s the real reason, the real reason is more as you said, because it’s not in their business interests to open source. And that’s totally fine. Usually what we explain to some of these companies is that, you know, if they open source some small parts of what they do, you know, maybe they can publish a research paper about some of the things that they do, maybe they can release a data set that has been kind of like partly public, maybe they can release like a small model and keep the big model proprietary. And if they do that, first it’s good for the world, right, for for the field because they they contribute, and we’ve seen many examples where it’s good for the companies themselves, right, because it helps them hire better people who want to contribute like that. It makes them be seen as more credible. If they release kind of like small artifacts, I think it builds up their credibility, it builds up their their visibility. Obviously with lots of examples of of companies like like Mistral, Cohere, like these kinds of companies who benefited tremendously from releasing things in in open source and managed to build like a big business, big businesses thanks thanks to it. So that’s kind of like usually what we explain. But again, it’s totally fine for business not to open source if they feel it’s not aligned with their strategy, if that’s not what they what they want to do.
Ksenia Se: 14:31 You’ve recently been flagging lobbying in DC against open source, and again, you’re very passionate about it. What’s your stance here? What do you want to tell them? What do you want us to do? Why is it important?
Clem Delangue: 14:46 Well first, it’s not the first time. We’ve had kind of like the same thing happening two three years ago for various reasons, but it looks like it’s it’s coming back. I think it would be a mistake for the US, but frankly for any country to try to slow down open source because it’s it’s really kind of like a chance for them because open source is is the foundation to to all technology and a country that leads open source is a country that can lead AI in in general, right? I think all the progress in AI that you’re seeing today and all the American leadership that you’re seeing, in my opinion, comes from the open source leadership of the US, right? Like if you remember of course like Google famously open sourced transformers
Ksenia Se: 15:33 Attention Is All You Need.
Clem Delangue: 15:35 which is then used by ChatGPT, right? The T in GPT. And that and that’s kind of like just one one example of all the emulation, the collaboration that happened in open source in the US that led to the leadership today. So if tomorrow the US slows down open source, automatically a few months, a few years later, they’re going to lose their AI leadership in general, which is which is not something something that we want. And then second, I mean, if you if you slow down open source, you increase concentration of of power, of capabilities, of revenue, and you run the risk of AI being dominated by one or two or three companies because if you remove open source, you prevent anyone from competing with the big guys, right? Without open source models, without open source datasets, without open source libraries, it’s impossible for anyone to do AI except OpenAI, Anthropic, and the big tech companies. And imagine a world where only a few companies can do AI, just like for example if you would have a world where only a few companies could do software, it would be quite quite scary. You want open source to create competition, create emulation, to create more jobs, you know, like if just a few companies do AI, they’re going to capture all the value and they’re not going to create enough jobs to compensate all the jobs that are destroyed. You want AI to to create more open source to create more growth, right? Like you want you want an ecosystem of small companies, small like businesses, medium businesses, big companies, everyone to be able to build AI and create value for them. Otherwise, you’re going to end up, yeah, with with a world where only a few companies are capturing all the value and that not that’s not going to create enough growth.
Ksenia Se: 17:25 You’ve recently had like a three months paternity leave, and three months is a long time in AI. Yeah. So how do you see the world now? What has changed?
Clem Delangue: 17:37 Yeah, so I came back a little bit earlier than than my planned three months because I was I was too excited to to come back, to be honest. And also I was lucky that my my wife did such an amazing job at kind of like taking care of the baby that I was lucky to to feel confident to be able to go go back to go back to work. Obviously, even even a week Just like a long time in AI these days, the biggest change has been kind of like the the advent, the total domination total kind of like a mind-blowing adoption of coding agents, right? That completely changed the way most most technology is built. We’ve seen that also with with Hugging Face because we’re seeing a bigger and bigger part of our usage coming from agents. I wouldn’t be surprised if by the end of this year, we had more agent users than human users of Hugging Face. Because we’re seeing, yeah, a lot of people obviously using agents who are pulling models from Hugging Face, pulling data sets, contributing to things. So that’s - that’s super exciting, right? It’s kind of like a crazy kind of like multiplier effect on on builders, multiplying force for builders. So it’s - it’s really cool to see.
Ksenia Se: 18:59 That’s very interesting because we were talking about the builders coming, new builders, maybe not AI engineers or machine engineers, people who tinker with all that stuff. But at the same moment, agents are coming to the same platform. So if we talk about Hugging Face, how structurally or technically do you need to change the platform?
Clem Delangue: 19:19 Yeah. Yeah, you have to adapt it. Agents are - are your new customers or your new users. So you put much more - we already did, but you put much more focus on your CLIs, on your APIs, on everything headless that agents can adopt easily. You make sure obviously all your documentation, your agents.md files or everything works seamlessly for agents. You make it - something that people sometimes underestimate, but you also make it token efficient for agents to use your platform and your tools, right? Because you don’t want - you don’t want agents to burn through too many tokens, especially with the prices of tokens these days. To use your platform, yeah, right, you want to make it very efficient in the way you build your APIs and your abstractions for agents. One good example of that is also with - with the ReachMini’s because when we started working on it last year, it was really, really hard for people to build apps for the ReachMini’s. And now we’re shipping like a new batch this week, next week. You’re going to - going to receive one. And now people can just - just talk to their agents to build any robotics app they want. So like this ReachMini is going to be like the first - first robot that is fully agent-native where people when they receive it, they can right away start building apps that - that they’re excited about with their agents in - in a few hours. So I’m super excited to see with the new batch what people are going to build with their agents.
Ksenia Se: 21:00 Hey, one of the first platforms making it possible. How does Hugging Face robotics ecosystem - does it involve your business strategy? What part of it?
Clem Delangue: 21:11 Yeah, I mean, robotics is is very important for us. It’s an extension of our platform. The community is very vibrant with LeRobot. We’re seeing like a really great excitement from the community that has become kind of like the most used library for open robotics. And we see it as kind of like a way to continue to enable empower AI builders, right? Like in the same way an AI builder should be able to train their own models, optimize their own models, they should be able to build their own robot, optimize their own robot. And frankly, it’s quite fun. And like you’ll see when you receive yours, just to assemble it yourself, to build it, to fix it when there’s a problem and to build new apps, it’s quite quite fun. We don’t always think - that’s probably one of our flaws or one of our qualities, it depends which perspective you take. For our investors, it’s probably a flaw, for other people, it’s probably a quality. But we usually not so good at thinking too much about how we monetize and how we generate revenue from from things. In general, I think we’re starting to see the right system where we have some sort of a freemium model, right? Where a lot of what we do is free, is open source. And then a small percentage of what we do, especially when it comes to bigger enterprise, to people using us a lot or people using it for things that are just costly, right? When they use us for storage, when they use us for tokens, when they buy a robot, then people pay for that. And hopefully, kind of like this part of the product of the platform funds like the free part and we can create the right kind of like flywheel where we continue to grow, to be profitable.
Ksenia Se: 23:02 Cause maybe, you know, when everybody will switch to their local models on their laptops, there is no business for you here. So you will just switch to be hardware and sell robots.
Clem Delangue: 23:13 Yeah, yeah, yeah. No, but that would be amazing. Local, I’m so excited about it. Because obviously when you run a model locally, it’s almost free, right? You already have the hardware most of the time. It’s private. When you’re talking about cyber security and then local is kind of like the biggest cyber security gain that you can ever have because you don’t send your data anywhere, it stays on your hardware, right? There’s a reason why Apple is kind of like the champion of local. And you know, it’s fast, it’s controllable, it’s you can kind of like really play and update the weights and transfer the weights. So I’m super excited about it. So we see the volume of downloads of local models from Hugging Face to really like, that really exploded. These days, part of the team we have Llama.cpp, which is like the most used runtime for local AI these days, makes it super easy to once you’ve downloaded a model to basically run it on any of your of your hardware. And I hope, I hope it will continue to grow. Like ultimately for me, we are like in the early stage of AI, right? And in this early stage of AI, for some reason, 99% of the workloads are APIs calls to massive proprietary models. Probably because it’s easier, because, you know, people feel more comfortable maybe to start with that. But ultimately, I think a very big part of your workloads are going to be either new open source models, are going to be smaller, more specialized models, are going to be local models, and maybe you’re going to use kind of like a big proprietary, costly API just for 5% of your tokens, right? And then 95% of your token is either going to be like more specialized open source models or local, local models, which, which would be, which would be nice.
Ksenia Se: 25:09 I had a conversation with Nathan Lambert recently, and he said that he does not believe that open models will catch up with closed models ever soon or maybe even ever. But how do you think about it?
Clem Delangue: 25:21 Comparing kind of like open weights with APIs is a little bit comparing apples to oranges because it’s not really the same systems. For example, behind an API, you have a bunch of tooling, you have a bunch of harnesses, you have a bunch of systems, sometimes you have several models behind an API, and so comparing it to a raw model is, is kind of like unfair. It’s like saying, “Okay, my, my engine is never going to be better than a car.” It’s not, it’s not the same thing. That’s first. And they don’t need to have the exact same accuracy because they provide different kind of like value. And I mean, a local model is free, is, is private, and so even if you don’t have 100% of the same accuracy on 100% of the tasks, it’s better for you to use that for some tasks because you save money, because you have more privacy. I mean, it’s fun to always compare the accuracy of the two. I think the gap has been shrinking. I think the community has been amazing at, at pushing the frontier in, in open source. But at the end of the day, I don’t think that’s what matters most. I think ultimately, especially with people looking more at the cost, people being worried about compute constraints, I think that will be kind of like positive forces for open source models no matter how close they are to kind of like accuracy in some benchmark. And obviously, benchmarks are benchmarks. Like, it’s not because one model is, is better at one big generalist chat benchmark that on your specific task, that it’s going to perform as well. Right. I think progressively you can let go of this thinking of, you know, is it in absolute terms better or worse, and look at specific things and basically use whatever you want to use, whatever you feel like is gonna give you the best trade-off between not only accuracy but cost, speed, privacy, control. One thing that people don’t talk a lot about is that open source gives you learning experience. It builds up your skills at training these models, which in the long run in my opinion is key for companies. As kind of like building features and products is starting to be trivial, right, with Cursor, with Loveable, with all these products, basically anyone can build websites and apps and features. Progressively I think what is gonna help you differentiate and be successful as a company is gonna move to the frontier, right? So maybe it’s gonna be your ability to train models yourself, optimize models yourself, fine-tune, post-train on your own datasets, and obviously that you can only do with open source. You can’t really do it with an API. So that’s another aspect that I think people should care more about.
Ksenia Se: 28:13 I always talk about AI literacy becoming as important as like writing, reading, but what you just said means that we not only need to write, read, maybe code, but also like train models.
Clem Delangue: 28:28 Yeah.
Ksenia Se: 28:29 What else do you think people don’t understand or what topics are we not talking enough about in this AI world?
Clem Delangue: 28:34 Yeah, what’s training models, optimizing models, building your data sets, creating kind of like multi-modal systems, running them locally, all this is kind of like the new, it’s the equivalent of building software like a few years ago. I think it’s the kind of skills that you’ll increasingly need to differentiate yourself. And robotics too, building robots.
Ksenia Se: 28:58 Maybe Hugging Face need to go to schools with Reachy Mini and training models.
Clem Delangue: 29:08 Yeah, with a bunch of professors, like Reachy Mini to kind of like help their students learn a little bit about robotics. Yeah, it’s good. I mean everything we do, we like to…
Ksenia Se: 29:22 But you’ve been playing with robotics probably more than many people. What is the equivalent of robotic slop? What is the equivalent?
Clem Delangue: 29:32 Of course, I mean humanoid robots. There’s a bunch of like fake AI marketing videos. Not many humanoid robots have actually been shipped, right, or have turned really useful yet. So there’s I think there’s a lot of slop there.
Ksenia Se: 29:44 From the marketing side.
Clem Delangue: 29:45 From the marketing side, yeah. Videos like be careful every time you see kind of like a video with a robot, you know, obviously use critical thinking because a lot of them are just marketing. …nowhere close to the real capabilities of these robots, or at least the kind of like real-life behavior of these robots. Try to see, I mean, I think this uh reachy is probably kind of like the the robot that is actually shipped the most uh these days, um and it’s pretty close to kind of like the frontier of what you can see for this level of cost, like you won’t have kind of like a more advanced robot for this kind of price, or at least I haven’t I haven’t seen seen that.
Ksenia Se: 30:36 It’s not a part of physical AI yet. It’s kind of this step from the chatbot towards robotics, I would say, right?
Clem Delangue: 30:43 Yeah, it’s kind of like a intermediary more kind of like today’s steps, like instead of thinking okay, humanoid robots who are kind gonna be more kind of like the norm maybe in three, four, five, six years, we still have a lot of things to I think solve before having like reliable, useful humanoid robots, to be honest.
Ksenia Se: 31:03 From the Hugging Face platform, what is missing? Data sets, what the bottlenecks you think your audience can solve for robotics?
Clem Delangue: 31:14 Yeah, data sets are missing, large data sets, especially because uh there’s sometimes costly to uh to gather and to host. So we’ve worked actually on a product recently to try to help with that, it’s called Hugging Face Buckets. It’s like a way to uh to store large data sets on the hub, some sort of kind of like S3 a little bit but more designed for AI so it’s using some of the technology we’ve been building for a few years called the D-Z-I-T, a company that we acquired, that kind of like deduplicates data sets and makes them kind of like much easier and cheaper to host. So hopefully that’s gonna help a little bit on the robotics data sets. Frankly, we need more builders, too.
Ksenia Se: 31:54 Interesting.
Clem Delangue: 31:55 Right? Like uh more people focusing on the topics because historically it’s been seen as a difficult topic to break into, right? And so a lot of builders are a little bit worried of getting into robotics but I think it has changed and so now more people, more builders that are like software engineers, maybe have been focusing more on software can play with robotics via Reachy and and kind of like start their learning curve and after a few months I think they can do really cool stuff especially with agents. So uh more builders, more datasets, more openness to fight the work in like silos, right? Like right now a lot of companies are working in silos without really sharing too much of their research. With more openness, more collaborativeness, I think you’ll accelerate the field of robotics uh similarly to what you’ve seen for LLMs.
Ksenia Se: 32:46 Mm-hmm. How do you feel about Hugging Face becoming a big part of the whole ecosystem and the establishment, if you will?
Clem Delangue: 32:56 Ah, establishment.
Ksenia Se: 32:58 Aren’t you afraid to be a company… It’s sort of a company that you fight against.
Clem Delangue: 33:03 Oh, it’s—it’s right. I mean, we’re still tiny compared to many other people. You know, very, very much community-driven, so I think that keeps us kind of like grounded in stuff that people need. We’ve 15 million AI builders using the platform now, one new repository created every, every 8 seconds on the platform. Like almost 3 million models, public models on the platform, almost a million datasets. But this field kind of like changes so fast that you know, I think it creates good forcing factors for us to keep keep kind of like changing, keep innovating. We’re just a 200 team member company. So we’re relatively—200.
Ksenia Se: 33:43 Hm.
Clem Delangue: 33:44 So we’re relatively small compared to our peers. We, I think we have a good setup to keep kind of like pushing, pushing ourselves to do new cool stuff, like the robotics stuff, the more storage infrastructure stuff that is useful. Keep pushing on the agent side to help everyone become more like AI builders and for agents to become AI builders themselves with our user users. So quite excited about it. My paternity leave was amazing also for that, for like taking a break and then coming back. Almost—it’s a bit of a cliché, but you know, in founder mode and with all the energy and all kind of like the freshness that sometimes founders can progressively lose. I think it was a good, good way for me to get that back.
Ksenia Se: 34:31 What was your biggest revelation?
Clem Delangue: 34:34 Obviously, I mean, you know, having five kids, having kids like gives you a bit more perspective on things. To kind of like not take things too, too seriously, not—sometimes not overthink things, but focus more on kind of like getting, getting things done and getting progress done. Focus on what’s, what’s important because it creates also more time pressure to work on really what’s most impactful so that you have time for your family and for your kids too. And these were the two biggest things. I’m sure progressively you uncover more, more insights progressively. It’s been maybe a bit too fresh to have too many, too many insights.
Ksenia Se: 35:11 Is it true that you recently turned down like a big investment from NVIDIA?
Clem Delangue: 35:17 I—we usually don’t really comment on, on kind of like these things because it’s more kind of like private conversations that, that happens. We’ve been lucky to sit in a very, very strategic position obviously in the community to always have kind of like good support from, from investors. I mean, probably could have raised like 10 times more money than, than what we’ve raised so far. And, you know, we’re, we’re lucky and grateful to be, to be in that position. At the end of the day, for something like us, I’m not really sure fundraising is, is what matters the most because it’s more kind of like a community platform that, that we built for, for the long run.
Ksenia Se: 35:52 For your freedom, probably.
Clem Delangue: 35:53 Yeah. Yeah. I mean, it’s always, always trade-offs, right, when, when you take actions, when you take strategic decisions for, for companies. Yeah, it’s true that sometimes you trade, you know, kind of like more money in the bank, bigger teams, or bigger kind of like spending for for more constraints, right? Like constraints to to return more money to investors, constraints to to go in certain directions. So it’s always, always trade-offs that you make as as founders.
Ksenia Se: 36:19 Concluding our conversation, couple more questions. What is that excites you the most in the upcoming couple of years? And I want to make a little note because I think the previous couple of years were kind of rushing into this whole bubble and still trying and researching and developing and this year feels, at least for me, more like moving into something more concrete. I don’t know how you feel about it and what excites you the most?
Clem Delangue: 36:47 Yeah, the field is definitely maturing, right? As as I mentioned, I think we go from like a place where everyone was just using the largest model proprietary from an API, to I think now people being more thoughtful about it and thinking okay, maybe for some things we can use these APIs, for others we can use more specialized open source models, and for others we can we can use local models that are that are free. And I’m excited for this, I think it’s the sign that the field is is maturing, it’s gonna be good because it’s gonna make it more sustainable, right? What you see with a lot of companies is that it’s just not sustainable to keep increasing kind of like their token spend on on large models for everything. So that’s probably the thing I’m most excited about with a part of these things, probably local, being local AI being the thing that that excites me the most.
Ksenia Se: 37:50 And to put people out of business?
Clem Delangue: 37:53 Oh no, I mean it’s the more people will do local AI, the more they’ll solve problem and see the value of AI, and then the more they’ll use everything else. So yeah, I think I think local AI is great. Things like llama.cpp is like, it’s amazing when you see how how people are using it, what they get, the kind of like tokens, token speed that they get on their laptop now. A lot of the open agent platforms have been kind of like sucking for local AI so far, which is kind of a paradox because they’re supposed to be open coding agents, open coding agent platforms, but it’s almost kind of like the the harness and the surroundings were open source but optimized for proprietary APIs, right? And so they were working really, really well with, you know, Claude, with kind of like a bunch of other models, and kind of like now working well at all for other kind of models. And people were sometimes assuming, ‘Oh it’s a models, the models suck.’ But a lot of the time it was more like the harnesses and all the tricks and systems that you have around the model. You can’t really expect to have something like… Can be like Pi, like Open-LLaMA and just switch from one model to another and for it to work instantly because there are so many things around that are kind of like making or breaking the performance and making or breaking the accuracy. But I think we’ve been working with them and they’ve been amazing at working with the community to progressively change that. And so I’m excited to see in the next few months if you have the combination of this plus the open source models plus local models, what people are going to be able to build with that.
Ksenia Se: 39:33 Yeah, it’s very exciting. What concerns you the most?
Clem Delangue: 39:39 I mean I think what we talked about in terms of kind of like renewed lobbying again against open source in the US I think is a bit concerning because it’s going totally towards kind of the wrong direction in my opinion and it’s a bit like destructive for the field. I feel like builders should focus on kind of like builders and there’s already enough challenges to compete against some of the biggest technology companies in the world that you don’t want in addition to that worry about, okay, is there going to be like regulation that is kind of like preventing me from running just kind of like a model locally? The same way like you’d think okay, if I want to write software and run software locally, I’m allowed to do that. Why would it be different for AI? That’s probably kind of like, not something that worries me because I think ultimately people will understand how important open source is and how good open source is for the world, but more bothers me because I feel like we have better things to do than to fight these kinds of things. And unfortunately we have to fight these kinds of things.
Ksenia Se: 40:54 And my last question is always about a book. What book influenced you a lot that you would love to share and it can be from your forming years or from recent years?
Clem Delangue: 41:02 I have one illustration from one of my favorite books here that I can take, from the Sisyphus Myth. It’s like one of my favorite books from Camus, which is basically this idea that you have Sisyphus that has been doomed to push this rock up a mountain and each time it reaches the top, then the ball goes down, the rock goes down and he has to start again. And the conclusion of the book, the last sentence is, you have to imagine Sisyphus happy, despite kind of like this curse, because the whole philosophy behind it is that despite the kind of like meaninglessness of the task, you know, Sisyphus is finding more happiness in the task itself because, you know, maybe the surrounding is nice, because there are much more meaningful… Two ways to look at the task and just kind of like reaching the task, reaching the top or reaching successfully the end of a task and it’s been… yeah, it’s been a good metaphor for me as a founder I think to think about like how to enjoy kind of like the the task of building itself, you know, not so much the the outcome or like where where you want to be ultimately but more kind of like enjoying the task and the process. I think you need it even more in AI right now because sometimes so many things are happening that I hear a lot of people who can feel like a bit nervous or worried or or stressed or kind of like overwhelmed, right? And they’re like so many things are happening, how how can I keep up, how can I follow, how can I compete in a way, right? Especially for us as parents sometimes to see like 20 year old in Silicon Valley like working 24/7 and you’re like fuck how how am I going to stay relevant in a world like that but adopting more of a mindset of just enjoying the task, enjoying the journey, the work is useful.
Ksenia Se: 43:09 And having fun seems like a big part.
Clem Delangue: 43:13 Yeah, having fun and doing it.
Ksenia Se: 43:15 Thank you so much. That was wonderful.
Clem Delangue: 43:18 Thanks for having me.
