Twenty minutes with the CEO of ElevenLabs, now reportedly valued at $22 billion
There is still a lot of work to be done, and the quality delta you can achieve just on the model level is still significant. If we think longer term, probably three, five years from now, those differences will be smaller. What we’d love to do, and be the first ones to do, is pass the Turing test for conversational AI. You need to combine intelligence, but you also need emotional intelligence. You need to understand the emotions of the other side, to be able to slow down or speak up. That hasn’t yet been done.
What percentage of the business is enterprise now?
We are now $600 million in ARR. Fifty-five-percent plus is classic enterprise, and a big percentage of the 45% are small and medium businesses, developers, builders, creators.
Like everyone else in AI, you’re increasingly competing with your customers. Decagon trained its voice product on you and now runs queries through its own models.
The lines become more blurry. As we think about model companies, platform companies, application companies, in the past you’d have very clear splits where one starts and ends. Today that line is much more blurry. In Anthropic’s case, what was a model company is definitely a platform and increasingly a wide set of applications. I think this will continue.
Your customers can choose the “reasoning layer” from a menu of options at ElevenLabs. What are you seeing in terms of the use of frontier lab models versus open-weight?
It’s less of a binary choice. In customer experience, if you’re calling in and it’s just informational, you’re not executing any actions — you can use a lot of the open-source models because your knowledge base defines what a good experience is. But if it’s financial services, you want to be authenticated, you want information about a transaction, maybe a refund. There’s no room for error. Here, frontier models will still lead.
Some of those open-weight models are Chinese. The U.S. government is a customer. European governments are customers. What are those conversations like?
Different. In each deployment, the models and the voices we deploy will depend on the case. If we work with the Polish government or the Brazilian government, they have their own set of requirements. It can be an open-weight model, a closed-source model, their own fine-tuned model. it’s a healthcare case. You have patients booking appointments across the public health system, and 18% never show up. The deployment is agents that call and remind you. They had a set of models optimized on their knowledge, and we integrate while keeping data residency.
Should businesses disclose when someone is talking to an agent rather than a human?
I think there should be disclosure at this time. Currently, people aren’t used to it, and the common pattern is you don’t want to feel cheated on that call. But in five years, when everybody has their own agent working on their behalf, you’ll be calling in and expecting an agent. Then I think we’ll shift as a society. There are good ways of doing it — if there’s a 30-minute wait for a human, offer the customer a choice. In almost all cases, they choose the agent and then they’re surprised by how good the experience is.
What are your gross margins, given what you’re paying for models and inference?
I’m going to give a vague answer. Given we have that research element, we’re able to fine-tune and constrain models in extremely smart ways. But if we can pass on any savings to the customer, we do that. The biggest thing is still proving the value and being there with the customer. So if we can invest and prove that value, we don’t mind the margins going lower to actually benefit together as the value gets created in the next five years.
You have millions of hours of customer service calls. Do you train on them? How much of your training data is synthetic?
In certain companies, we created the models together. They wanted a specific model for their use case. Otherwise, the big part of the training hasn’t been so much the volume of data, it was annotating the data. We have thousands of people internally on a contracting basis helping us annotate not only what was said, but when people were speaking, how they said things, what emotions were used. We had to bring voice coaches in to be able to detect accents accurately.
It’s been reported you’re looking at 2028 for an IPO. Can you confirm that?
We’d love to create a company that stands the test of time. We are preparing the foundation to be able to do it in the next years. But whether we do it will depend on the time and place.
‘Years’ is very vague.
Backstage: Where do you land on whether the frontier labs should slow down?
Everybody is aligned to work together on finding a way to pace. Whether they should be public about it, and how much of the media conversation or regulation it should involve, that’s another topic. But yes, we should all take the right precautions as we deploy the technology. We don’t train the text models and the intelligence side of models, which is the core key of the debate.
Could ElevenLabs be exposed the way Hugging Face was?
We’re a step further, because we don’t deploy self-replicating or recurrent parts of the intelligence of agents. Our technology doesn’t allow you to let agents create more agents. Every customer goes through KYC. Cybersecurity risk is definitely a risk for the wider world, but we have a good set of precautions in place.
Discover more from NAIRAVOICE.COM.NG
Subscribe to get the latest posts sent to your email.