Four years old. $22 billion valuation.
The voice behind Klarna’s phone support, Deutsche Telekom’s customer service, and Poland’s healthcare reminders.
And when asked about gross margins, CEO Mati Staniszewski just smiles and changes the subject.
In an interview at the Nrth conference in Toronto (formerly Elevate), Staniszewski confirmed that ElevenLabs is now pacing at $600 million in annual recurring revenue.
Enterprise customers make up more than 55% of that. The rest comes from small businesses, developers, and creators who use the platform for audiobooks, dubbing, and music.
That’s up from $500 million reported in May. Twenty percent growth in four months. Investors reportedly value the company at $22 billion, making it one of the most valuable AI startups in Europe or anywhere else.
But the interview’s most revealing moments weren’t about the numbers.
They were about the questions Staniszewski dodged, the ones he answered honestly, and the way he sees his company’s role in an industry that just spent the summer watching AI models escape their labs.
The margins question came up directly. Staniszewski wouldn’t give a number. What he did say was telling: he doesn’t mind margins getting squeezed if it means expanding market share.
ElevenLabs isn’t alone in voice AI anymore, and some of its competitors used to be its customers.
The Customer-Turned-Competitor Problem
Decagon, a conversational AI platform, trained its voice product on ElevenLabs’ technology.
Now it runs queries through its own models and competes directly with the company that enabled it.
It’s the classic platform risk: build a tool good enough that your customers don’t need you anymore.
Staniszewski was philosophical about it. “The lines become more blurry,” he said.
“In Anthropic’s case, what was a model company is definitely a platform and increasingly a wide set of applications. I think this will continue.”
He’s not wrong. But it’s easier to be philosophical when you’re growing 20% per quarter.
Whether that composure holds if Decagon or another customer-competitor starts taking enterprise accounts is another question.
Should Companies Tell You When You’re Talking to a Bot?
Yes, according to Staniszewski.
And his reasoning was more nuanced than you’d expect from someone whose business depends on making AI voices sound indistinguishable from humans.
“Currently, people aren’t used to it, and the common pattern is you don’t want to feel cheated on that call,” he said. “But in five years, when everybody has their own agent working on their behalf, you’ll be calling in and expecting an agent.”
His suggested approach: tell the customer there’s a 30-minute wait for a human, offer the AI as an alternative, and let them choose.
“In almost all cases, they choose the agent and then they’re surprised by how good the experience is.”
That’s a practical middle ground. Transparency now, normalization later. It also sidesteps the privacy lawsuits that have hit competitors.
Otter.ai is facing a lawsuit over meeting transcription practices. Granola has faced similar allegations. Companies that skip disclosure are learning that courts take a dim view of it.
Frontier vs. Open-Weight: It Depends on the Stakes
Staniszewski offered a clean framework for when customers should use expensive frontier models versus cheaper open-weight alternatives.
Informational calls where the customer just needs a basic answer? Open-weight models work fine. “Your knowledge base defines what a good experience is.”
Financial services where the customer needs authentication, transaction details, or a refund? “There’s no room for error. Here, frontier models will still lead.”
When government customers are involved, things get more complex.
Different deployments get different models depending on the country’s requirements.
Poland’s public healthcare system, for example, uses agents that call patients to remind them about appointments (18% were no-shows before the deployment).
That system runs on models fine-tuned on local knowledge with data residency requirements met.
Some of those open-weight models are Chinese. Staniszewski acknowledged the geopolitical sensitivity without offering specifics on how ElevenLabs navigates it for government clients.
The IPO Question
Reports have placed ElevenLabs’ IPO target around 2028.
Staniszewski wouldn’t confirm. “We’d love to create a company that stands the test of time. We are preparing the foundation to be able to do it in the next years.”
The interviewer pressed: “‘Years’ is very vague.”
Staniszewski laughed.
At $600 million ARR and a $22 billion valuation, the math is approaching public-market territory. But Staniszewski seems to want optionality rather than a deadline. ElevenLabs was valued at $11 billion just seven months ago. Doubling in half a year gives you leverage to wait.
Could ElevenLabs Get Hacked Like Hugging Face?
The question came backstage, and Staniszewski had a direct answer.
ElevenLabs doesn’t train the reasoning models that caused the Hugging Face incident and Anthropic’s accidental breaches.
“Our technology doesn’t allow you to let agents create more agents,” he said. Every customer goes through KYC verification.
On the broader question of whether frontier labs should slow down, Staniszewski aligned with the emerging consensus: “Everybody is aligned to work together on finding a way to pace.”
But he also drew a clear boundary. ElevenLabs makes voices, not intelligence. The pacing debate is about the labs building the brains.
ElevenLabs is building the mouth. Different risks, different conversation.
The Turing Test Is Still Years Away
Here’s the admission that surprised me.
Staniszewski predicted last year that audio models would be commoditized within a couple of years.
He’s walking that back slightly. “There is still a lot of work to be done, and the quality delta you can achieve just on the model level is still significant.”
His new target is more ambitious: passing the Turing test for conversational AI. Not just sounding human. Understanding emotions. Knowing when to slow down or speak up. Reading the room.
“You need to combine intelligence, but you also need emotional intelligence,” he said. “That hasn’t yet been done.”
Three to five years, he estimates.
In the meantime, ElevenLabs has thousands of contractors annotating not just what was said in training data, but how it was said, including emotions, accents, and timing.
They brought in voice coaches to detect accents accurately.
At $600 million in revenue and 400 employees, ElevenLabs has bought itself time to get there. Whether it gets there before its customers do is the race that actually matters.

