Skip to content
tokynstudio
journal
21 September 2026 · agents

Why we build voice agents on ElevenLabs

Not because of the voice - plenty of vendors deliver that now. Because of latency across a full conversational turn, telephony without bolt-on parts, and an EU region. Plus where we advise against it, and what the partnership actually changes.

by tokyn studio · 6 min read

Why we build voice agents on ElevenLabs - AgentsAI

TL;DR. We build voice agents on ElevenLabs, and we are now an official implementation partner. The reason is not voice quality, which several vendors now deliver, but everything around it: latency across a full conversational turn, telephony without bolt-on parts, access to the customer's own systems, testability before go-live, and an EU region for clients who need one. What gets misread about this, and where we advise against it, is further down. Last updated: 2026-09.

what decides a phone call is not the voice

The most common misconception about voice agents is that this is about sound quality. It isn't. Natural-sounding speech synthesis is not a distinction in 2026, it is table stakes. Several vendors deliver voices that nobody picks out over a phone line.

What carries a conversation, or sinks it, is the silence in between. A caller does not hear the voice, they hear the pause. If nothing happens for a second after their question, they know they are talking to a machine before the first word arrives. Equally decisive: what happens when they cut in mid-sentence. People do that constantly.

So we do not judge platforms by listening samples. We judge them by what happens between the microphone and the answer.

the 75 milliseconds everyone quotes

ElevenLabs cites roughly 75 milliseconds for its Flash model. The number is accurate and almost always read wrong.

Those 75 ms are time to first audio byte from the speech model. A full conversational turn contains more: the agent has to detect that the caller has finished, transcribe the audio, let the language model think, possibly query a business system, and only then speak. Network and phone carrier sit on top.

Choosing a voice platform by its TTS figure therefore compares one component, not the outcome. What we measure instead is the time from the end of the caller's question to the first audible word, on a real phone line, with the customer's real data lookups in place. That number is always well above 75 ms, and it is the only one the caller experiences.

ElevenLabs still wins here, but for a different reason than the one on the spec sheet: because speech recognition, turn detection, model access and output all sit on one platform, half the handoffs that usually cost time simply disappear. The shortest chain wins, not the fastest single part.

the four things that convinced us

Telephony is not an add-on. The agent runs across telephony, web, mobile, WhatsApp, SMS and chat, with integrations for Twilio, Genesys and Amazon Connect. For mid-market projects that is the difference between "runs on the phone system" and "runs in a demo browser".

Access to what is already there. Through MCP, ordinary APIs and webhooks, the agent reaches CRM, ticketing, calendar and knowledge base. That is the actual lever for us, because an agent that can read appointments but not write them saves nobody any work. Why MCP is the right standard for this, we wrote up in AI agents and MCP.

Languages that switch mid-conversation. ElevenLabs states 70+ languages and more than 16,000 voices, with switching inside a single call. That is the vendor's figure, not our measurement. In our projects only one case of it usually matters, but it matters often: the caller switches to English and the agent follows without anyone pressing a key.

Testing before go-live. Simulated conversations and automated tests against agent behaviour are part of the platform. It sounds unremarkable and it is where we save the most time. Otherwise a voice agent can only be checked by calling it, and that does not scale to fifty conversation paths.

the data protection part, honestly counted

For clients with EU requirements we work in an Enterprise account with EU data residency. That is a genuine argument, and it is narrower than it sounds in sales material.

What holds: storage sits in the selected region. What also holds, and is rarely said out loud: processing can still leave that region, including via group affiliates and subprocessors. Keeping processing actually inside the EU requires additionally enabling Zero Retention Mode and configuring API access accordingly. That is not a default, it is a configuration step.

Three further constraints belong on the table before anything is promised to a client. Data residency is reserved for Enterprise customers. Not every language model is available in the isolated region, because that depends on the respective providers. And some features, dubbing among them, are not there at all. The isolated account is also a separate workspace with its own keys, not a switch inside the existing one.

None of this is disqualifying. But it is the difference between "sits in the EU" and "stays in the EU", and that difference belongs in the assessment before it surfaces in a data protection impact assessment.

what the partnership changes

We are an official implementation partner for ElevenLabs. What that means in practice is short, and deliberately unspectacular.

It does not mean a different feature set. The platform is the same one anyone can buy. It means a direct line when something in a client project gets stuck that public support will not resolve, and it means we know earlier what is coming. For projects with a fixed go-live date the second is often worth more than the first, because you then avoid building on features that get renamed four weeks later.

What it explicitly does not mean: that we recommend ElevenLabs because of it. The order was the other way round. We were building on the platform before there was a partnership.

when we advise against it

There are three cases where we steer clients away from ElevenLabs.

When the use case is not a conversation at all. Anyone who needs an announcement that changes once a year needs an audio file, not a conversational platform.

When the requirement is genuinely on-premise processing. An EU region is not the same as your own data centre. Where that is mandated, the path runs through locally hosted models, with everything that costs in quality and effort.

And when the real bottleneck sits elsewhere. If the data is not maintained, or nobody has decided what the agent is allowed to commit to, the speech platform is not the problem to solve first. That is the same lesson as in Copilot as a company GPT: the technology is rarely the hard part.

hear it yourself

Our AI agents page has three voice agents you can call directly. They run on exactly the stack this piece is about. The first sentence of the call states that this is an AI voice; the portraits on the page are AI-generated and labelled as such.

If you are weighing whether a voice agent would carry in your own organisation: 30 minutes, no pitch deck. We will also say so if we think it is the wrong route.

sources

related service

Agents

next step

your case, concretely - let's talk.

30 minutes, no pitch deck. We look at your use case and tell you honestly whether - and how - it's worth doing.

Why we build voice agents on ElevenLabs · tokyn studio