Quick answer
Voxovo AI ships with Cartesia as its default TTS provider, delivering a fully managed voice stack that runs on Voxovo’s low‑latency infrastructure. If you prefer to bring your own TTS service, you can plug in ElevenLabs via an API key, turning Voxovo into a BYOK (Bring Your Own Kit) platform for text‑to‑speech. The choice between the two paths is fundamentally about control versus convenience.
The Cartesia stack is a turnkey solution: it integrates seamlessly with Voxovo’s Speed, Smart, and Ultra pipelines, requires no external key management, and benefits from Voxovo’s presence buffer that eliminates dead air on every call. Because the TTS is managed, latency is predictable and the turn‑time scorecards show a consistent 10–20 ms advantage over most external providers. However, the trade‑off is that you are locked into Cartesia’s voice library and pricing model.
With ElevenLabs as a BYOK TTS, you gain granular control over voice selection, model version, and cost per character. Integration requires adding a private API key to your Voxovo workspace, which means your team handles key rotation and uptime monitoring. The benefit is that you can tailor the voice to your brand or regulatory requirements while still using Voxovo’s core pipeline and presence buffer.
How Voxovo AI works on Voxovo
Both stacks benefit from the same presence buffer, a pre‑cached filler audio that guarantees callers never hear dead air while the system processes LLM or tool calls. This buffer works independently of the TTS provider, so switching from Cartesia to ElevenLabs does not impact the dead‑air mitigation strategy.
Live Sheet and Live Calendar are first‑party data tools that sit inside the call flow; they are agnostic to the TTS layer. Whether you’re using Cartesia or ElevenLabs, your agent can pull CRM data, schedule appointments, or create tickets in real time, improving agent productivity and reducing friction.
The Voxovo 1.0 modes—Speed, Smart, Ultra—define the overall pipeline performance. Cartesia is the default voice for all three modes, offering the lowest turn latency for Speed, balanced accuracy for Smart, and the best voice realism for Ultra. ElevenLabs can be paired with any mode, but you may need to test each to confirm that the perceived latency matches the managed stack’s performance.
When you look beyond Voxovo, competitors such as Vapi, Retell, and Bland emphasize different strengths. Vapi focuses on developer APIs and assistants, Retell on ultra‑low‑latency telephony, and Bland on outbound dialer scale. None of these vendors provide the same blend of BYO numbers, mid‑call tool integration, and agency‑centric workspace support that Voxovo offers.