Voice

How to choose a voice for your BPO or contact center AI agent

Most BPOs spend more time choosing their CRM than choosing their voice. The voice is the first three seconds of every call, and it sets the tone for everything that follows.

Cover image for bpo voice choice

Most BPOs spend more time choosing their CRM than choosing their voice. That is a mistake.

The voice is the first three seconds of every call. It is the first thing a caller hears, and it sets the tone for everything that follows. A collections call delivered in a warm, chatty voice feels dismissive. A support call delivered in a flat monotone feels unhelpful. An outbound sales call that takes 800 milliseconds to respond feels like talking to a wall.

Gartner predicts conversational AI will cut contact center labor costs by $80 billion in 2026. That savings only materializes if the voice keeps callers on the line long enough for the AI to resolve the issue. Get the voice wrong and callers hang up, the AI never gets a chance to help, and the savings evaporate.

The voice is not the platform

Most content about AI voice agents for contact centers is about choosing a platform. Which vendor has the best STT, which has the lowest latency, which has the most languages. That matters, but it is the wrong first question.

The first question is: what should the agent sound like?

This is not an aesthetic decision. Regal AI found that contact center teams prefer expressive voices in demos but switch to consistency in production. AssemblyAI found that natural-sounding voices are rated as more trustworthy and competent, which directly improves CSAT. Cartesia's blind tests showed their Sonic-2 model was preferred over ElevenLabs Flash V2 by 61.4% to 38.6%. The voice you pick changes completion rates, satisfaction scores, and how callers perceive your brand.

The problem is that most teams pick a voice by listening to a demo on studio headphones, then deploy it on an 8 kHz phone line and wonder why it sounds thin. Or they pick one voice for every call type and use a warm, friendly tone for collections, which makes the call feel like it is not being taken seriously.

What actually matters when choosing a voice

Five things, in order of impact:

Match the voice to the call type. This is the biggest lever. A support call needs warmth and patience because the caller is already frustrated. A collections call needs authority and clarity because the caller is stressed and needs to hear numbers correctly. A sales call needs energy and confidence because you are asking for their time. An appointment scheduling call needs efficiency because the caller wants to get in and out. Using one voice for all four is like sending the same salesperson to a funeral and a birthday party.

Test on a real phone line, not headphones. Phone audio strips high frequencies and compresses the signal. A voice that sounds rich and natural on studio monitors can sound muffled or robotic on a phone. If your BPO serves mobile callers, test on a mobile connection. If it serves landlines, test on a landline. The delivery channel determines the voice quality, not the demo player.

Prioritize consistency over expressiveness. Expressive voices win demos and lose in production. A voice that delivers the same sentence differently every time feels erratic to repeat callers and creates problems for compliance-heavy scripts. Regal AI's production data showed teams consistently chose consistency once they saw real call volume. Pick a voice that sounds the same on the 10,000th call as it does on the 1st.

Check pronunciation on your actual scripts. Generic demo scripts do not reveal how a voice handles your terminology. Test with your longest, most complex scripts. Include names, addresses, product terms, and industry language. Cartesia supports IPA for specialized terms like prescription drug names. Deepgram offers domain-tuned accuracy for regulated industries. Mispronouncing a caller's name or a product term in the first 10 seconds of a call erodes trust for the rest of the conversation.

Measure latency at production concurrency. A voice that responds in 200 ms at 10 concurrent calls may degrade at 100. The TTS time-to-first-audio is a significant portion of your total pipeline budget. If your total latency (STT plus LLM plus TTS) exceeds 500 ms, callers perceive a delay. Cartesia targets 40 ms TTFB. Telnyx NaturalHD runs on the same network as the call, so there is no cross-vendor network hop. ElevenLabs has higher quality but higher latency. The right tradeoff depends on your call type: support calls can tolerate slightly more latency than sales calls where you are asking for someone's time.

Voice profiles by BPO call type

Here is what works, based on what BPO teams deploying voice agents are reporting:

Customer support: Warm tone, moderate pace, high warmth. The caller is frustrated, so the voice should de-escalate. Think of the best human support agent you have talked to. They were not perky. They were calm, clear, and patient. Astra from Telnyx NaturalHD hits this profile. Luna is too energetic for support. Andersen is too cold.

Outbound sales: Energetic tone, quick pace, moderate warmth. You are interrupting someone's day, so the voice needs to sound like it belongs there. Confidence matters more than warmth. A slightly faster pace creates momentum without feeling pushy. Luna works here. Astra works if the pitch is technical. Andersen is too stiff for cold outreach.

Collections: Authoritative tone, steady pace, moderate warmth, high clarity on numbers. The caller is stressed about the topic. A warm voice feels dismissive. An energetic voice feels aggressive. You need clear, measured delivery that sounds professional without sounding cold. Andersen is built for this. The voice needs to pronounce dollar amounts and account numbers without ambiguity.

Appointment scheduling: Efficient tone, moderate pace, moderate warmth. The caller wants to get in, get the appointment, and get out. The voice should sound competent and organized. Luna works well here because it sounds helpful without being slow. Clear pronunciation of dates and times is non-negotiable.

Configure your voice agent

The dashboard below lets you set up a BPO voice agent and see the cost and latency impact in real time. Pick a call type, set your monthly volume, choose your priority, and listen to the recommended voice.

What the providers actually offer

The TTS market for BPO has consolidated around four providers, each with a clear position.

Telnyx NaturalHD runs on the same network as the call. Three voices: Astra (warm, professional), Luna (friendly, approachable), Andersen (authoritative, clear). Pay-as-you-go pricing of $0.000048 per character ($48 per million characters, roughly $0.036 per minute at a 750-character-per-minute speaking rate). The advantage is end-to-end latency control since there is no cross-vendor network hop between the TTS provider and the telephony layer.

ElevenLabs has the largest voice library and the highest voice quality in blind tests. 32 languages. Voice cloning from 30 seconds of audio. The tradeoff is latency, which runs higher than Cartesia or Telnyx. Best for BPOs that prioritize naturalness over response speed and need multilingual coverage.

Cartesia leads on latency at approximately 40 ms time-to-first-audio, optimized for 8 kHz phone interactions. Voice cloning from 3 seconds of audio. Accent localization so an American voice can speak with a French accent. Sonic-2 was preferred over ElevenLabs Flash V2 in blind tests. Best for BPOs where real-time responsiveness is the top priority.

Deepgram Aura-2 is built for regulated industries. Domain-tuned speech accuracy for healthcare, finance, and compliance-heavy verticals. English-only. On-premise deployment for data residency. Best for BPOs where terminology accuracy and on-premise deployment matter more than multilingual coverage.

How to test before you commit

Run a blind listening test. Record the same script with three to five voice candidates. Play them to a panel that matches your caller demographics. Do not label the voices. Ask which sounds most natural, trustworthy, and appropriate for the call type. The voice that wins the demo is not always the voice that wins the blind test.

Then A/B test in a pilot. Route 10% of calls to the new voice and 90% to the existing one. Measure completion rate, average handle time, and CSAT over two weeks. A voice that wins a blind test may lose in production because it does not handle real caller interruptions, background noise, or edge-case scripts.

Test at production concurrency. A voice that responds in 200 ms at 10 concurrent calls may degrade at 100. Test at the concurrency level your BPO actually operates at during peak hours. Latency spikes during peak load are what cause callers to hang up, not average latency.

The mistake most teams make

They optimize for voice quality only. They pick the most natural-sounding voice, deploy it across every call type, test it on headphones instead of a phone line, and never measure latency at production concurrency. Then they wonder why completion rates are lower than the demo suggested.

The voice is an operational decision, not an aesthetic one. Match it to the call type. Test it on the delivery channel. Prioritize consistency over expressiveness. Measure latency at production load. The three seconds at the start of every call are worth more than the three weeks you will spend choosing the CRM.

Share on Social
Abhishek Sharma
Abhishek Sharma
Sr Technical Product Marketing Manager

Senior Technical Product Marketing Manager