Playbook · Business case to production

Voice agents, from business case to production. With the code.

Eight parts. Each one pairs a decision you need to make with the API call that makes it real. Pick your industry and every example on the page adapts.

8Parts, strategy to scale
5Industries, every example adapts
<200 msRound-trip latency
terminal
# install the Telnyx CLI, then create your first agentgo install github.com/team-telnyx/telnyx-cli/cmd/telnyx@latestexport TELNYX_API_KEY="YOUR_API_KEY"telnyx ai:assistants create \  --name "Front Desk" \  --model moonshotai/Kimi-K2.6 \  --instructions "You book, confirm and reschedule patient appointments for Riverside Family Clinic."

Examples for

Highlighted values change with your pick.

Part 01

Strategy

The goal, the platform decision and the business case. 10 minute read.

Pick one primary goal

Cost, coverage or conversion. An agent tuned for all three delivers none. Write the goal as a number you can read from the call log.

  • Coverage: answer the after-hours scheduling calls the clinic misses today.
  • Cost: take reminder and reschedule calls off the front desk.
  • Conversion: not the goal for v1. Do not optimize for it yet.
goal.yaml
# one goal, one metric, read from the call logindustry: Healthcaregoal: Reduce no showsmetric: no_show_ratebaseline: 18%target: 11%window: 90 daysowner: ops_leadnot_optimizing_for: [handle_time, csat]   # yet

Platform or stitched stack

A stitched stack takes telephony from one vendor, speech and the model from three more, and an orchestrator on top. It can win on model choice. It loses on vendor hops, on who owns an incident, and on the latency budget.

Where a stitched stack still wins: a team that already runs its own inference fleet, or a use case that needs a model no platform hosts. Telnyx supports that path too. Point the assistant at your own LLM endpoint.
ConcernStitched, 4 to 5 vendorsTelnyx
Media pathPublic internet between vendorsTelnyx's private global network, with edge PoPs in 9 regions
Round tripThe sum of every vendor hopUnder 200 ms round trip, GPU clusters co-located with telephony PoPs
Numbers and complianceSeparate number and compliance vendorsLicensed carrier in 45+ countries, numbers and voice coverage in 140+ countries
Model choiceAny modelTelnyx-hosted models, or your own LLM endpoint
Who you page at 2 a.m.Four or five status pagesOne

Part 02

Discovery

Call journeys, hotspots and the scope of v1. 8 minute read.

Score use cases before you write a prompt

Rate each call type on volume, how repetitive it is, and the blast radius when the agent gets it wrong. Start with the highest volume and the lowest blast radius. Everything else waits for v2.

Use real call reasons. Pull a month of dispositions from your current system. Once the pilot is live, AI Insights tags every conversation, so you can rescore with real data.
Call typeVolumeRepetitiveBlast radiusVerdict
Appointment rescheduleHighHighLowStart here
Prescription refill statusHighHighMediumv2
New patient intakeMediumMediumMediumv2
Clinical questionsLowLowHighNever

Scope v1 in one file

Supported intents, refused intents and the escalation rule. If a request is not on the list, the agent transfers. This file becomes the first version of your instructions.

scope.md
# scope.md  (this becomes v1 of the instructions)## Supported intents- reschedule_appointment- confirm_appointment- cancel_appointment- clinic_hours_and_directions## Refused- clinical advice- billing disputes- anything about another patient## Escalation ruleTransfer to the front desk queue after one clarifying question.## Done whenno_show_rate moves from 18% toward 11% in 90 days

Part 03

Design

Instructions, persona, dynamic variables and escalation. 14 minute read.

Write instructions the model can obey

Short sections, one rule per line, and the escalation path spelled out. Callers interrupt and change topic. The instructions decide what happens when they do.

  • Tool first: never confirm an action before the tool returns.
  • One clarifying question, then transfer.
  • Never read back a full card, account or record number.
POST /v2/ai/assistants
curl -X POST https://api.telnyx.com/v2/ai/assistants \  -H "Authorization: Bearer $TELNYX_API_KEY" \  -H "Content-Type: application/json" \  -d '{    "name": "Front Desk",    "model": "moonshotai/Kimi-K2.6",    "instructions": "You book, confirm and reschedule patient appointments for Riverside Family Clinic. The caller is {{full_name}}. Never confirm an action before the tool returns. Ask one clarifying question, then transfer.",    "greeting": "Thanks for calling Riverside Family Clinic. This call may be recorded. How can I help?",    "voice_settings": { "voice": "Telnyx.KokoroTTS.af_heart" },    "transcription": { "model": "deepgram/flux" },    "telephony_settings": { "noise_suppression": "krisp" },    "dynamic_variables": { "full_name": "there" }  }'

Personalize at call time, not in the prompt

Caller name, location and department arrive as dynamic variables with each call. Send them in the API request, in a SIP header, or from a webhook your server answers when the call starts. The prompt stays short and identical for every caller.

POST /v2/texml/ai_calls/{connection_id}
curl -X POST https://api.telnyx.com/v2/texml/ai_calls/{connection_id} \  -H "Authorization: Bearer $TELNYX_API_KEY" \  -H "Content-Type: application/json" \  -d '{    "From": "+13125550100",    "To": "+13125550123",    "AIAssistantId": "{assistant_id}",    "AIAssistantDynamicVariables": {      "full_name": "Maria Alves",      "location": "Riverside Family Clinic",      "department": "Primary Care"    }  }'# the instructions reference them as {{full_name}}, {{location}}, {{department}}# inbound: send X-Full-Name as a SIP header, or answer the dynamic variables webhook

Part 04

Build

Latency budget, tool contracts, telephony and data. 18 minute read.

Spend the latency budget on purpose

A turn feels natural at under about a second, voice to voice. Every component spends from that budget. Measure each one with your real prompt, not a test prompt.

What Telnyx removes: the hops between vendors. GPU clusters are co-located with telephony PoPs and voice traffic stays on Telnyx's private global network, for under 200 ms round-trip latency.

Your turn latency

930 ms

Inside the 1,000 ms budget, with 70 ms of headroom for tool calls.

200 msEndpointing and transcription
400 msTime to first token with your real prompt and tool calls
180 msTime to first audio byte
150 msTelnyx private global network, under 200 ms round trip

Example values for planning. Replace them with your own measurements.

Treat every tool as a contract

Purpose, typed inputs, a timeout, and an error taxonomy mapped to spoken responses. Any gap you leave undefined, the model fills by guessing.

  • An idempotency key on every tool that changes state.
  • A timeout on every tool. Slow work runs async with a filler line.
  • Every error code gets a spoken response and a test.
tools on the assistant
"tools": [  {    "type": "webhook",    "webhook": {      "name": "book_appointment",      "description": "Book a confirmed appointment slot for a verified patient",      "url": "https://api.riversideclinic.example/v1/appointments",      "method": "POST",      "timeout_ms": 3000,      "headers": [{ "name": "Authorization", "value": "Bearer YOUR_TOOL_TOKEN" }],      "body_parameters": {        "type": "object",        "properties": {          "patient_id": { "type": "string" },          "slot_id": { "type": "string" },          "idempotency_key": { "type": "string" }        },        "required": ["patient_id", "slot_id", "idempotency_key"]      }    }  },  {    "type": "transfer",    "transfer": {      "from": "+13125550100",      "targets": [{ "name": "Front desk", "to": "+13125550199" }]    }  }]# error taxonomy, each mapped to a spoken response in the instructions# NO_AVAILABILITY       offer the next two slots# PATIENT_NOT_FOUND     verify date of birth once, then transfer# EHR_TIMEOUT           retry once, then offer a callback

Telephony is invisible until it fails

Outbound answer rates collapse when numbers get flagged. On Telnyx, caller ID name, STIR/SHAKEN signing and number reputation sit with the same carrier that runs the agent.

  • Warm up new numbers over two weeks.
  • Channels = peak calls per hour × handle time in hours, plus 30%.
  • Recording consent per jurisdiction, stated in the greeting.

Concurrency calculator

Provision for peak plus 30%. Assumes calls spread evenly across the window.

10 channels

numbers and greeting
# caller ID name on each outbound number, 15 characters maxcurl -X PATCH https://api.telnyx.com/v2/phone_numbers/{phone_number_id}/voice \  -H "Authorization: Bearer $TELNYX_API_KEY" \  -H "Content-Type: application/json" \  -d '{ "cnam_listing": { "cnam_listing_enabled": true, "cnam_listing_details": "RIVERSIDE CARE" } }'# recording consent in the greeting, background noise suppressedcurl -X POST https://api.telnyx.com/v2/ai/assistants/{assistant_id} \  -H "Authorization: Bearer $TELNYX_API_KEY" \  -H "Content-Type: application/json" \  -d '{    "greeting": "Thanks for calling Riverside Family Clinic. This call may be recorded. How can I help?",    "telephony_settings": { "noise_suppression": "krisp" }  }'

Map the data surface

Audio, transcript, model context, tool logs and storage. Classify each one and set its control before launch.

Keep PHI on Telnyx-hosted models. That is what the Telnyx HIPAA architecture guide recommends. Confirm a BAA covers your Voice AI workload before PHI reaches the agent.
DataClassControl
Audio recordingPHIConsent in the greeting, retention period set
TranscriptPHIAccess by role, every read audited
LLM contextPHITelnyx-hosted model, nothing kept after the call
Tool logsPHIPatient ID only, never date of birth
StoragePHIBAA in place with every vendor in the path

Part 05

Test

Tool tests, conversation tests and the pilot. 8 minute read.

Test tools, then turns, then calls

Test each tool against every error code first. Then run scripted conversations scored against a rubric. Then put real calls on it: a pilot on a slice of inbound, with exit criteria written down before it starts.

  • Tool tests run in your own test suite.
  • Conversation tests run on Telnyx, by phone or web call.
  • Run the suite against every new version before you promote it.
POST /v2/ai/assistants/tests
# conversation test: a scripted caller, scored against a rubriccurl -X POST https://api.telnyx.com/v2/ai/assistants/tests \  -H "Authorization: Bearer $TELNYX_API_KEY" \  -H "Content-Type: application/json" \  -d '{    "name": "Front Desk happy path",    "destination": "+13125550100",    "telnyx_conversation_channel": "phone_call",    "instructions": "Caller wants to move a Tuesday appointment to Thursday afternoon.",    "rubric": [      { "name": "Tool first", "criteria": "Never confirms a booking before book_appointment returns" },      { "name": "Refusal", "criteria": "Declines clinical questions and offers a transfer" }    ],    "max_duration_seconds": 180  }'# run it against a new version before you promote that versioncurl -X POST https://api.telnyx.com/v2/ai/assistants/tests/{test_id}/runs \  -H "Authorization: Bearer $TELNYX_API_KEY" \  -H "Content-Type: application/json" \  -d '{ "destination_version_id": "{version_id}" }'

Part 06

Launch

Graduated rollout, the readiness gate and change management. 8 minute read.

Roll out by percentage, not by date

Save the new instructions as a version, send a slice of traffic to it, and watch the same dashboards you will watch for the life of the agent. Promote when the numbers hold. Rolling back is one call.

versions and canary deploy
# save new instructions as a version, without promoting itcurl -X POST https://api.telnyx.com/v2/ai/assistants/{assistant_id} \  -H "Authorization: Bearer $TELNYX_API_KEY" \  -H "Content-Type: application/json" \  -d '{ "instructions": "<v2 instructions for Front Desk>", "promote_to_main": false }'# send 10% of calls to the new version; the rest stay on maincurl -X POST https://api.telnyx.com/v2/ai/assistants/{assistant_id}/canary-deploys \  -H "Authorization: Bearer $TELNYX_API_KEY" \  -H "Content-Type: application/json" \  -d '{ "rules": [ { "serve": { "rollout": [ { "version_id": "{version_id}", "weight": 10 } ] } } ] }'# numbers hold for a week: promote. Otherwise delete the canary deploy.curl -X POST https://api.telnyx.com/v2/ai/assistants/{assistant_id}/versions/{version_id}/promote \  -H "Authorization: Bearer $TELNYX_API_KEY"

Hold a readiness gate

No agent reaches production until every item is checked. Add a week to the plan for it. It saves six.

0 of 6 checked

Not yet

Part 07

Operate and improve

Three layers of monitoring, structured insights and the weekly loop. 12 minute read.

Monitor at three layers

Infrastructure: call setup and latency per component. Conversation: containment, transfer reasons and silence gaps. Business: the number from Part 01. Alert on the first two. Review the third weekly.

event stream and tracing
# stream conversation events to your monitoring, trace every turn in Langfusecurl -X POST https://api.telnyx.com/v2/ai/assistants/{assistant_id} \  -H "Authorization: Bearer $TELNYX_API_KEY" \  -H "Content-Type: application/json" \  -d '{    "websocket_settings": {      "enabled": true,      "url": "wss://events.riversideclinic.example/telnyx",      "auth_ref": "events_token"    },    "observability_settings": {      "status": "enabled",      "host": "https://cloud.langfuse.com",      "public_key_ref": "langfuse_public_key",      "secret_key_ref": "langfuse_secret_key"    }  }'# thresholds live in your own alerting, for example:#   p95 turn latency  > 1000 ms     page#   transfer rate     > 35%         review#   no_show_rate     weekly business review

Mine every call with structured insights

Define the fields you want from each conversation as a JSON schema, and send the results to a webhook. Your weekly review reads a table, not a folder of recordings.

AI Insights
# 1. define the fields to extract from every conversationcurl -X POST https://api.telnyx.com/v2/ai/conversations/insights \  -H "Authorization: Bearer $TELNYX_API_KEY" \  -H "Content-Type: application/json" \  -d '{    "name": "no_show_drivers",    "instructions": "Extract the fields below from the conversation.",    "json_schema": {      "type": "object",      "properties": {        "reschedule_reason": { "type": "string" },        "new_slot_offered": { "type": "boolean" },        "transferred": { "type": "boolean" },        "transfer_reason": { "type": "string" }      }    }  }'# 2. group insights and deliver results to your webhookcurl -X POST https://api.telnyx.com/v2/ai/conversations/insight-groups \  -H "Authorization: Bearer $TELNYX_API_KEY" \  -H "Content-Type: application/json" \  -d '{ "name": "weekly_review", "webhook": "https://api.riversideclinic.example/telnyx/insights" }'# 3. assign the insight to the group, then set it on the assistant:#    "insight_settings": { "insight_group_id": "{group_id}" }curl -X POST https://api.telnyx.com/v2/ai/conversations/insight-groups/{group_id}/insights/{insight_id}/assign \  -H "Authorization: Bearer $TELNYX_API_KEY"

Part 08

Scale

From one agent to five, handling drift, and building the skill in house. 8 minute read.

Start the next agent from a clone

The second agent starts from the first: shared tools, shared escalation rules, a different scope file. A clone copies everything except telephony and messaging settings. Version every change and keep the last version ready to promote.

Already built somewhere else? Import assistants from Vapi, ElevenLabs or Retell, with their configuration.
clone and import
# the second agent starts from the firstcurl -X POST https://api.telnyx.com/v2/ai/assistants/{assistant_id}/clone \  -H "Authorization: Bearer $TELNYX_API_KEY"# rename it, then assign numbers: clones skip telephony and messaging settingscurl -X POST https://api.telnyx.com/v2/ai/assistants/{new_assistant_id} \  -H "Authorization: Bearer $TELNYX_API_KEY" \  -H "Content-Type: application/json" \  -d '{ "name": "Front Desk Dental" }'# already running on Vapi, ElevenLabs or Retell? import itcurl -X POST https://api.telnyx.com/v2/ai/assistants/import \  -H "Authorization: Bearer $TELNYX_API_KEY" \  -H "Content-Type: application/json" \  -d '{ "provider": "vapi", "api_key_ref": "vapi_api_key" }'

Handle drift

Prompts drift as people add rules. Backends drift as APIs change. Callers drift as the business changes. A monthly review of transfer reasons catches all three.

  • A length budget for instructions. Anything over it moves to a tool or a knowledge source.
  • Tool contracts versioned with the backend, not with the prompt.
  • One named owner per agent, written in the assistant description.
tooling for your coding agent
# Telnyx skills for your coding agentnpx skills add team-telnyx/ai --skill <SKILL> --agent <AGENT># or, in Claude Code, the Telnyx plugin/plugin marketplace add team-telnyx/ai/plugin install telnyx-ai@telnyx# remote MCP server, authenticated with your API keyhttps://api.telnyx.com/v2/mcpAuthorization: Bearer $TELNYX_API_KEY

Where it runs

Every part of this playbook runs on infrastructure Telnyx owns.

One platform, not a list of vendor recommendations. Carrier, network and inference sit together and are exposed through the APIs above.

Edge PoPs
9 regions
Round-trip latency
<200 ms
Licensed carrier
45+ countries
Numbers and voice coverage
140+ countries

Carrier

A licensed carrier

Telnyx originates calls as a licensed carrier in 45+ countries, with numbers and voice coverage in 140+ countries.

Network

A private global network

Voice traffic stays on Telnyx's private global network, with edge PoPs in 9 regions and under 200 ms round-trip latency.

Inference

GPUs beside the media plane

GPU clusters are co-located with telephony PoPs, so speech, model and voice run where the call does.

Identity

Trusted caller ID

A-level STIR/SHAKEN attestation on eligible US outbound calls, plus Branded Calling and Number Reputation.

Edge Compute

Your code next to the call

Functions, KV, Object Storage and Inference, colocated with the carrier facilities your calls terminate on.

Compliance

Compliance built in

SOC 2 Type II, ISO/IEC 27001, PCI DSS and GDPR. For PHI, confirm BAA coverage with your account team.

Who this is for

Read the parts that match your job.

Executive leadership

Business case, risk boundaries and the platform decision.

Product managers

Use case selection, scoping and the roadmap from v1 to v5.

Engineering leaders

Latency budget, tool contracts, telephony and observability.

Operations and CX

Escalation, QA, change management and frontline adoption.

Start with the quickstart. Ship the pilot this month.

Portal, CLI or API. The same assistant either way.