Where your model runs decides how fast it answers, what each request costs, and how often it fails. This guide compares the three ways to run AI inference, shows where the milliseconds go in a single request, and explains how to match a serving tier and region to each workload.

Takeaways
Inference infrastructure is the serving layer that turns a trained model into live responses: accelerators, a serving engine, the network path to the user, autoscaling, and the API in front of it. The decisions that move latency and cost are where each workload's GPU sits and which serving tier it runs on, which matters more than which cluster a team rents.
Inference is the forward pass that turns new input into output: a prompt goes in, tokens come out. Inference infrastructure is everything that makes that pass answer a live request quickly, reliably, and at a predictable cost, from the GPU that runs the model to the API your application calls.
Training builds a model's weights in a large job that eventually ends. Fine-tuning adapts those weights to a narrower task, also as a bounded job. Inference and serving never end while users keep sending requests, so capacity has to follow user demand instead of a project schedule.
That one difference explains most AI inference infrastructure decisions. The system is sized for peaks nobody can schedule, judged on latency and cost per token, and kept running indefinitely.
Every inference request passes through five layers, in this order:
Large language model inference also runs in two phases. Prefill processes the whole prompt and produces the first token. Decode then generates each following token one at a time, and every latency metric later in this article depends on that split.
Telnyx Inference is one complete instance of this stack. It is an OpenAI-compatible inference API that runs a curated library of open-weight models on GPUs Telnyx owns, with zero data retention and no infrastructure for the customer to manage.

Three realistic routes exist for running inference in production, and they differ mainly in who operates each layer of the stack above.
Which inference route fits my team?
A hosted inference API is the route most teams already use, usually with a proprietary model vendor. The application sends an HTTP request and gets tokens back, and the provider runs every layer underneath. The convenience is real. The risk sits in the vendor's model roadmap: when a proprietary model is retired, every team on it gets a deadline.
A hosted API serving open-weight models keeps the convenience and lowers that risk, because the API shape stays standard and the models are not tied to one lab. On Telnyx, pricing is per token at up to 75% less than proprietary alternatives, there are no minimums, and there are no GPU rental fees or compute surcharges. Capacity autoscales from 0 to 1000s of requests per second, and inference runs in-region across the Americas, Europe, MENA, and APAC with zero data retention.

This route also works for teams that keep their own orchestration. They can call a hosted LLM endpoint alongside speech-to-text and text-to-speech endpoints without replacing the platform they already run.
The second route rents GPU instances or a managed Kubernetes service from a large cloud provider and runs a serving engine on top. The team gains control over model choice, batching settings, and deployment topology.
It also inherits the work: requesting GPU quota, choosing instance types, operating the serving engine, and paying for idle capacity at nearly the same rate as busy capacity. Getting from provisioned GPUs to a first production request takes real engineering time, and the network path still runs from the cloud region to wherever the user happens to be.
The third route owns or leases the hardware outright. It offers full control over every layer and the lowest unit cost when utilization stays high and steady.
The trade is that the team owns everything: cluster provisioning, the Kubernetes lifecycle, GPU drivers, high-performance networking, model storage, autoscaling, and upgrades. The Kubernetes guide to GPU scheduling in Kubernetes shows only the starting point: install vendor drivers on every node, run a device plugin, and request GPUs as a schedulable resource.
A single inference request spends time in six places, and GPU placement changes only some of them.
Each request moves through the same sequence:
Model time in steps three and four scales with prompt and output length. Network time in steps one and six is fixed per hop, so a cross-region or multi-provider path pays it two or more times. For short interactive replies, the network can take a larger share of the budget than the model does. Our guide to inference latency breaks down each component and includes a script for measuring it from your users' region.
"We often obsess over LLM inference speeds, but in Voice AI, the network is often the silent killer." Ian Reither, COO at Telnyx |
Telnyx runs models on GPUs colocated with the network that serves the user, in-region across the Americas, Europe, MENA, and APAC. Requests make no inter-provider hops and no cross-border routing, which removes the extra network segments from the budget without touching model time.
For interactive work, the Priority tier selects faster serving capacity for low-latency workloads, including voice orchestration, without changing the model or the context window. Four measurements show how any provider will behave in production:
A voice agent runs speech-to-text, an LLM, and text-to-speech inside a single conversational turn, then delivers audio back to the caller. When each piece comes from a different vendor, every handoff adds a network round trip and its own jitter.
Teams that run their own orchestration can still cut those hops by placing the LLM endpoint on the same network as speech and telephony. Our write-up on co-located voice AI infrastructure explains how keeping audio, transcription, model, and speech on one network holds voice-to-voice response under a second.
Efficiency in inference is a utilization problem first and a tier-matching problem second, and Telnyx exposes both as choices made per request.
A GPU costs nearly the same idle as busy, and inference demand arrives when users arrive. A single team rarely keeps a dedicated cluster busy around the clock, so much of the bill pays for waiting.
The useful measure is tokens served per dollar. A hosted provider that pools many customers' traffic on GPUs it owns reaches utilization a single tenant cannot, and that shows up in the per-token price. Our guide to inference cost optimization works through a full example that takes a $50,000 monthly bill to about $12,400.
Paying for low latency makes sense only where latency earns revenue. Telnyx offers three serving tiers on the same OpenAI-compatible endpoint, selected with the service_tier field on each request.
Telnyx Inference serving tiers
| Tier | What it is for | How to select it |
|---|---|---|
| Default | Standard serving capacity at standard rates. All published pricing comparisons and benchmarks refer to this tier. | Omit service_tier |
| Priority | Faster serving capacity for low-latency interactive workloads, including voice orchestration. Same model and context window. Live on select models today. | "service_tier": "priority" |
| Flex | Lowest rates where minutes of end-to-end latency is acceptable, such as offline evaluation, document processing, and background summarization. Requests stay synchronous. Available in the US on select models. | "service_tier": "flex" |

Because the tier is one field in the request body, a team can route a voice agent to Priority and a nightly summarization job to Flex without adding a second inference provider.
Owned GPU infrastructure scales from 0 to 1000s of requests per second without capacity planning. Billing is per token, so low-volume and bursty workloads pay only for what they use, with no minimums to meet. Prompts that repeat, such as system prompts and few-shot templates, get cached input at up to 88% off.
Structural pricing: Up to 75% less than proprietary alternatives, cached input at up to 88% off, and no GPU rental fees, compute surcharges, or minimums. Per-token rates vary by model and serving tier, and every rate is listed on the inference API pricing page.
Reliable inference infrastructure delivers four things: capacity when traffic spikes, no cold-start penalty, output the application can parse every time, and an API that survives a model's retirement.
Benchmarks rarely show the failures teams meet in production. A cold start turns a fast call into a multi-second wait. A traffic spike exhausts reserved capacity, and noisy neighbors push P99 latency up.
Observability gaps hide which layer failed. Then a provider deprecates the model an application depends on and sets a migration deadline weeks away.
"Model quality is a solved problem across several providers. The hard part is running those models at the edge, co-located with telephony, so the full chain, STT to inference to TTS to network delivery, stays on one infrastructure. Nobody else runs the full chain end-to-end. Vapi, Retell, Bland, they resell other people's TTS and STT, which means multiple vendors, multiple hops, compounded latency. We host seven TTS engines on our own GPUs, Natural, NaturalHD, Ultra, Kokoro, Qwen3TTS, xAI Grok, Inworld, and route to seven more through the same API. The model is the easy part. The infrastructure underneath it is what makes it production-grade." Dr. Sonam Gupta, Developer Advocate at Telnyx |
Telnyx handles the first two failures with owned GPU infrastructure that serves concurrent requests and scales automatically, with no capacity planning and no cold starts. Structured output with JSON mode and regex constraints keeps every response in the application's schema, so downstream code never parses free text. Compute stays in-region with zero data retention, and Telnyx does not train on customer data.
A model roadmap is part of the infrastructure. Telnyx serves open-weight models from z.ai, DeepSeek, Minimax, and others through one OpenAI-compatible API, so replacing a retired model is a change to the model name in the request. Any provider, including a current one, can be tested against the same checklist:
Six practices cover most of the decisions in this guide:
Run the same prompts against each candidate from a machine in your users' region, at realistic concurrency and prompt lengths. A latency chart measured next to the GPU leaves out the network hop your users actually pay for.
The request below is the documented Telnyx Inference call to /v2/ai/chat/completions. An existing OpenAI SDK reaches the same endpoint by changing its base URL and credentials.
Omitting service_tier uses Default. Adding "service_tier": "priority" or "service_tier": "flex" to the JSON body selects the other tiers with the same moonshotai/Kimi-K3 model ID.
Inference on the same network as speech and telephony removes handoffs for voice and audio workloads. The ai-podcast-producer-python project in the Telnyx code examples repository shows the pattern: it records a multi-host podcast over a conference call, transcribes each speaker with speech-to-text, generates show notes and chapters with inference, and produces text-to-speech intro and outro bumpers. Teams with their own orchestration can adopt the same approach one endpoint at a time.
"What we see is a clear expansion path: teams land on inference for the models and the economics, and then realize the same infrastructure already runs STT, TTS, and telephony. The integration cost to go from text inference to voice agents is zero because it's the same network. We deliberately don't lead with that on the inference product because it's not why people show up. But it's why they stay." Regis David Souza, Staff Software Engineer at Telnyx |
FAQ
A proprietary model we depend on is being retired in a few weeks. How do we move to different inference infrastructure without rewriting the application?
Move to an OpenAI-compatible provider that serves open-weight models, then change the base URL, credentials, and model name in your existing SDK. Before cutting over, test tool calls, structured output, and context length on the new model, because those vary by model. Telnyx Inference uses the same chat completions request shape, so the application code itself stays the same.
Can we use an inference provider only for the LLM, speech-to-text and text-to-speech pieces and keep our own orchestration?
Yes. A hosted inference API is a set of HTTP endpoints, so it plugs into orchestration you already run. Telnyx runs LLM inference, speech-to-text, and text-to-speech on the same network, and each one can be called on its own. Keeping all three in one region removes vendor handoffs inside each voice turn without replacing your orchestration layer.
We are hitting latency and token-throughput limits on our current cloud provider's endpoint. What actually changes when inference runs on in-region GPUs?
Two things change. In-region GPUs remove cross-region and inter-provider network hops, which often dominate the latency of short interactive replies. Autoscaling on owned capacity removes the need to reserve throughput ahead of demand. Model compute time stays roughly the same, so measure TTFT and P99 from your users' region to see the real difference for your workload.
What is inference in AI, and which infrastructure decisions follow from it?
Inference is running a trained model on new input to produce output, such as generating a reply to a prompt. Because it runs every time a user sends a request, the decisions that follow are about serving: where the GPU sits relative to users, how capacity scales with traffic, what each token costs, and which API shape the application depends on.
What is AI inference versus training, and why do they need different infrastructure?
Training adjusts a model's weights over large datasets in a job that eventually finishes, so it needs maximum compute for a fixed period. Inference uses the finished weights to answer live requests and never finishes while users keep sending them. Training infrastructure is sized for a job. Inference infrastructure is sized for unpredictable demand and judged on latency and cost per token.
Give each workload the serving tier and in-region GPU it needs instead of renting a cluster for all of them. Move an existing OpenAI SDK app to Telnyx Inference by changing the base URL, with up to 75% savings against proprietary models.
Start buildingRelated articles
What is an inference engine? Types, uses, and vLLM

Open-Source Models Are Catching Up to Frontier

How to scan AGENTS.md and CLAUDE.md for indirect prompt injection

WhatsApp Business API Cost in 2026 After October 1

VoIP infrastructure explained: find the layer behind every bad call

The 9 best WhatsApp API providers in 2026
