Inference

Inference Infrastructure: What to Run Where and Why It Matters

Where your model runs decides how fast it answers, what each request costs, and how often it fails. This guide compares the three ways to run AI inference, shows where the milliseconds go in a single request, and explains how to match a serving tier and region to each workload.

Inference infrastructure feature image

Takeaways

Inference infrastructure is the serving layer that turns a trained model into live responses: accelerators, a serving engine, the network path to the user, autoscaling, and the API in front of it. The decisions that move latency and cost are where each workload's GPU sits and which serving tier it runs on, which matters more than which cluster a team rents.

  • Teams have three realistic routes: a hosted inference API on provider-owned GPUs (fastest to production, least to operate), a hyperscaler GPU cloud (more control, plus quota limits and a serving stack to run), or self-managed bare metal (full control and the full operational burden).
  • Latency is a budget of network round trip, queueing, time to first token, and decode. The network hop is the part most benchmarks never measure.
  • A hosted API with in-region GPUs and per-request serving tiers lets a team change a base URL and one request field instead of hiring a platform team.

What Inference infrastructure is

Inference is the forward pass that turns new input into output: a prompt goes in, tokens come out. Inference infrastructure is everything that makes that pass answer a live request quickly, reliably, and at a predictable cost, from the GPU that runs the model to the API your application calls.

Training, fine-tuning, inference and serving

Training builds a model's weights in a large job that eventually ends. Fine-tuning adapts those weights to a narrower task, also as a bounded job. Inference and serving never end while users keep sending requests, so capacity has to follow user demand instead of a project schedule.

That one difference explains most AI inference infrastructure decisions. The system is sized for peaks nobody can schedule, judged on latency and cost per token, and kept running indefinitely.

The layers between a GPU and a response

Every inference request passes through five layers, in this order:

  1. Accelerator and memory: The GPU holds the model weights and the KV cache, which stores intermediate attention results for each active request.
  2. Serving engine: Software such as vLLM, TensorRT-LLM, or Triton batches requests and manages KV cache memory. The vLLM documentation covers continuous batching, prefix caching, and PagedAttention, and our inference engine explainer traces where these engines came from.
  3. Network path: The route between the GPU and the user, including every provider boundary the request crosses.
  4. Autoscaling: The logic that adds and removes capacity as traffic changes.
  5. API: The endpoint, authentication, and request format your code depends on.

Large language model inference also runs in two phases. Prefill processes the whole prompt and produces the first token. Decode then generates each following token one at a time, and every latency metric later in this article depends on that split.

Telnyx Inference is one complete instance of this stack. It is an OpenAI-compatible inference API that runs a curated library of open-weight models on GPUs Telnyx owns, with zero data retention and no infrastructure for the customer to manage.

Diagram: The layers between a GPU and a response

Try Telnyx Inference to serve models on in-region GPUs with per-request serving tiers behind a hosted API.

Top AI inference infrastructure solutions compared

Three realistic routes exist for running inference in production, and they differ mainly in who operates each layer of the stack above.

Which inference route fits my team?

  • Spiky traffic, stock open-weight model: call Telnyx Inference and pay per token. Capacity scales from 0 to 1000s of requests per second, with no minimums or GPU rental fees.
  • Regulated data, users outside the US: use Telnyx Inference in the user's region. GPUs run in the Americas, Europe, MENA and APAC with zero data retention.
  • On a proprietary model facing retirement: move to Telnyx Inference and change the base URL. Open-weight models behind an OpenAI-compatible API cost up to 75% less per token.
  • Custom fine-tune, platform team in place: rent hyperscaler GPUs and run vLLM or TensorRT-LLM yourself. You gain control but inherit GPU quota limits and a serving stack to operate.
  • Must own every layer, ops staff to match: go self-managed bare metal. You control accelerators, network path and autoscaling, and you carry the full operational burden for all five layers.

AI inference layers

Hosted inference APIs on provider-owned GPUs

A hosted inference API is the route most teams already use, usually with a proprietary model vendor. The application sends an HTTP request and gets tokens back, and the provider runs every layer underneath. The convenience is real. The risk sits in the vendor's model roadmap: when a proprietary model is retired, every team on it gets a deadline.

A hosted API serving open-weight models keeps the convenience and lowers that risk, because the API shape stays standard and the models are not tied to one lab. On Telnyx, pricing is per token at up to 75% less than proprietary alternatives, there are no minimums, and there are no GPU rental fees or compute surcharges. Capacity autoscales from 0 to 1000s of requests per second, and inference runs in-region across the Americas, Europe, MENA, and APAC with zero data retention.

Diagram: Hosted inference APIs on provider-owned GPUs

This route also works for teams that keep their own orchestration. They can call a hosted LLM endpoint alongside speech-to-text and text-to-speech endpoints without replacing the platform they already run.

Hyperscaler GPU clouds and managed Kubernetes

The second route rents GPU instances or a managed Kubernetes service from a large cloud provider and runs a serving engine on top. The team gains control over model choice, batching settings, and deployment topology.

It also inherits the work: requesting GPU quota, choosing instance types, operating the serving engine, and paying for idle capacity at nearly the same rate as busy capacity. Getting from provisioned GPUs to a first production request takes real engineering time, and the network path still runs from the cloud region to wherever the user happens to be.

Self-managed bare metal

The third route owns or leases the hardware outright. It offers full control over every layer and the lowest unit cost when utilization stays high and steady.

The trade is that the team owns everything: cluster provisioning, the Kubernetes lifecycle, GPU drivers, high-performance networking, model storage, autoscaling, and upgrades. The Kubernetes guide to GPU scheduling in Kubernetes shows only the starting point: install vendor drivers on every node, run a device plugin, and request GPUs as a schedulable resource.

Top AI infrastructure for fast inference: where the milliseconds go

A single inference request spends time in six places, and GPU placement changes only some of them.

The end-to-end latency budget

Each request moves through the same sequence:

  1. Network round trip from the client to the API
  2. Queueing while the provider batches requests
  3. Prefill, which ends at time to first token (TTFT)
  4. Decode, which costs time per output token multiplied by output length
  5. Post-processing
  6. Return path to the user

Inference latency budget

Model time in steps three and four scales with prompt and output length. Network time in steps one and six is fixed per hop, so a cross-region or multi-provider path pays it two or more times. For short interactive replies, the network can take a larger share of the budget than the model does. Our guide to inference latency breaks down each component and includes a script for measuring it from your users' region.

Why in-region GPUs change the result

Telnyx runs models on GPUs colocated with the network that serves the user, in-region across the Americas, Europe, MENA, and APAC. Requests make no inter-provider hops and no cross-border routing, which removes the extra network segments from the budget without touching model time.

For interactive work, the Priority tier selects faster serving capacity for low-latency workloads, including voice orchestration, without changing the model or the context window. Four measurements show how any provider will behave in production:

  • Time to first token (TTFT): what a chat user or caller waits through before anything appears
  • Time per output token: how fast the rest of the reply streams
  • P99 end-to-end latency under concurrent load: the slowest requests, measured with varied input lengths, which benchmark averages hide
  • Voice turn time: end of caller speech to first synthesized audio

Real-time voice: three inference hops in one turn

A voice agent runs speech-to-text, an LLM, and text-to-speech inside a single conversational turn, then delivers audio back to the caller. When each piece comes from a different vendor, every handoff adds a network round trip and its own jitter.

Teams that run their own orchestration can still cut those hops by placing the LLM endpoint on the same network as speech and telephony. Our write-up on co-located voice AI infrastructure explains how keeping audio, transcription, model, and speech on one network holds voice-to-voice response under a second.

The most efficient AI infrastructure for inference under load

Efficiency in inference is a utilization problem first and a tier-matching problem second, and Telnyx exposes both as choices made per request.

Utilization decides cost per token

A GPU costs nearly the same idle as busy, and inference demand arrives when users arrive. A single team rarely keeps a dedicated cluster busy around the clock, so much of the bill pays for waiting.

The useful measure is tokens served per dollar. A hosted provider that pools many customers' traffic on GPUs it owns reaches utilization a single tenant cannot, and that shows up in the per-token price. Our guide to inference cost optimization works through a full example that takes a $50,000 monthly bill to about $12,400.

Inference savings

Match the serving tier to the workload

Paying for low latency makes sense only where latency earns revenue. Telnyx offers three serving tiers on the same OpenAI-compatible endpoint, selected with the service_tier field on each request.

Telnyx Inference serving tiers

TierWhat it is forHow to select it
DefaultStandard serving capacity at standard rates. All published pricing comparisons and benchmarks refer to this tier.Omit service_tier
PriorityFaster serving capacity for low-latency interactive workloads, including voice orchestration. Same model and context window. Live on select models today."service_tier": "priority"
FlexLowest rates where minutes of end-to-end latency is acceptable, such as offline evaluation, document processing, and background summarization. Requests stay synchronous. Available in the US on select models."service_tier": "flex"

Diagram: Match the serving tier to the workload

Because the tier is one field in the request body, a team can route a voice agent to Priority and a nightly summarization job to Flex without adding a second inference provider.

Scaling without capacity planning

Owned GPU infrastructure scales from 0 to 1000s of requests per second without capacity planning. Billing is per token, so low-volume and bursty workloads pay only for what they use, with no minimums to meet. Prompts that repeat, such as system prompts and few-shot templates, get cached input at up to 88% off.

Structural pricing: Up to 75% less than proprietary alternatives, cached input at up to 88% off, and no GPU rental fees, compute surcharges, or minimums. Per-token rates vary by model and serving tier, and every rate is listed on the inference API pricing page.

The most reliable infrastructure for AI inference in production

Reliable inference infrastructure delivers four things: capacity when traffic spikes, no cold-start penalty, output the application can parse every time, and an API that survives a model's retirement.

What breaks after launch

Benchmarks rarely show the failures teams meet in production. A cold start turns a fast call into a multi-second wait. A traffic spike exhausts reserved capacity, and noisy neighbors push P99 latency up.

Observability gaps hide which layer failed. Then a provider deprecates the model an application depends on and sets a migration deadline weeks away.

Owned capacity, no cold starts, structured output

Telnyx handles the first two failures with owned GPU infrastructure that serves concurrent requests and scales automatically, with no capacity planning and no cold starts. Structured output with JSON mode and regex constraints keeps every response in the application's schema, so downstream code never parses free text. Compute stays in-region with zero data retention, and Telnyx does not train on customer data.

Model retirement is an infrastructure risk

A model roadmap is part of the infrastructure. Telnyx serves open-weight models from z.ai, DeepSeek, Minimax, and others through one OpenAI-compatible API, so replacing a retired model is a change to the model name in the request. Any provider, including a current one, can be tested against the same checklist:

  • Capacity scales with concurrency without a reservation
  • The first request after an idle period carries no cold-start penalty
  • Output can be constrained with JSON mode or regex
  • The model library and API shape survive a single model's retirement
  • Data stays in-region with zero retention

AI inference infrastructure best practices

Six practices cover most of the decisions in this guide:

  1. Measure before choosing. Record TTFT, time per output token, and P99 under load for each workload type.
  2. Keep the API shape. An OpenAI-compatible interface turns a provider or model change into a base URL and model-name edit.
  3. Pick a serving tier per request. Default, Priority, and Flex replace a separate cluster for each workload.
  4. Constrain output. JSON mode or regex constraints stop parse failures before they reach the application.
  5. Keep inference in-region. Place compute where users and data already are.
  6. Plan for model churn. Treat model choice as an infrastructure lever and test new open-weight releases as they ship.

Measure before you choose

Run the same prompts against each candidate from a machine in your users' region, at realistic concurrency and prompt lengths. A latency chart measured next to the GPU leaves out the network hop your users actually pay for.

Keep the API shape, change the base URL

The request below is the documented Telnyx Inference call to /v2/ai/chat/completions. An existing OpenAI SDK reaches the same endpoint by changing its base URL and credentials.

Shell
curl -i -X POST "https://api.telnyx.com/v2/ai/chat/completions" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "moonshotai/Kimi-K3", "messages": [{"role": "user", "content": "Hello, World!"}] }'

Omitting service_tier uses Default. Adding "service_tier": "priority" or "service_tier": "flex" to the JSON body selects the other tiers with the same moonshotai/Kimi-K3 model ID.

Put inference where the rest of the pipeline already runs

Inference on the same network as speech and telephony removes handoffs for voice and audio workloads. The ai-podcast-producer-python project in the Telnyx code examples repository shows the pattern: it records a multi-host podcast over a conference call, transcribes each speaker with speech-to-text, generates show notes and chapters with inference, and produces text-to-speech intro and outro bumpers. Teams with their own orchestration can adopt the same approach one endpoint at a time.

FAQ

A proprietary model we depend on is being retired in a few weeks. How do we move to different inference infrastructure without rewriting the application?

Move to an OpenAI-compatible provider that serves open-weight models, then change the base URL, credentials, and model name in your existing SDK. Before cutting over, test tool calls, structured output, and context length on the new model, because those vary by model. Telnyx Inference uses the same chat completions request shape, so the application code itself stays the same.

Can we use an inference provider only for the LLM, speech-to-text and text-to-speech pieces and keep our own orchestration?

Yes. A hosted inference API is a set of HTTP endpoints, so it plugs into orchestration you already run. Telnyx runs LLM inference, speech-to-text, and text-to-speech on the same network, and each one can be called on its own. Keeping all three in one region removes vendor handoffs inside each voice turn without replacing your orchestration layer.

We are hitting latency and token-throughput limits on our current cloud provider's endpoint. What actually changes when inference runs on in-region GPUs?

Two things change. In-region GPUs remove cross-region and inter-provider network hops, which often dominate the latency of short interactive replies. Autoscaling on owned capacity removes the need to reserve throughput ahead of demand. Model compute time stays roughly the same, so measure TTFT and P99 from your users' region to see the real difference for your workload.

What is inference in AI, and which infrastructure decisions follow from it?

Inference is running a trained model on new input to produce output, such as generating a reply to a prompt. Because it runs every time a user sends a request, the decisions that follow are about serving: where the GPU sits relative to users, how capacity scales with traffic, what each token costs, and which API shape the application depends on.

What is AI inference versus training, and why do they need different infrastructure?

Training adjusts a model's weights over large datasets in a job that eventually finishes, so it needs maximum compute for a fixed period. Inference uses the finished weights to answer live requests and never finishes while users keep sending them. Training infrastructure is sized for a job. Inference infrastructure is sized for unpredictable demand and judged on latency and cost per token.

Choose a tier and a region per workload

Give each workload the serving tier and in-region GPU it needs instead of renting a cluster for all of them. Move an existing OpenAI SDK app to Telnyx Inference by changing the base URL, with up to 75% savings against proprietary models.

Start building
Share on Social
Eli Mogul
Eli Mogul
Content Writer & Editor

Eli is the content writer and editor at Telnyx. Born and raised in Chicago, Eli attended the University of Missouri where he obtained a BA in Journalism. Eli joined Telnyx in August of 2025. In his spare time, you'll find Eli reading, playing video games, or running.