Inference latency is a systems problem, not just a model problem. The model is one of three layers, and for real-time AI it is rarely the layer with the most latency to give back.

Takeaways
Inference latency is the time between sending a request to a model and receiving a usable response. It comes from four places: the network, the provider's queue, prefill, and decode. Your provider's serverless region coverage decides the network hop.
You picked your inference provider the sensible way. You compared TTFT and tokens-per-second charts, chose the fastest one, and shipped. Then users in Frankfurt or Singapore started reporting replies that felt two or three times slower than the chart promised.
The chart wasn't wrong. It measured a different request than the one your users send. This guide breaks inference latency into its four components, shows how to measure each one from where your users actually are, and explains which levers cut which milliseconds.
Inference latency is the time between sending a request to a model and receiving a usable response. What counts as usable depends on the product, which is why a single latency number rarely tells the whole story.
Which latency number should you optimize?
Pick the workload closest to yours to see the metric that decides it.
Optimize time-to-first-token first.
TTFT is what a chat user waits for. It runs from the moment the request is sent to the moment the first token arrives, so it covers three of the four components: the network, the provider's queue, and prefill.
Because prefill processes the entire prompt, TTFT grows with prompt length and provider load. Tokens per second matters only after the first token lands.
Before: choosing the provider at the top of the tokens per second chart. After: choosing the provider with the lowest TTFT at your real prompt length.
TTFT is the whole budget. Guard it.
A caller hears silence until the first token arrives, so TTFT decides whether the agent feels live or broken. Every component of TTFT counts here: the network hop, the queue wait, and prefill.
The network hop is set by your provider's serverless region coverage, so check where the GPU runs before you compare chips.
Before: a fast GPU in a US region serving callers abroad. After: a serverless region close to the caller so the network hop stops eating the budget.
Optimize tokens per second.
Once output starts streaming, decode throughput decides how long the reader waits for the rest. It depends mostly on model size, the GPU, and how aggressively the provider batches requests.
A provider can win on TTFT and lose on throughput, so compare both numbers rather than one blended score.
Before: picking the provider with the quickest first token for a code assistant. After: picking the provider with the highest decode throughput, since most of the wait is decode.
Optimize end-to-end latency.
End-to-end latency runs from the request to the final token. It equals TTFT plus the time decode needs to produce every remaining token, and NVIDIA's benchmarking metrics note it includes queueing, batching, and network latency.
A tool-calling agent cannot act until the full response is complete, so the last token matters more than the first.
Before: judging an agent pipeline by TTFT alone. After: timing each step from request to last token and summing across the chain.
Region placement matters more than the chip.
Benchmarks run from a US client sitting near the GPU measure three of the four components and never see the out-of-region network penalty. That is why users in Frankfurt or Singapore can see replies two or three times slower than the chart promised.
Few providers offer serverless inference outside the US, so check region coverage first and measure TTFT from where your users actually are.
Before: trusting a TTFT chart measured from a US client. After: running the same TTFT test from a Frankfurt or Singapore client against each provider's nearest serverless region.
Follow one request in order. It leaves the client and crosses the network to the provider. It waits in a queue until the provider can batch it onto a GPU. The GPU runs prefill, processing the entire prompt and emitting the first token. Decode then generates the remaining tokens one at a time until the response is complete.
Three metrics come out of that timeline, and most comparison pages blur them together.
Time-to-first-token (TTFT) runs from the moment you send the request to the moment the first token arrives. It covers the network, the queue, and prefill, so it grows with prompt length and provider load.
End-to-end latency runs from the request to the final token. It equals TTFT plus the time decode needs to produce every remaining token. NVIDIA's benchmarking metrics define it the same way and note that it includes queueing, batching, and network latency.
Tokens per second, sometimes reported as decode throughput or its inverse, time per output token, describes how fast output streams once it starts. It depends mostly on model size, the GPU, and how aggressively the provider batches requests.
Because TTFT and throughput depend on different things, you need both numbers to predict how a product will feel. Our TTFT vs E2E benchmark reports them separately for exactly this reason: the metric that matters depends on what you are building.
| Metric | What it measures | Workload that feels it most |
|---|---|---|
| Time-to-first-token (TTFT) | Request sent to first token received | Chat interfaces and voice agents |
| Tokens per second | Output speed after the first token | Long answers and code generation |
| End-to-end latency | Request sent to last token received | Batch jobs and tool-calling agents |
Throughput isn't responsiveness. A provider that streams 400 tokens per second can still feel slower in a chat window than one streaming 200, if its TTFT is twice as long. The user waits for the first token, not the last. |
Every inference latency number is the sum of four components. Each one scales with something different, and a different party controls it.
The request travels from the client to the provider's edge, and the response travels back. This component scales with physical distance and path quality. You don't control it directly; the location of your provider's GPU does.
Distance adds delay, and shared paths add jitter, which is variation in delay from one request to the next. Jitter often hurts more than a stable offset. A voice agent or streaming interface can plan around a consistent 80ms delay, but it can't plan around a delay that swings between 40ms and 300ms.
"When you are actually using an inference or a TTS model that is hosted in a third-party vendor's servers, then they typically use a public cloud network. With that, that brings a lot of latency and jitter." Abhishek Sharma, Senior Technical Marketing Manager at Telnyx |
Providers group requests into batches to keep GPUs busy. Your request waits until a slot opens. Queue time scales with the provider's load and batching policy, which is why the same model on the same provider can be fast at 3 a.m. and slow at noon.
During prefill, the GPU processes every token in your prompt before it can emit anything. Prefill time scales with prompt length, so a 20,000-token context costs far more TTFT than a 500-token one.
During decode, the model produces output tokens one at a time. Decode time scales with output length and model size.
Here's how to attribute your own measured latency:
Most published comparisons start measuring at step two. A benchmark with no stated client location is a benchmark of three components out of four.
Provider comparisons tend to group the market into bands. Groq and Cerebras lead on speed with custom silicon. Together AI and Fireworks AI compete on price and model breadth with optimized GPU clouds. These bands are useful shorthand, but they describe hardware, not the experience of a user 6,000 kilometers from the data center.
Before you trust any latency chart, check four things in this order:
A chart missing any of these can't be compared to your production traffic. Client location is the one most often left out.
We publish our own per-model benchmarks with the methodology attached, and we report TTFT and end-to-end latency separately. The results show why that split matters. In our GLM-5.3 latency benchmarks, a competitor led on time to first answer while Telnyx posted the lowest completed-response times in most workloads. Neither result is the whole picture on its own.
Our MiniMax M3 benchmark runs six prompt scenarios, from 1,000 to 100,000 input tokens, and reports end-to-end latency, TTFT, and throughput for each. Use these articles for current per-model numbers rather than any figure quoted secondhand, including in this guide.
The network component is the one most benchmarks skip, and it's the one your provider's serverless footprint decides.
A user in Frankfurt calling a serverless GPU in Virginia pays the transatlantic round trip on every request. No amount of GPU optimization removes it, because the time is spent in fiber, not on silicon. For a chat product, that delay lands directly on TTFT. For a voice agent, it lands on every turn of the conversation.
Most providers solve the distance problem with dedicated deployments: reserved capacity in a chosen region, usually on a different pricing model. That works, but it means the serverless tier most teams start on stays in the US.
"It's both, but the dirty secret is most competitors don't offer serverless in-region at all. Together is US-concentrated with no advertised EU or APAC serverless. Fireworks has 15-plus regions but dedicated deployments only outside the US. We run serverless inference in the Americas, Europe, and APAC today, with Dubai and São Paulo launching next. Data stays where your users are because our GPUs are physically there, not because you toggled a setting." Regis David Souza, Staff Software Engineer at Telnyx |
That last sentence of the quote carries two arguments. The first is latency: a GPU in the user's region shortens the physical path, which cuts both delay and jitter. The second is data residency. A residency setting on a provider that routes traffic to another continent is a policy promise. A GPU that physically sits in the region is an architectural fact.
The Telnyx GPU network places compute alongside our global points of presence, so adding a region adds capacity close to users instead of adding another network hop. If most of your users live outside the US, test your provider from their region before you compare chips.
Work through the levers in the same order the request meets them. That way you fix the component that is actually slow instead of tuning decode when the problem is the network.
Network: Serve users from a serverless GPU in their region. This is a provider decision, and it's the only lever that removes distance.
Queueing: If your traffic is predictable, dedicated capacity removes the wait for a batch slot. The trade-off is cost and commitment.
Prefill: Shorter prompts reach the first token sooner. Trim retrieved context to what the model needs, and keep static instructions in a stable system prompt so providers that support prompt caching can reuse them.
Decode: Stream tokens so users see output at TTFT instead of end-to-end latency. Cap max_tokens for replies that should be short. Use the smallest model that passes your quality bar, since smaller models generate each token faster.
Several of these levers also lower cost, because fewer prompt and output tokens mean a smaller bill.
The most useful benchmark is the one you run from where your users are. The script below sends streamed chat completions through the OpenAI-compatible interface, times the first content chunk, and reports p50 and p95 across 20 runs. Existing OpenAI SDK code only needs a new base URL and API credential.
Before you run it:
pip install openai (Python 3.8 or later).TELNYX_API_KEY.Record these with every run. Model name, input and output token counts, the region the request left from, TTFT, tokens per second after the first token, and p50 and p95 across repeated runs. Without the client region, your own measurement has the same blind spot as the charts you are trying to check. |
Model size is the biggest decode lever you control. Because the script reads the model from an environment variable, you can run it against two models back to back and compare TTFT and throughput on your real prompts. In production, the same pattern applies: keep the model ID in configuration rather than code, so you can route simple requests to a smaller, faster model without a redeploy.
Prefill cost grows with every token you send. Audit what goes into each request. Conversation history, retrieved documents, and tool schemas tend to grow over time, and every one of those tokens is processed before the first output token appears.
Voice is where inference latency stops being a metric and becomes the product. The caller hears every millisecond as silence.
A voice turn runs in order. The caller speaks, speech-to-text (STT) transcribes, the LLM produces a reply, text-to-speech (TTS) synthesizes it, and audio returns to the caller. TTFT is the gate for the LLM step, because TTS can start speaking the first sentence while the model is still generating the rest. Throughput matters less here than how fast the first words arrive.
People notice. Our guide to voice AI latency notes that callers expect a response within about 500 milliseconds and that anything over a second feels awkward.
"In the agent era, latency, reliability and compliance aren't nice-to-haves, they are the product. If your intelligent system lags 800ms, drops context when a vendor hiccups or fails compliance checks, you're already out." Ian Reither, COO at Telnyx |
When STT, the LLM, and TTS run on three different vendors, every handoff crosses a vendor boundary. Each crossing adds a network round trip and its own jitter, so a three-vendor stack pays the network component three times per turn.
Running the whole chain on co-located infrastructure removes those handoffs. Regis David Souza sees this play out with customers: teams land on Telnyx inference for the models and the economics, then realize the same infrastructure already runs STT, TTS, and telephony. As he puts it, "The integration cost to go from text inference to voice agents is zero because it's the same network."
FAQ
Inference latency is the time between sending a request to an AI model and receiving a usable response. It includes the network trip to the provider, time waiting in the provider's queue, prefill (processing the prompt), and decode (generating output tokens).
Start by finding which component is slow. Serve users from a GPU in their region to cut network time, use dedicated capacity to avoid queueing, shorten prompts to speed up prefill, and stream output, cap output length, or use a smaller model to cut decode time.
It's the time from capturing a frame to getting the model's prediction. The same four components apply, but the budget is set by frame rate. At 30 frames per second, each frame has about 33 milliseconds before the next arrives, which is why real-time vision often runs on hardware near the camera.
No. Time-to-first-token is one measure of inference latency. It stops at the first token. End-to-end latency continues until the final token, so it also includes decode time for the full response.
It depends on the interface. For voice agents, callers expect a reply within about 500 milliseconds of finishing speaking, and anything over a second feels awkward. That budget covers the whole turn, including STT and TTS, so the LLM's share is only part of it.
You can now split any inference latency number into network, queue, prefill, and decode. The network hop is the one you fix by running on a serverless GPU in your users' region. Run the script above from where your users are, point it at Telnyx Inference, and compare the TTFT you see with the chart you picked your provider from.
Try Telnyx InferenceRelated articles
Why your voice AI agent is slow on phone calls, and how to fix it
The 11 best conversational AI platforms in 2026
How to evaluate Voice AI the way real conversations work

NVIDIA B300 pricing in 2026: per hour, per system and per token

We tested 3 open-weight models for multi-turn agents

Inference Infrastructure: What to Run Where and Why It Matters
