Conversational AI

Inference latency: where milliseconds go and how to cut them

Inference latency is a systems problem, not just a model problem. The model is one of three layers, and for real-time AI it is rarely the layer with the most latency to give back.

Inference Latency feature

Takeaways

Inference latency is the time between sending a request to a model and receiving a usable response. It comes from four places: the network, the provider's queue, prefill, and decode. Your provider's serverless region coverage decides the network hop.

  • Time-to-first-token (TTFT) and tokens per second measure different things. TTFT is what a chat or voice user waits for; tokens per second is how fast the rest arrives. A provider can win one and lose the other.
  • Benchmarks run from a US client sitting near the GPU measure three of the four components and never see the out-of-region network penalty.
  • Few providers offer serverless inference outside the US, so the region your users live in can matter more than the chip your provider runs.

You picked your inference provider the sensible way. You compared TTFT and tokens-per-second charts, chose the fastest one, and shipped. Then users in Frankfurt or Singapore started reporting replies that felt two or three times slower than the chart promised.

The chart wasn't wrong. It measured a different request than the one your users send. This guide breaks inference latency into its four components, shows how to measure each one from where your users actually are, and explains which levers cut which milliseconds.

What inference latency is and how it is measured

Inference latency is the time between sending a request to a model and receiving a usable response. What counts as usable depends on the product, which is why a single latency number rarely tells the whole story.

Which latency number should you optimize?

Pick the workload closest to yours to see the metric that decides it.

A chat interface where users watch the reply appear

Optimize time-to-first-token first.

TTFT is what a chat user waits for. It runs from the moment the request is sent to the moment the first token arrives, so it covers three of the four components: the network, the provider's queue, and prefill.

Because prefill processes the entire prompt, TTFT grows with prompt length and provider load. Tokens per second matters only after the first token lands.

Before: choosing the provider at the top of the tokens per second chart. After: choosing the provider with the lowest TTFT at your real prompt length.

A voice agent that has to answer a caller

TTFT is the whole budget. Guard it.

A caller hears silence until the first token arrives, so TTFT decides whether the agent feels live or broken. Every component of TTFT counts here: the network hop, the queue wait, and prefill.

The network hop is set by your provider's serverless region coverage, so check where the GPU runs before you compare chips.

Before: a fast GPU in a US region serving callers abroad. After: a serverless region close to the caller so the network hop stops eating the budget.

Long answers or code generation

Optimize tokens per second.

Once output starts streaming, decode throughput decides how long the reader waits for the rest. It depends mostly on model size, the GPU, and how aggressively the provider batches requests.

A provider can win on TTFT and lose on throughput, so compare both numbers rather than one blended score.

Before: picking the provider with the quickest first token for a code assistant. After: picking the provider with the highest decode throughput, since most of the wait is decode.

Batch jobs or tool-calling agents

Optimize end-to-end latency.

End-to-end latency runs from the request to the final token. It equals TTFT plus the time decode needs to produce every remaining token, and NVIDIA's benchmarking metrics note it includes queueing, batching, and network latency.

A tool-calling agent cannot act until the full response is complete, so the last token matters more than the first.

Before: judging an agent pipeline by TTFT alone. After: timing each step from request to last token and summing across the chain.

Users in Frankfurt, Singapore, or anywhere outside the US

Region placement matters more than the chip.

Benchmarks run from a US client sitting near the GPU measure three of the four components and never see the out-of-region network penalty. That is why users in Frankfurt or Singapore can see replies two or three times slower than the chart promised.

Few providers offer serverless inference outside the US, so check region coverage first and measure TTFT from where your users actually are.

Before: trusting a TTFT chart measured from a US client. After: running the same TTFT test from a Frankfurt or Singapore client against each provider's nearest serverless region.

Follow one request in order. It leaves the client and crosses the network to the provider. It waits in a queue until the provider can batch it onto a GPU. The GPU runs prefill, processing the entire prompt and emitting the first token. Decode then generates the remaining tokens one at a time until the response is complete.

Three metrics come out of that timeline, and most comparison pages blur them together.

Inference latency

Time-to-first-token vs end-to-end latency

Time-to-first-token (TTFT) runs from the moment you send the request to the moment the first token arrives. It covers the network, the queue, and prefill, so it grows with prompt length and provider load.

End-to-end latency runs from the request to the final token. It equals TTFT plus the time decode needs to produce every remaining token. NVIDIA's benchmarking metrics define it the same way and note that it includes queueing, batching, and network latency.

Tokens per second and decode throughput

Tokens per second, sometimes reported as decode throughput or its inverse, time per output token, describes how fast output streams once it starts. It depends mostly on model size, the GPU, and how aggressively the provider batches requests.

Because TTFT and throughput depend on different things, you need both numbers to predict how a product will feel. Our TTFT vs E2E benchmark reports them separately for exactly this reason: the metric that matters depends on what you are building.

MetricWhat it measuresWorkload that feels it most
Time-to-first-token (TTFT)Request sent to first token receivedChat interfaces and voice agents
Tokens per secondOutput speed after the first tokenLong answers and code generation
End-to-end latencyRequest sent to last token receivedBatch jobs and tool-calling agents
Explore Telnyx Inference to check how serverless region coverage shapes the network hop your users wait through.

The four places inference latency comes from

Every inference latency number is the sum of four components. Each one scales with something different, and a different party controls it.

Inference latency hops

Network hop and jitter

The request travels from the client to the provider's edge, and the response travels back. This component scales with physical distance and path quality. You don't control it directly; the location of your provider's GPU does.

Distance adds delay, and shared paths add jitter, which is variation in delay from one request to the next. Jitter often hurts more than a stable offset. A voice agent or streaming interface can plan around a consistent 80ms delay, but it can't plan around a delay that swings between 40ms and 300ms.

Queueing and batching on the provider

Providers group requests into batches to keep GPUs busy. Your request waits until a slot opens. Queue time scales with the provider's load and batching policy, which is why the same model on the same provider can be fast at 3 a.m. and slow at noon.

Prefill and decode on the GPU

During prefill, the GPU processes every token in your prompt before it can emit anything. Prefill time scales with prompt length, so a 20,000-token context costs far more TTFT than a 500-token one.

During decode, the model produces output tokens one at a time. Decode time scales with output length and model size.

Here's how to attribute your own measured latency:

  1. Network: scales with distance and path quality. Controlled by your provider's region footprint.
  2. Queueing: scales with provider load and batch policy. Controlled by the provider and your capacity tier.
  3. Prefill: scales with prompt tokens. Controlled by you.
  4. Decode: scales with output tokens and model size. Controlled by you and your model choice.

Most published comparisons start measuring at step two. A benchmark with no stated client location is a benchmark of three components out of four.

Inference latency benchmarks: TTFT and decode throughput across providers

Provider comparisons tend to group the market into bands. Groq and Cerebras lead on speed with custom silicon. Together AI and Fireworks AI compete on price and model breadth with optimized GPU clouds. These bands are useful shorthand, but they describe hardware, not the experience of a user 6,000 kilometers from the data center.

How to read a provider latency chart

Before you trust any latency chart, check four things in this order:

  1. Model and quantization. Two providers serving the "same" model at different precisions aren't comparable.
  2. Input and output token counts. A 100-token prompt flatters TTFT; a 50,000-token prompt exposes prefill.
  3. Client location. If the chart doesn't say where the requests came from, assume they came from next to the GPU.
  4. p50 or p95. The median tells you the typical experience. The 95th percentile tells you what your unhappiest users feel, and it's where jitter shows up.

A chart missing any of these can't be compared to your production traffic. Client location is the one most often left out.

Telnyx published benchmarks by model

We publish our own per-model benchmarks with the methodology attached, and we report TTFT and end-to-end latency separately. The results show why that split matters. In our GLM-5.3 latency benchmarks, a competitor led on time to first answer while Telnyx posted the lowest completed-response times in most workloads. Neither result is the whole picture on its own.

Our MiniMax M3 benchmark runs six prompt scenarios, from 1,000 to 100,000 input tokens, and reports end-to-end latency, TTFT, and throughput for each. Use these articles for current per-model numbers rather than any figure quoted secondhand, including in this guide.

Why region placement decides inference latency for users outside the US

The network component is the one most benchmarks skip, and it's the one your provider's serverless footprint decides.

A user in Frankfurt calling a serverless GPU in Virginia pays the transatlantic round trip on every request. No amount of GPU optimization removes it, because the time is spent in fiber, not on silicon. For a chat product, that delay lands directly on TTFT. For a voice agent, it lands on every turn of the conversation.

Serverless in-region: who actually offers it

Most providers solve the distance problem with dedicated deployments: reserved capacity in a chosen region, usually on a different pricing model. That works, but it means the serverless tier most teams start on stays in the US.

Telnyx serverless map

GPU placement versus network path

That last sentence of the quote carries two arguments. The first is latency: a GPU in the user's region shortens the physical path, which cuts both delay and jitter. The second is data residency. A residency setting on a provider that routes traffic to another continent is a policy promise. A GPU that physically sits in the region is an architectural fact.

The Telnyx GPU network places compute alongside our global points of presence, so adding a region adds capacity close to users instead of adding another network hop. If most of your users live outside the US, test your provider from their region before you compare chips.

How to reduce inference latency in production

Work through the levers in the same order the request meets them. That way you fix the component that is actually slow instead of tuning decode when the problem is the network.

Network: Serve users from a serverless GPU in their region. This is a provider decision, and it's the only lever that removes distance.

Queueing: If your traffic is predictable, dedicated capacity removes the wait for a batch slot. The trade-off is cost and commitment.

Prefill: Shorter prompts reach the first token sooner. Trim retrieved context to what the model needs, and keep static instructions in a stable system prompt so providers that support prompt caching can reuse them.

Decode: Stream tokens so users see output at TTFT instead of end-to-end latency. Cap max_tokens for replies that should be short. Use the smallest model that passes your quality bar, since smaller models generate each token faster.

Several of these levers also lower cost, because fewer prompt and output tokens mean a smaller bill.

Measure inference latency from your users' region

The most useful benchmark is the one you run from where your users are. The script below sends streamed chat completions through the OpenAI-compatible interface, times the first content chunk, and reports p50 and p95 across 20 runs. Existing OpenAI SDK code only needs a new base URL and API credential.

Before you run it:

  • Install the SDK with pip install openai (Python 3.8 or later).
  • Create an API credential in the Telnyx Mission Control Portal and export it as TELNYX_API_KEY.
  • Run the script from a machine in your users' region, such as a small cloud instance in Frankfurt or Singapore. A laptop at your desk measures your network path, not theirs.
Python
import osimport statisticsimport timefrom openai import OpenAIclient = OpenAI( api_key=os.environ["TELNYX_API_KEY"], base_url="https://api.telnyx.com/v2/ai/openai",)MODEL = os.environ.get("MODEL", "meta-llama/Meta-Llama-3.1-8B-Instruct")PROMPT = "Explain what a SIP trunk is in three sentences."RUNS = 20def measure_once(): start = time.perf_counter() first_token_at = None chunks = 0 stream = client.chat.completions.create( model=MODEL, messages=[{"role": "user", "content": PROMPT}], max_tokens=256, stream=True, ) for chunk in stream: if not chunk.choices: continue if chunk.choices[0].delta.content: if first_token_at is None: first_token_at = time.perf_counter() chunks += 1 end = time.perf_counter() ttft = first_token_at - start decode_time = end - first_token_at # Content chunks approximate output tokens. tokens_per_sec = (chunks - 1) / decode_time if chunks > 1 else 0.0 return ttft, end - start, tokens_per_secdef p95(values): return statistics.quantiles(values, n=20)[18]results = [measure_once() for _ in range(RUNS)]ttfts, e2es, rates = zip(*results)print(f"model={MODEL} runs={RUNS}")print(f"TTFT p50={statistics.median(ttfts):.3f}s p95={p95(ttfts):.3f}s")print(f"E2E p50={statistics.median(e2es):.3f}s p95={p95(e2es):.3f}s")print(f"tokens/s p50={statistics.median(rates):.1f}")

Pick the smallest model that passes, and switch at runtime

Model size is the biggest decode lever you control. Because the script reads the model from an environment variable, you can run it against two models back to back and compare TTFT and throughput on your real prompts. In production, the same pattern applies: keep the model ID in configuration rather than code, so you can route simple requests to a smaller, faster model without a redeploy.

Cut prompt tokens and reuse context

Prefill cost grows with every token you send. Audit what goes into each request. Conversation history, retrieved documents, and tool schemas tend to grow over time, and every one of those tokens is processed before the first output token appears.

Inference latency for voice agents: the budget that decides the product

Voice is where inference latency stops being a metric and becomes the product. The caller hears every millisecond as silence.

Where the LLM sits in the STT, LLM, TTS chain

A voice turn runs in order. The caller speaks, speech-to-text (STT) transcribes, the LLM produces a reply, text-to-speech (TTS) synthesizes it, and audio returns to the caller. TTFT is the gate for the LLM step, because TTS can start speaking the first sentence while the model is still generating the rest. Throughput matters less here than how fast the first words arrive.

People notice. Our guide to voice AI latency notes that callers expect a response within about 500 milliseconds and that anything over a second feels awkward.

Why one network beats three vendors

When STT, the LLM, and TTS run on three different vendors, every handoff crosses a vendor boundary. Each crossing adds a network round trip and its own jitter, so a three-vendor stack pays the network component three times per turn.

Vendor stack comparison

Running the whole chain on co-located infrastructure removes those handoffs. Regis David Souza sees this play out with customers: teams land on Telnyx inference for the models and the economics, then realize the same infrastructure already runs STT, TTS, and telephony. As he puts it, "The integration cost to go from text inference to voice agents is zero because it's the same network."

FAQ

What is inference latency?

Inference latency is the time between sending a request to an AI model and receiving a usable response. It includes the network trip to the provider, time waiting in the provider's queue, prefill (processing the prompt), and decode (generating output tokens).

How can you optimize inference latency?

Start by finding which component is slow. Serve users from a GPU in their region to cut network time, use dedicated capacity to avoid queueing, shorten prompts to speed up prefill, and stream output, cap output length, or use a smaller model to cut decode time.

What is inference latency in real-time computer vision?

It's the time from capturing a frame to getting the model's prediction. The same four components apply, but the budget is set by frame rate. At 30 frames per second, each frame has about 33 milliseconds before the next arrives, which is why real-time vision often runs on hardware near the camera.

Is inference latency the same as time-to-first-token?

No. Time-to-first-token is one measure of inference latency. It stops at the first token. End-to-end latency continues until the final token, so it also includes decode time for the full response.

What is a good inference latency for a real-time AI agent?

It depends on the interface. For voice agents, callers expect a reply within about 500 milliseconds of finishing speaking, and anything over a second feels awkward. That budget covers the whole turn, including STT and TTS, so the LLM's share is only part of it.

Find the hop that's eating your latency budget

You can now split any inference latency number into network, queue, prefill, and decode. The network hop is the one you fix by running on a serverless GPU in your users' region. Run the script above from where your users are, point it at Telnyx Inference, and compare the TTFT you see with the chart you picked your provider from.

Try Telnyx Inference
Share on Social
Eli Mogul
Eli Mogul
Content Writer & Editor

Eli is the content writer and editor at Telnyx. Born and raised in Chicago, Eli attended the University of Missouri where he obtained a BA in Journalism. Eli joined Telnyx in August of 2025. In his spare time, you'll find Eli reading, playing video games, or running.