Inference

vLLM inference: how it works and when to stop self-hosting

Run vLLM inference on one GPU, tune the config flags that matter, and see when moving the same client to a managed endpoint costs less.

vllm Inference feature image

Takeaways

vLLM inference is the process of serving an open-weight LLM through vLLM, an open-source engine that exposes an OpenAI-compatible API. This guide gets one running on a single GPU and shows the one-line change that moves the same client to a managed endpoint.

  • PagedAttention stores the KV cache in 16-token blocks, and continuous batching lets the scheduler mix prefill and decode work in one engine step.
  • One Docker container is the fastest first deployment. Model size and the gpu_memory_utilization fraction decide whether it starts, and --enforce-eager decides how long startup takes.
  • A handful of config flags trade latency for throughput, and the KV cache block formula tells you how many sequences fit on your GPU.
  • When running your own GPUs stops paying, the same OpenAI-compatible client can point at a managed endpoint instead: in-region GPUs, up to 75% less than proprietary alternatives, and no GPUs to manage, through Telnyx Inference.

What vLLM inference is and why it exists

vLLM inference means running an open-weight large language model through the vLLM engine, which schedules requests, manages GPU memory for the KV cache, and returns tokens through an OpenAI-compatible server. You bring the GPUs; vLLM turns them into an endpoint your application can call.

If you want the same open-weight models without owning that hardware, Telnyx Inference runs them on GPUs Telnyx owns, in-region across the Americas, Europe, MENA, and APAC, with no infrastructure to manage.

What is vLLM? vLLM is an open-source inference engine that batches requests, pages the KV cache, and exposes an OpenAI-compatible chat completions API for open-weight LLMs. It is software you run on your own GPUs, not a hosted API.

vLLM is an inference engine, not a model and not a platform

vLLM started as a research project at UC Berkeley's Sky Computing Lab and is now a hosted project of the PyTorch Foundation. Its job is narrow: take a model checkpoint, load it onto accelerators, and serve tokens as efficiently as the hardware allows. Red Hat expands the name as "virtual large language model," though the project itself rarely uses the long form.

Two boundaries trip people up:

  1. vLLM is not an orchestration layer. Projects like llm-d, which Red Hat launched in May 2025 with Google Cloud, IBM Research, NVIDIA, and CoreWeave, sit on top of vLLM to handle Kubernetes scheduling and routing across many instances. Those are separate projects, not vLLM features.
  2. vLLM is not a hosted service. You provision the GPUs, pick the model, size the memory, and keep it running. That operational work is the real cost, and the rest of this guide shows where it lands. For the broader picture of how serving engines fit into a stack, see our guide to the inference engine.

What vLLM inference runs on

vLLM supports text, multimodal, embedding, and reward models across a wide hardware range:

  • NVIDIA GPUs: the primary serving target, with the deepest kernel support.
  • AMD Instinct GPUs: supported through ROCm, with official Docker images.
  • Google TPUs: supported through a dedicated TPU backend.
  • AWS Trainium and Inferentia: supported through the Neuron backend.
  • Intel Gaudi and XPU: supported, with official XPU images.
  • CPUs (x86, ARM): functional for testing, not a production serving target.

vllm hardware families

Try Telnyx Inference to point your existing OpenAI-compatible vLLM client at a managed endpoint with a one-line change.

How vLLM architecture serves a request

Every serving stack, including the GPUs behind a managed API, has to get these mechanisms right. Your users only see them as time to first token and tokens per second, which is why they matter for inference latency.

The engine step: schedule, forward pass, postprocess

vLLM runs a loop. Each iteration, or engine step, does three things:

  1. Schedule: pick which requests run this step and allocate memory for them.
  2. Forward pass: run the model once over every scheduled token and sample the next token for each request.
  3. Postprocess: append the new tokens, detokenize, and check whether each request is finished.

The scheduler holds a waiting queue and a running queue, and serves them first come, first served or by priority. A request finishes when one of four stop conditions fires:

  • It reaches max_model_len or its own max_tokens limit.
  • It samples the end-of-sequence token (unless ignore_eos is set, which is useful for benchmarking fixed output lengths).
  • It samples a token listed in stop_token_ids. That token stays in the output.
  • Its output matches a stop string. The stop string is removed from the output.

These four rules explain most "why did generation stop early" tickets.

vllm engine step

PagedAttention and the KV cache block pool

A transformer keeps a key and value vector for every token it has already processed. That KV cache grows with every output token, and storing it in one contiguous slab per request wastes memory on padding and fragmentation.

PagedAttention stores the KV cache in fixed-size blocks instead, 16 tokens per block by default. At startup, vLLM carves the reserved GPU memory into a pool of free blocks, often hundreds of thousands of them depending on VRAM. Requests borrow blocks as they grow and return them when they finish. When the pool runs dry, the scheduler preempts lower-priority requests and recomputes them later.

The block math: each block, per layer, takes 2 × block_size (default 16) × num_kv_heads × head_size × dtype_bytes, where dtype_bytes is 2 for bf16. A 17-token prompt needs ceil(17/16) = 2 blocks, so the second block sits 15/16 empty until decode fills it.

Paged KV cache is now standard across major serving engines. Where engines differ is scheduling: how they mix work, prioritize requests, and reuse cached blocks. The anatomy of vLLM post from the vLLM team walks through the V1 engine internals in detail.

Prefill, decode, and continuous batching

Inference has two phases with opposite bottlenecks:

  • Prefill runs over every prompt token at once. It is compute-bound, and its cost scales with prompt length. It determines time to first token.
  • Decode runs over one new token per request. It is memory-bandwidth-bound, because the GPU reloads all model weights to produce a single token. It determines tokens per second.

The V1 scheduler mixes both phases in the same step and serves decode requests first, so streaming users keep getting tokens while new prompts are admitted. The older V0 engine could only run one phase per step.

Continuous batching makes this work. vLLM flattens every scheduled sequence into one long sequence, and position indices and attention masks keep each request attending only to its own tokens. New requests join between steps, and nobody waits for the slowest request in a fixed batch.

Prefix caching, chunked prefill, and speculative decoding

Three extensions build on the block pool:

  • Prefix caching hashes each complete 16-token block of a prompt. When a later request shares that prefix, such as a long system prompt, vLLM reuses the cached blocks instead of recomputing them. It is on by default and speeds up prefill only, not decode.
  • Chunked prefill caps how many prompt tokens one request can push through a single step. Set long_prefill_token_threshold to a positive integer and a 50,000-token prompt can no longer block everyone else.
  • Speculative decoding proposes several draft tokens cheaply and verifies them in one forward pass. The V1 engine supports n-gram, EAGLE, and Medusa proposers.

vLLM inference examples: offline, server, and OpenAI-compatible client

Because vLLM and Telnyx Inference both speak the OpenAI chat completions format, the client in the second example is the client in the third, with a different base URL and API key.

Offline vLLM inference example in Python

The smallest complete vLLM program loads a model and generates for a list of prompts. Use this pattern for batch jobs and benchmarks, not for serving users.

Python
from vllm import LLM, SamplingParamsprompts = [ "Explain PagedAttention in one sentence.", "What does a KV cache store?",]sampling_params = SamplingParams(temperature=0.7, max_tokens=128)llm = LLM(model="Qwen/Qwen3-0.6B")outputs = llm.generate(prompts, sampling_params)for output in outputs: print(output.outputs[0].text)

vllm serve and the OpenAI-compatible request

vllm serve starts the same engine behind an OpenAI-compatible HTTP server. Any OpenAI SDK works against it once you point the SDK's base URL at your host.

Shell
vllm serve Qwen/Qwen3-0.6Bcurl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "Qwen/Qwen3-0.6B", "messages": [{"role": "user", "content": "Hello, World!"}] }'

The same request against Telnyx Inference

Here is the same chat completions request sent to Telnyx, exactly as published on the product page:

Shell
curl -i -X POST "https://api.telnyx.com/v2/ai/chat/completions" \ -H "Authorization: Bearer $TELNYX_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "moonshotai/Kimi-K3", "messages": [{"role": "user", "content": "Hello, World!"}] }'

Only three things changed: the base URL, the credential, and the model name. The messages array, the request shape, and your SDK stay the same.

Two engine toggles come up often when self-hosting. To turn off prefix caching, for example when you need fully cold measurements:

Python
llm = LLM(model="Qwen/Qwen3-0.6B", enable_prefix_caching=False)

To try n-gram speculative decoding, which helps most when outputs repeat spans of the prompt, such as code edits or document extraction:

Python
speculative_config = {"method": "ngram", "prompt_lookup_max": 5, "prompt_lookup_min": 3, "num_speculative_tokens": 3}llm = LLM(model="Qwen/Qwen3-0.6B", speculative_config=speculative_config)

Structured output works per request on the vLLM server. On Telnyx Inference, JSON mode and regex constraints do the same job.

vLLM Docker quickstart: one GPU, one container, one endpoint

If the model you want is already in the Telnyx library, your quickstart is the curl above. This section is for when you need your own weights or your own hardware.

Why won't my vLLM container start?

Pick the error or symptom you see in the container logs.

CUDA out of memory while loading weights

Weights alone exceed the GPU memory budget

BF16 weights take about 2 bytes per parameter, so an 8B model needs roughly 16 GB before any KV cache exists. With the default gpu_memory_utilization of 0.9, a 24 GB card gives vLLM about 21.6 GB, which leaves only around 5 GB for cache and activations. A 70B model at about 140 GB cannot fit on a single 80 GB GPU at all.

Switch to a 4-bit AWQ or GPTQ checkpoint, which is roughly a quarter of the BF16 size, or split the model across cards with --tensor-parallel-size.

Before: a BF16 Llama 3.1 70B on one 80 GB H100 fails at load. After: a 4-bit AWQ build of the same model, about 40 GB of weights, starts on that card with room left for cache.

Max seq len is larger than the KV cache can store

Cap context length at what the cache holds

vLLM refuses to start if a single request at full context length cannot fit in the KV cache. Llama 3.1 models default to a 131072-token context, which a 24 GB card running an 8B model cannot hold. The error message prints the exact number of tokens your cache can store.

Set --max-model-len at or below that printed number, or raise --gpu-memory-utilization toward 0.95 if nothing else shares the GPU. Most chat workloads never come close to 131072 tokens, so a lower cap costs you nothing in practice.

Before: default context of 131072 tokens, startup fails. After: --max-model-len 16384, the server starts and serves more sequences at once.

Free memory is less than desired GPU memory utilization

Another process already holds GPU memory

gpu_memory_utilization is a fraction of the card's total memory, not of what is currently free. If a notebook or a second container already holds 6 GB on a 24 GB card, the default 0.9 asks for 21.6 GB that no longer exists.

Run nvidia-smi to find the process and stop it, or lower the fraction so the request fits. Two vLLM instances sharing one GPU each need their own fraction, and the two values together must stay under 1.0.

Before: 6 GB used by Jupyter, vLLM at 0.9 fails. After: --gpu-memory-utilization 0.7 requests 16.8 GB and starts cleanly.

401 or gated repo error while downloading the model

Accept the license and pass a token

Llama, Gemma, and several Mistral checkpoints are gated on Hugging Face. The container has no credentials by default, so the download fails even when the model name is correct.

Accept the license on the model page with your Hugging Face account, then pass your token into the container. Mount the Hugging Face cache directory as well, so the 16 GB of weights for an 8B model download once instead of on every restart.

Before: docker run --gpus all vllm/vllm-openai --model meta-llama/Llama-3.1-8B-Instruct. After: add -e HF_TOKEN=$HF_TOKEN -v ~/.cache/huggingface:/root/.cache/huggingface to the same command.

Stuck for minutes after the weights finish loading

It is capturing CUDA graphs, not hung

After loading, vLLM profiles memory, compiles the model, and captures CUDA graphs for a range of batch sizes. On first run this can take several minutes with no new log lines, which looks like a hang but is not.

--enforce-eager skips graph capture so the endpoint comes up faster, at the cost of slower per-token decode. Use it for development loops and drop it for production. Mounting ~/.cache/vllm keeps the compile cache between restarts, which cuts later startups even with graphs enabled.

Before: every container restart repeats the full compile step. After: -v ~/.cache/vllm:/root/.cache/vllm reuses the cached compile output on the next start.

Pick a model that fits the GPU

Before the container will start, the weights have to fit. In FP16 or BF16, weights take about 2 bytes per parameter. INT4 quantization cuts that to roughly half a byte per parameter, plus some overhead:

  • A 70B model: about 140 GB in FP16, so it needs more than one 80 GB GPU unless quantized.
  • A 72B model: about 144 GB in FP16, or under 48 GB in INT4, which fits one 80 GB card.
  • A 176B model: about 352 GB in FP16, which needs tensor parallelism across several GPUs.

Those numbers cover weights only. The KV cache needs room on top, so leave headroom. Prefer a checkpoint the model vendor has already quantized over quantizing on load, which adds startup time and CPU and GPU work.

Run the vLLM Docker image

The official image, vllm/vllm-openai, runs the OpenAI-compatible server. This is the command from the vLLM Docker documentation, with --gpu-memory-utilization added after the image tag, where the docs say extra engine arguments go:

Shell
docker run --runtime nvidia --gpus all \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=$HF_TOKEN" \ -p 8000:8000 \ --ipc=host \ vllm/vllm-openai:latest \ --model Qwen/Qwen3-0.6B \ --gpu-memory-utilization 0.8

Mounting the Hugging Face cache keeps downloaded weights between container restarts, and --ipc=host gives PyTorch the shared memory it uses between processes. Once the logs show the server is up, the local curl from the previous section works against localhost:8000. For a model that needs several GPUs on one node, add --tensor-parallel-size with the GPU count.

Self-host versus Telnyx GPUs

Why the first request is slow and when it fails to start

At startup, vLLM checks that the GPU has enough free VRAM for the gpu_memory_utilization fraction (0.8 means 80% of total VRAM). It loads the weights, runs a dummy forward pass, and takes a memory snapshot to count how many KV blocks fit. Then it captures CUDA graphs for a range of batch sizes, which cuts kernel launch overhead during serving.

Startup troubleshooting

  • Out of memory at start: the weights plus the reserved fraction exceed the card. Lower --gpu-memory-utilization or pick a smaller quantized checkpoint.
  • Slow first request: that is warmup. Exclude it from any benchmark, and measure TTFT versus end-to-end latency separately.
  • Slow startup: --enforce-eager skips CUDA graph capture. Startup gets faster, and each token gets a little slower.
  • trust_remote_code: this runs code from the model repository with full permissions. Enable it only for sources you trust.

vLLM config: the knobs that change latency and throughput

On Telnyx Inference, these knobs collapse into one request field, service_tier. When you self-host, they are yours to set.

Memory and batch knobs

Every knob here pulls from the same KV block pool that gpu_memory_utilization reserved. More concurrent sequences times longer context means more blocks, and the pool is fixed at startup.

  • gpu_memory_utilization: the fraction of VRAM vLLM reserves for weights and KV cache. Higher means more blocks and more concurrency, with less room for anything else on the card.
  • max_model_len: the longest context a single request can use. Raising it lets fewer full-length requests fit at once.
  • max_num_seqs: the cap on requests running concurrently. Raising it trades per-request latency for total throughput, until the block pool runs out and preemption starts.
  • max_num_batched_tokens: the cap on tokens processed in one engine step. It bounds how much prefill work can crowd into a step alongside decode.
  • --enforce-eager: skips CUDA graph capture. Faster startup, slower tokens.
  • enable_prefix_caching: on by default. Leave it on when requests share long prompts.
  • long_prefill_token_threshold: a positive integer enables chunked prefill, so long prompts stop monopolizing steps.

To estimate capacity, multiply the per-block size from the architecture section by your layer count, divide the reserved VRAM left after weights by that number, and you have the block pool. Divide the pool by max_model_len / 16 to see how many full-context sequences fit.

Attention backend and quantization

vLLM picks an attention kernel to match the model's positional encoding. FlashAttention is the usual fast path for RoPE models such as Qwen, while other backends cover ALiBi models. You can change the kernel at serve time, but not the model's attention mechanism, so check vLLM's attention backend support matrix before forcing one.

For quantization, the rule from the sizing section holds: a pre-quantized checkpoint starts faster and more predictably than quantizing on load.

Parallelism: tensor first, pipeline second, replicas for the rest

When the weights do not fit one GPU:

  1. Tensor parallelism splits each layer across GPUs on the same node. Use it first, because bandwidth inside a node is far higher than between nodes.
  2. Pipeline parallelism splits layers across nodes. Use it when one node still is not enough.
  3. Replicas scale throughput. vLLM's built-in data parallelism keeps replicas in lockstep, running dummy steps on idle replicas, which only mixture-of-experts models require. For dense models, run independent vLLM instances behind a load balancer.

What a managed endpoint replaces these knobs with

On Telnyx Inference, one field on each request selects serving capacity:

  • Default: standard capacity at standard rates. Omit service_tier and this is what you get.
  • Priority: faster serving capacity for interactive work, at a different rate. It does not change the model or the context window.
  • Flex: the lowest rates, for background jobs where minutes of end-to-end latency is acceptable.

vLLM inference vs a managed API: when self-hosting stops paying

Telnyx Inference serves a curated library of open-weight models on GPUs Telnyx owns, behind OpenAI-compatible endpoints you reach by changing the base URL. Pricing is per token, at up to 75% less than proprietary alternatives, with Cached input at up to 88% off and no GPU rental fees, compute surcharges, or minimums.

What you take on when you self-host vLLM inference

Everything earlier in this guide becomes your job:

  • Sizing models against VRAM and reserving the right gpu_memory_utilization fraction.
  • Waiting through the profiling pass and CUDA graph warmup on every new instance.
  • Watching for recompute preemption when the block pool runs out under load.
  • Sharding with tensor parallelism once weights outgrow one card.
  • Provisioning GPUs ahead of every traffic peak, and paying for them while idle.

That last line is usually where the math turns. Our guide to inference cost optimization covers how idle GPU time drives the real cost per token.

What changes when the endpoint is Telnyx

The managed path removes the GPU operations, not the choices that matter. Inference runs in-region across the Americas, Europe, MENA, and APAC, with LATAM coming soon, on Telnyx's GPU network. Data stays in-region with zero data retention and no training on your data. Capacity autoscales from zero to thousands of requests per second with no cold starts.

What you give up is just as clear. You cannot load your own fine-tuned or LoRA weights, and you do not control the hardware.

The decision in one table

Self-hosted vLLM compared with Telnyx Inference

ConcernSelf-hosted vLLMTelnyx Inference
Who runs the GPUsYouTelnyx, on GPUs it owns
ModelsAny weights you can load, including fine-tunesCurated open-weight library from z.ai, DeepSeek, Minimax, and more
ScalingCapacity you provision, warmup per instanceAutoscaling from zero to thousands of requests per second, no cold starts
Region and dataWherever your hardware isIn-region across the Americas, Europe, MENA, and APAC, zero data retention
PricingGPU hours, busy or idlePer token, up to 75% less than proprietary alternatives, no minimums
Latency tuningThe config knobs aboveservice_tier: Default, Priority, or Flex
Client changeNoneChange the base URL and API key

Pricing comparisons refer to the Default tier. Per-token rates vary by model and tier; see inference API pricing for the current lineup.

Inference decision

The migration is the request you already have. Point your OpenAI SDK's base URL at Telnyx, set your Telnyx API key, keep the messages array, and choose a model from the pricing page. Set service_tier to Priority for interactive traffic or Flex for background jobs.

vLLM inference FAQ

What is vLLM?

vLLM is an open-source inference engine for large language models. It schedules requests, manages GPU memory with PagedAttention, and exposes an OpenAI-compatible API server. It started at UC Berkeley and is now a hosted project of the PyTorch Foundation.

How does vLLM's PagedAttention allocate KV cache blocks?

At startup, vLLM reserves a fraction of VRAM and divides it into fixed blocks of 16 tokens each. Each request takes blocks from a shared pool as its context grows and returns them when it finishes. If the pool runs out, the scheduler preempts lower-priority requests and recomputes them later.

How does prefix caching work in vLLM and how do you disable it?

vLLM hashes each complete 16-token block of a prompt. When a new request starts with the same blocks, vLLM reuses the stored KV cache instead of recomputing it, which speeds up prefill. It is on by default; set enable_prefix_caching=False to turn it off.

What parallelism strategies does vLLM support and when should each be used?

vLLM supports tensor, pipeline, data, and expert parallelism. Use tensor parallelism across GPUs on one node first, pipeline parallelism across nodes when one node is not enough, and independent replicas behind a load balancer for extra throughput on dense models. Built-in data parallelism matters mainly for mixture-of-experts models.

What hardware does vLLM run on?

vLLM runs on NVIDIA and AMD GPUs, Google TPUs, AWS Trainium and Inferentia, Intel Gaudi and XPU accelerators, and CPUs. NVIDIA GPUs have the deepest support; CPUs work for testing but are not a serving target.


Diagram: What hardware does vLLM run on?

\Run open-weight models without running GPUs

You now have a vLLM server on one GPU and the one-line change that moves the same client to Telnyx-hosted GPUs. Sign up to test open-weight models on in-region GPUs, at up to 75% less than proprietary alternatives, with no infrastructure to manage. Start building with Telnyx Inference

Move your vLLM client to a managed endpoint

When running your own GPUs stops paying, point the same OpenAI-compatible client at Telnyx and get in-region GPUs at up to 75% less than proprietary alternatives.

Try Telnyx Inference
Share on Social
Eli Mogul
Eli Mogul
Content Writer & Editor

Eli is the content writer and editor at Telnyx. Born and raised in Chicago, Eli attended the University of Missouri where he obtained a BA in Journalism. Eli joined Telnyx in August of 2025. In his spare time, you'll find Eli reading, playing video games, or running.