Run vLLM inference on one GPU, tune the config flags that matter, and see when moving the same client to a managed endpoint costs less.

Takeaways
vLLM inference is the process of serving an open-weight LLM through vLLM, an open-source engine that exposes an OpenAI-compatible API. This guide gets one running on a single GPU and shows the one-line change that moves the same client to a managed endpoint.
gpu_memory_utilization fraction decide whether it starts, and --enforce-eager decides how long startup takes.vLLM inference means running an open-weight large language model through the vLLM engine, which schedules requests, manages GPU memory for the KV cache, and returns tokens through an OpenAI-compatible server. You bring the GPUs; vLLM turns them into an endpoint your application can call.
If you want the same open-weight models without owning that hardware, Telnyx Inference runs them on GPUs Telnyx owns, in-region across the Americas, Europe, MENA, and APAC, with no infrastructure to manage.
vLLM started as a research project at UC Berkeley's Sky Computing Lab and is now a hosted project of the PyTorch Foundation. Its job is narrow: take a model checkpoint, load it onto accelerators, and serve tokens as efficiently as the hardware allows. Red Hat expands the name as "virtual large language model," though the project itself rarely uses the long form.
Two boundaries trip people up:
vLLM supports text, multimodal, embedding, and reward models across a wide hardware range:
Every serving stack, including the GPUs behind a managed API, has to get these mechanisms right. Your users only see them as time to first token and tokens per second, which is why they matter for inference latency.
vLLM runs a loop. Each iteration, or engine step, does three things:
The scheduler holds a waiting queue and a running queue, and serves them first come, first served or by priority. A request finishes when one of four stop conditions fires:
max_model_len or its own max_tokens limit.ignore_eos is set, which is useful for benchmarking fixed output lengths).stop_token_ids. That token stays in the output.These four rules explain most "why did generation stop early" tickets.
A transformer keeps a key and value vector for every token it has already processed. That KV cache grows with every output token, and storing it in one contiguous slab per request wastes memory on padding and fragmentation.
PagedAttention stores the KV cache in fixed-size blocks instead, 16 tokens per block by default. At startup, vLLM carves the reserved GPU memory into a pool of free blocks, often hundreds of thousands of them depending on VRAM. Requests borrow blocks as they grow and return them when they finish. When the pool runs dry, the scheduler preempts lower-priority requests and recomputes them later.
Paged KV cache is now standard across major serving engines. Where engines differ is scheduling: how they mix work, prioritize requests, and reuse cached blocks. The anatomy of vLLM post from the vLLM team walks through the V1 engine internals in detail.
Inference has two phases with opposite bottlenecks:
The V1 scheduler mixes both phases in the same step and serves decode requests first, so streaming users keep getting tokens while new prompts are admitted. The older V0 engine could only run one phase per step.
Continuous batching makes this work. vLLM flattens every scheduled sequence into one long sequence, and position indices and attention masks keep each request attending only to its own tokens. New requests join between steps, and nobody waits for the slowest request in a fixed batch.
Three extensions build on the block pool:
long_prefill_token_threshold to a positive integer and a 50,000-token prompt can no longer block everyone else.Because vLLM and Telnyx Inference both speak the OpenAI chat completions format, the client in the second example is the client in the third, with a different base URL and API key.
The smallest complete vLLM program loads a model and generates for a list of prompts. Use this pattern for batch jobs and benchmarks, not for serving users.
vllm serve starts the same engine behind an OpenAI-compatible HTTP server. Any OpenAI SDK works against it once you point the SDK's base URL at your host.
Here is the same chat completions request sent to Telnyx, exactly as published on the product page:
Only three things changed: the base URL, the credential, and the model name. The messages array, the request shape, and your SDK stay the same.
Two engine toggles come up often when self-hosting. To turn off prefix caching, for example when you need fully cold measurements:
To try n-gram speculative decoding, which helps most when outputs repeat spans of the prompt, such as code edits or document extraction:
Structured output works per request on the vLLM server. On Telnyx Inference, JSON mode and regex constraints do the same job.
If the model you want is already in the Telnyx library, your quickstart is the curl above. This section is for when you need your own weights or your own hardware.
Why won't my vLLM container start?
Pick the error or symptom you see in the container logs.
Weights alone exceed the GPU memory budget
BF16 weights take about 2 bytes per parameter, so an 8B model needs roughly 16 GB before any KV cache exists. With the default gpu_memory_utilization of 0.9, a 24 GB card gives vLLM about 21.6 GB, which leaves only around 5 GB for cache and activations. A 70B model at about 140 GB cannot fit on a single 80 GB GPU at all.
Switch to a 4-bit AWQ or GPTQ checkpoint, which is roughly a quarter of the BF16 size, or split the model across cards with --tensor-parallel-size.
Before: a BF16 Llama 3.1 70B on one 80 GB H100 fails at load. After: a 4-bit AWQ build of the same model, about 40 GB of weights, starts on that card with room left for cache.
Cap context length at what the cache holds
vLLM refuses to start if a single request at full context length cannot fit in the KV cache. Llama 3.1 models default to a 131072-token context, which a 24 GB card running an 8B model cannot hold. The error message prints the exact number of tokens your cache can store.
Set --max-model-len at or below that printed number, or raise --gpu-memory-utilization toward 0.95 if nothing else shares the GPU. Most chat workloads never come close to 131072 tokens, so a lower cap costs you nothing in practice.
Before: default context of 131072 tokens, startup fails. After: --max-model-len 16384, the server starts and serves more sequences at once.
Another process already holds GPU memory
gpu_memory_utilization is a fraction of the card's total memory, not of what is currently free. If a notebook or a second container already holds 6 GB on a 24 GB card, the default 0.9 asks for 21.6 GB that no longer exists.
Run nvidia-smi to find the process and stop it, or lower the fraction so the request fits. Two vLLM instances sharing one GPU each need their own fraction, and the two values together must stay under 1.0.
Before: 6 GB used by Jupyter, vLLM at 0.9 fails. After: --gpu-memory-utilization 0.7 requests 16.8 GB and starts cleanly.
Accept the license and pass a token
Llama, Gemma, and several Mistral checkpoints are gated on Hugging Face. The container has no credentials by default, so the download fails even when the model name is correct.
Accept the license on the model page with your Hugging Face account, then pass your token into the container. Mount the Hugging Face cache directory as well, so the 16 GB of weights for an 8B model download once instead of on every restart.
Before: docker run --gpus all vllm/vllm-openai --model meta-llama/Llama-3.1-8B-Instruct. After: add -e HF_TOKEN=$HF_TOKEN -v ~/.cache/huggingface:/root/.cache/huggingface to the same command.
It is capturing CUDA graphs, not hung
After loading, vLLM profiles memory, compiles the model, and captures CUDA graphs for a range of batch sizes. On first run this can take several minutes with no new log lines, which looks like a hang but is not.
--enforce-eager skips graph capture so the endpoint comes up faster, at the cost of slower per-token decode. Use it for development loops and drop it for production. Mounting ~/.cache/vllm keeps the compile cache between restarts, which cuts later startups even with graphs enabled.
Before: every container restart repeats the full compile step. After: -v ~/.cache/vllm:/root/.cache/vllm reuses the cached compile output on the next start.
Before the container will start, the weights have to fit. In FP16 or BF16, weights take about 2 bytes per parameter. INT4 quantization cuts that to roughly half a byte per parameter, plus some overhead:
Those numbers cover weights only. The KV cache needs room on top, so leave headroom. Prefer a checkpoint the model vendor has already quantized over quantizing on load, which adds startup time and CPU and GPU work.
The official image, vllm/vllm-openai, runs the OpenAI-compatible server. This is the command from the vLLM Docker documentation, with --gpu-memory-utilization added after the image tag, where the docs say extra engine arguments go:
Mounting the Hugging Face cache keeps downloaded weights between container restarts, and --ipc=host gives PyTorch the shared memory it uses between processes. Once the logs show the server is up, the local curl from the previous section works against localhost:8000. For a model that needs several GPUs on one node, add --tensor-parallel-size with the GPU count.
At startup, vLLM checks that the GPU has enough free VRAM for the gpu_memory_utilization fraction (0.8 means 80% of total VRAM). It loads the weights, runs a dummy forward pass, and takes a memory snapshot to count how many KV blocks fit. Then it captures CUDA graphs for a range of batch sizes, which cuts kernel launch overhead during serving.
Startup troubleshooting
--gpu-memory-utilization or pick a smaller quantized checkpoint.--enforce-eager skips CUDA graph capture. Startup gets faster, and each token gets a little slower.trust_remote_code: this runs code from the model repository with full permissions. Enable it only for sources you trust.On Telnyx Inference, these knobs collapse into one request field, service_tier. When you self-host, they are yours to set.
Every knob here pulls from the same KV block pool that gpu_memory_utilization reserved. More concurrent sequences times longer context means more blocks, and the pool is fixed at startup.
gpu_memory_utilization: the fraction of VRAM vLLM reserves for weights and KV cache. Higher means more blocks and more concurrency, with less room for anything else on the card.max_model_len: the longest context a single request can use. Raising it lets fewer full-length requests fit at once.max_num_seqs: the cap on requests running concurrently. Raising it trades per-request latency for total throughput, until the block pool runs out and preemption starts.max_num_batched_tokens: the cap on tokens processed in one engine step. It bounds how much prefill work can crowd into a step alongside decode.--enforce-eager: skips CUDA graph capture. Faster startup, slower tokens.enable_prefix_caching: on by default. Leave it on when requests share long prompts.long_prefill_token_threshold: a positive integer enables chunked prefill, so long prompts stop monopolizing steps.To estimate capacity, multiply the per-block size from the architecture section by your layer count, divide the reserved VRAM left after weights by that number, and you have the block pool. Divide the pool by max_model_len / 16 to see how many full-context sequences fit.
vLLM picks an attention kernel to match the model's positional encoding. FlashAttention is the usual fast path for RoPE models such as Qwen, while other backends cover ALiBi models. You can change the kernel at serve time, but not the model's attention mechanism, so check vLLM's attention backend support matrix before forcing one.
For quantization, the rule from the sizing section holds: a pre-quantized checkpoint starts faster and more predictably than quantizing on load.
When the weights do not fit one GPU:
On Telnyx Inference, one field on each request selects serving capacity:
service_tier and this is what you get.Telnyx Inference serves a curated library of open-weight models on GPUs Telnyx owns, behind OpenAI-compatible endpoints you reach by changing the base URL. Pricing is per token, at up to 75% less than proprietary alternatives, with Cached input at up to 88% off and no GPU rental fees, compute surcharges, or minimums.
Everything earlier in this guide becomes your job:
gpu_memory_utilization fraction.That last line is usually where the math turns. Our guide to inference cost optimization covers how idle GPU time drives the real cost per token.
The managed path removes the GPU operations, not the choices that matter. Inference runs in-region across the Americas, Europe, MENA, and APAC, with LATAM coming soon, on Telnyx's GPU network. Data stays in-region with zero data retention and no training on your data. Capacity autoscales from zero to thousands of requests per second with no cold starts.
What you give up is just as clear. You cannot load your own fine-tuned or LoRA weights, and you do not control the hardware.
Self-hosted vLLM compared with Telnyx Inference
| Concern | Self-hosted vLLM | Telnyx Inference |
|---|---|---|
| Who runs the GPUs | You | Telnyx, on GPUs it owns |
| Models | Any weights you can load, including fine-tunes | Curated open-weight library from z.ai, DeepSeek, Minimax, and more |
| Scaling | Capacity you provision, warmup per instance | Autoscaling from zero to thousands of requests per second, no cold starts |
| Region and data | Wherever your hardware is | In-region across the Americas, Europe, MENA, and APAC, zero data retention |
| Pricing | GPU hours, busy or idle | Per token, up to 75% less than proprietary alternatives, no minimums |
| Latency tuning | The config knobs above | service_tier: Default, Priority, or Flex |
| Client change | None | Change the base URL and API key |
Pricing comparisons refer to the Default tier. Per-token rates vary by model and tier; see inference API pricing for the current lineup.
service_tier to Priority for interactive traffic or Flex for background jobs. vLLM is an open-source inference engine for large language models. It schedules requests, manages GPU memory with PagedAttention, and exposes an OpenAI-compatible API server. It started at UC Berkeley and is now a hosted project of the PyTorch Foundation.
At startup, vLLM reserves a fraction of VRAM and divides it into fixed blocks of 16 tokens each. Each request takes blocks from a shared pool as its context grows and returns them when it finishes. If the pool runs out, the scheduler preempts lower-priority requests and recomputes them later.
vLLM hashes each complete 16-token block of a prompt. When a new request starts with the same blocks, vLLM reuses the stored KV cache instead of recomputing it, which speeds up prefill. It is on by default; set enable_prefix_caching=False to turn it off.
vLLM supports tensor, pipeline, data, and expert parallelism. Use tensor parallelism across GPUs on one node first, pipeline parallelism across nodes when one node is not enough, and independent replicas behind a load balancer for extra throughput on dense models. Built-in data parallelism matters mainly for mixture-of-experts models.
vLLM runs on NVIDIA and AMD GPUs, Google TPUs, AWS Trainium and Inferentia, Intel Gaudi and XPU accelerators, and CPUs. NVIDIA GPUs have the deepest support; CPUs work for testing but are not a serving target.

\Run open-weight models without running GPUs
You now have a vLLM server on one GPU and the one-line change that moves the same client to Telnyx-hosted GPUs. Sign up to test open-weight models on in-region GPUs, at up to 75% less than proprietary alternatives, with no infrastructure to manage. Start building with Telnyx Inference
When running your own GPUs stops paying, point the same OpenAI-compatible client at Telnyx and get in-region GPUs at up to 75% less than proprietary alternatives.
Try Telnyx InferenceRelated articles
AI DLP Test: Six Data Types, Six Blocks, $0 Charged

The best open source LLMs in 2026 for coding, agents, and voice

Open-Source Models Are Catching Up to Frontier
Activate your eSIM manually with an SM-DP+ address

IoT Devices: Definition, Examples, and How to Connect Them

Edge computing applications: 12 real use cases and latency budgets
