Inference

Inference Cost Optimization: How to Cut Your AI Bill by up to 75%

Four levers to cut inference costs by up to 75%: open-source models, cached input, batching, owned infrastructure. The worked example takes a $50,000 monthly bill to about $12,400.

Four levers: open-source models, cached input, batching, owned infrastructure. Pull all four and your inference bill can drop by up to 75%.

That is not a promotional discount or a limited-time offer. It is the structural difference between renting compute from a chain of providers and running on infrastructure owned end to end, because most inference providers rent GPUs from cloud platforms, stack their margin on top, and pass the compounded cost to you in every token.

One owned stack, one margin. Price follows physics.

If your inference bill is growing faster than revenue, the problem is usually the cost structure underneath the usage. Here are the four levers that fix it, with the rates to run your own numbers.

How do you optimize inference costs?

Inference cost optimization comes down to four levers: switching to open-source model rates, caching repeated input tokens, batching latency-tolerant work, and running on infrastructure the provider owns. Pulled together on Telnyx, they take a $50,000 monthly bill to about $12,400 in the worked example below, and cached input alone cuts up to 98 percent off the uncached input rate. Every Telnyx rate in this article comes from the published pricing page, and the proprietary comparisons use OpenAI's current published rates.

Lever 1: switch to open-source inference models

Up to 75% savings versus proprietary alternatives, and the rates below are the receipt for that claim.

The gap between open-source and frontier proprietary models has closed. A curated frontier of open-source models now matches or exceeds proprietary quality on production workloads, from reasoning and function calling to multilingual generation and real-time voice. The quality argument for paying proprietary premiums is yesterday's conversation.

Open weights mean no vendor lock-in. You control the deployment, the fine-tuning, and the data pipeline, so a vendor that deprecates a model or rewrites its terms cannot strand your roadmap.

Here is the full lineup with published pay-as-you-go chat completions rates, per 1M tokens, from the Telnyx Inference pricing page. The second number in the middle column is the cached input rate:

ModelInput / cached input per 1MOutput per 1M
GLM-5.3, flagship frontier intelligence$1.250 / $0.240$4.000
GLM-5.3 Flash, fast and efficient inference$0.135 / $0.027$0.450
DeepSeek V4.1 Flash, low-latency multimodal inference$0.300 / $0.006$1.200
Kimi K3, multimodal intelligence with 1M context$2.700 / $0.270$13.500
Kimi K2.6, highest intelligence and voice AI$0.665 / $0.080$4.000
GLM-5.2, next-gen efficient reasoning$1.000 / $0.200$4.000
MiniMax-M3, cheapest while maintaining high intelligence$0.270 / $0.080$1.100
DeepSeek V4 Flash, ultra-fast, cost-efficient inference$0.130 / $0.030$0.260
Qwen 3.8 27B, lightweight open-weight reasoning$0.400 / $0.050$3.000

These are the standard rates before volume tiers, with no prompt-size multipliers and no minimums behind them. Kimi K2.6, the voice AI model, runs $0.665 input against $4.000 output, while Kimi K3, Moonshot AI's multimodal open model with a 1M-token context window, tops the lineup at $2.700 input and $13.500 output.

Picking the right model per workload matters more than any single rate. For most teams the cheapest capable model wins the bulk of the traffic while the flagship carries the hard calls. The API is OpenAI-compatible, so the switch is a base URL change, and the Inference API quickstart shows the whole swap in a few lines of code.

Lever 2: use cached input for inference

If your application sends repeated context, system prompts, few-shot examples, or any fixed prefix across requests, you are paying full price for computation the provider already performed. Cached input stores the computed representation of repeated prefixes and reuses them across requests, so the same prompt costs a fraction of list rate on every pass after the first. On Telnyx the caching is automatic with no configuration, and cache writes are not charged, so building up a cache costs nothing.

A voice agent with a 2000-token system prompt running 100,000 conversations a month re-sends 200 million input tokens of identical prefix. Uncached on Kimi K2.6, those tokens cost 200M × $0.665, or $133 a month. Cached at $0.080, the same prefix traffic costs $16, an 88 percent cut on that line item.

The discount varies by model. DeepSeek V4.1 Flash caches at $0.006 against a $0.300 input rate, which is 98 percent off, Kimi K3 caches at 90 percent off, and MiniMax-M3 at 70 percent off still beats full price on any repeated prefix.

Caching only pays where prefixes repeat, so point it at system prompts, few-shot templates, and stable knowledge-base context rather than one-off inputs. A document pipeline that never sees the same document twice writes cache entries nothing will read, and those requests simply pay the uncached input rate.

A cache miss costs nothing extra, because the first request over a new prefix pays the normal input price and writes the cache for the requests behind it. Cached prefixes persist for a limited window, so traffic that arrives days apart can re-pay the input rate and rebuild.

For agents that maintain conversation state across turns, stateful actors keep memory on the same infrastructure as the GPUs, and KV provides fast key-value storage for context windows and session data on the same network. The caching layer is invisible to your code, and the savings show up on the invoice.

Lever 3: batch your latency-tolerant inference work

Batching is the lever most cost guides skip. Grouping many requests into one forward pass lets the GPU amortize fixed costs across the group instead of paying them per request, and major API providers price batch jobs below interactive requests for exactly that reason.

The trade is latency. Batching fits document processing, classification, enrichment, and nightly generation, and it does not belong in interactive chat or voice, where the user is waiting.

On Telnyx there is no separate batch tier to navigate. Latency-tolerant work belongs on the efficient models, where GLM-5.3 Flash inputs at $0.135 and DeepSeek V4 Flash at $0.130 already price bulk work the way a batch tier would, and volume tiers apply as spend grows. Throughput compounds the saving, because the MiniMax-M3 benchmark measured 144.9 median tokens per second on Telnyx against 119.8 on Together AI, and more tokens per GPU hour mean a lower effective cost per token for exactly the workloads that batch well.

Lever 4: choose an inference provider that owns infrastructure

No stacked cloud margins. One owned stack, one margin.

Most inference providers do not own GPUs. They rent compute from cloud platforms, add their margin, and resell it to you, and some orchestrate across ten or more clouds, so every layer in that chain takes a margin that shows up in your per-token price. When a provider drops their price, it usually traces back to a better cloud negotiation rather than a change in the underlying cost structure, which makes it an advantage their next renewal can take away.

This applies to other open-source inference providers too. Together AI, Fireworks, and Baseten serve the same open-source models Telnyx does, but on rented GPU capacity, so their per-token price carries a cloud margin stacked on their own. Telnyx owns the GPU infrastructure: the models run on hardware we bought, in facilities we operate, on a network we control, so the price reflects the cost of running compute plus one margin.

The latency claim is measured rather than asserted. In the GLM-5.2 provider benchmark, the same model ran on Telnyx, Baseten, Together AI, and Fireworks across six prompt profiles, and Telnyx posted the lowest overall p50 end-to-end latency at 6.00 seconds, ahead of Baseten at 6.19, Together AI at 7.49, and Fireworks at 14.98.

In the MiniMax-M3 benchmark, Telnyx ran the same open model 12 percent faster than Together AI and 21 percent faster than Fireworks on the 1k-in, 1k-out profile, with an overall 14.72 second p95 against Fireworks at 54.06. Together AI edged time-to-first-token in that run, and we publish that too, because full-response time is the latency your users feel.

Ownership also shows up outside the price. We control the hardware, the network, and the SLA, so when something breaks there is one accountable party instead of finger-pointing across three vendors.

Our GPUs sit in the Americas, Europe, APAC, and MENA, with a strict in-region mode that fails closed rather than route a request outside your chosen region. Zero data retention: your prompts and responses are not persisted after inference completes, no request content is logged, and your data is processed and discarded. We do not train models on your data. Storage and inference share one network with zero egress between them.

A competitor can match a price. They cannot buy the structure that produced it.

Pulling all four inference cost levers

Here is the math behind the up-to-75-percent claim, worked line by line so you can replicate it on your own volumes. The starting bill is whatever your invoice says, and the after side uses the published rates above with round, illustrative volumes for a mid-size product: contact center support chat, a document pipeline, voice agents, and a hard-reasoning line for agent planning. The one assumption to challenge is the 85 percent cache-hit rate on support chat, which is conservative for a stable system prompt plus a knowledge-base prefix.

For the before side, the same volumes run about $45,200 a month on GPT-5.6 Sol at its current promotional rates of $4.00 input and $20.00 output per million tokens, and about $62,500 on GPT-5.5 at $5.00 and $30.00, so the $50,000 starting bill sits inside that proprietary band.

Workload, monthly volumesModelCost per month
Support chat, 2.0B input at 85% cache, 400M outputKimi K2.6$1,935.50
Document processing, 2.0B input, 300M outputGLM-5.3$3,700.00
Voice AI agents, 100k conversations with 2000-token prompts (200M cached) plus 300M input, 250M outputKimi K2.6$1,215.50
Hard reasoning, 800M input, 250M outputKimi K3$5,535.00
Total$12,386.00

The voice line is the one worth reading twice, because it stacks two levers at once. The 200M tokens of repeated prompts cache at $0.080 for $16.00, the 300M tokens of conversational input run at $0.665 for $199.50, and the 250M output tokens at $4.000 cost $1,000, for $1,215.50 against a workload that would otherwise pay full input price on all 500M tokens.

That lands at $12,386 a month, a cut of roughly 75 percent for this mix, with the hardest work deliberately routed to Kimi K3, the most expensive model on the page. Lean the mix harder on the efficient models and the cut goes deeper, because the levers compound: the model switch resets the base rate, and the cache discount applies to the open-source rates rather than the proprietary ones.

[BRANDED ASSET: bar_chart - monthly inference cost for the same product workload, $50,000 on frontier proprietary APIs versus $12,386 on Telnyx with all four levers pulled]

Start by running the numbers on your own bill. Pull your last month of token volumes, apply the published rates from the open-source model pricing page, and calculate what cached input saves on your repeated context. The gap between what you pay now and what you could pay is the budget for everything else your team wants to build.

Then swap your base URL and test. Your existing code works. Sign up free to get started.

Share on Social
Fiona McDonnell
Fiona McDonnell
Product Marketing Lead

Fiona McDonnell is Product Marketing Lead at Telnyx. Rather than writing about AI infrastructure and messaging compliance from the outside, she's run go-to-market directly for Telnyx's inference products and its trust-focused messaging and voice capabilities. That means she's see