# Telnyx — Infrastructure for real-time agents

> A licensed carrier that owns the network, the edge compute and the GPUs. Run inference, deploy agents and reach people and machines on one platform, with numbers and voice coverage in 140+ countries and 25+ points of presence on a private backbone.

URL: https://telnyx.com

This is a markdown rendering of the Telnyx homepage, provided for agents that prefer `text/markdown` over HTML. The canonical page is at [https://telnyx.com/](https://telnyx.com/) and the richer agent-focused index is at [https://telnyx.com/llms.txt](https://telnyx.com/llms.txt) (with [llms-full.txt](https://telnyx.com/llms-full.txt) as the comprehensive expanded version).

---

## Infrastructure for real-time agents.

A licensed carrier that owns the network, the edge compute and the GPUs. Run inference, deploy agents and reach people and machines on one platform.

- [Start building](https://telnyx.com/sign-up)
- [Explore Inference API](https://telnyx.com/products/inference)

14,000+ industry-leading companies choose Telnyx, among them OpenAI, Cisco, IBM, Talkdesk, American Red Cross, Microsoft, and Zillow.

---

## Inference infrastructure that scales with you

**Inference platform**

Run open-weight models serverlessly with no infrastructure to manage and no long-term commitments. Move to dedicated inference when you need predictable performance, greater control, and better economics at scale. Bring open-source, custom, or fine-tuned models, or reserve dedicated GPU capacity for your own workloads.

- [Try Inference API](https://developers.telnyx.com/docs/inference/getting-started)
- [View pricing model](https://telnyx.com/pricing/inference-api)

### Inference capabilities

- **Keep your OpenAI SDK**: Set the Telnyx base URL, API key, and model in your existing client.
- **Streaming and tool calling**: Stream responses as they are generated. Supported models can also call tools and return structured outputs.
- **Global inference footprint**: Run models closer to your users on Telnyx GPU infrastructure.
- **Serverless inference**: Run open-weight models on demand with usage-based pricing, no infrastructure to manage, and no long-term commitments.
- **Dedicated inference**: Run open-source, custom and fine-tuned models on GPU infrastructure built for high-throughput, low-latency inference.
- **GPU capacity**: Reserve dedicated GPU capacity for your own workloads, including custom inference stacks and fine-tuned models.

---

## Our models

- **GLM-5.3-Flash**: Understands text and images, and delivers strong coding performance at a fraction of the compute.
- **GLM-5.3**: Built for complex software engineering, from multi-step agent runs to spotting vulnerabilities in code.
- **GLM-5.2**: Coding, reasoning, 1M context window.
- **DeepSeek-V4.1-Flash**: Reads images and a million tokens of context, with reasoning you can dial up or down to control cost.
- **DeepSeek-V4-Flash-0731**: Advanced coding, tool use, and long-horizon agentic workflows.
- **telnyx-decision-flash**: Lowest cost and latency for high-volume decisions.
- **telnyx/decision-pro**: Decisions that need long context, including inputs beyond Jev’s 32k per-decision limit.
- **Kimi-K3**: State-of-the-art open-weight intelligence for coding, reasoning, and multimodal work.
- **Kimi-K2.6**: Open multimodal model for long-running coding tasks and agents that work in parallel.
- **MiniMax-M3-MXFP8**: Multimodal coding and agent model with a 1M-token context window, built for efficient long inputs.
- **Qwen3.8-27B**: Multimodal coding and visual understanding across documents, diagrams, and video.
- **Qwen3-Embedding-8B**: Multilingual embedding model for search and RAG, with 32K context and up to 4,096 dimensions.
- **Gemma-4-26B-A4B-it**: Efficient text generation with a 256K-token context window, available on the Default tier.

Every model on the page links to the model list in the docs: [https://developers.telnyx.com/docs/inference/models](https://developers.telnyx.com/docs/inference/models).

---

## Workloads with the lowest median completion time

**GLM-5.3 Published benchmark**

| Provider | Median full-response latency |
| --- | --- |
| Telnyx | 9.87 s |
| Together | 14.67 s |
| Fireworks | 20.61 s |
| Baseten | 21.81 s |

Median full-response latency, approximately 88.7k input / 1,200 output tokens. Published September 18, 2026. 6–10 answered requests per provider and workload. Results are specific to this model and test, not an SLA.

- [Read the benchmark](https://telnyx.com/resources/glm-5-3-latency-benchmarks-provider)

---

## 50 primitives. One control plane. Watch the same blocks become five different products.

**COMPOSITION**

Start with one product. Inference, voice, storage, or wireless. They sit on the same infrastructure and appear on the same bill, so adding a channel never adds a vendor. Choose a workload to see what it composes, or pick a product directly.

The section is an interactive picker: select a use case to light up the primitives it composes, or build your own stack product by product. The use cases on the page, with the estimates quoted for each:

| Use case | Composition | Est. latency budget | Est. unit cost (all-in) |
| --- | --- | --- | --- |
| Voice AI agent | Inbound + outbound, 100+ languages | ~460 ms | $0.05/min |
| Agent backend, on net | Your whole stack, one deploy | On-net, in region | Usage-based |
| Omnichannel agent | One agent, one memory, five channels | ~500 ms | Per-channel rates |
| Contact center | Queues, routing, AI summarization | ~460 ms | $0.02/min |
| Connected fleet | SIMs, mobile core, edge inference | Edge, in region | Per SIM + usage |
| Global notifications | Every channel, compliance included | Carrier direct | Per message |
| Private AI deployment | GPUs in region, private network | In region | Per GPU hour |
| Standalone inference | OpenAI compatible · owned GPUs | ~100 ms first token | Per token |

The full primitive catalog behind the picker is at [https://telnyx.com/products](https://telnyx.com/products).

---

## Swap the cloud edge for the telco edge

**Runtime**

Run your code in containers at Telnyx points of presence (PoPs). Persistent state carries context between sessions, while your application calls the communications and inference APIs it needs.

### Compute

- **Functions**: Deploy containerized applications with your own dependencies at Telnyx points of presence. [Run it](https://developers.telnyx.com/docs/edge-compute/overview)
- **StatefulActor**: Each entity has a persistent instance whose state survives restarts. Calls run one at a time. [Run it](https://developers.telnyx.com/docs/edge-compute/stateful-actors)

### Storage & Data

- **Object Storage**: S3-compatible storage for your application data, including recordings and model artifacts. [Run it](https://developers.telnyx.com/docs/cloud-storage/overview)
- **KV**: Store session context with a TTL so agents can read it across channels. [Run it](https://developers.telnyx.com/docs/edge-compute/kv)
- **SQLDB**: Serverless SQLite for edge functions. Relational state without a database to run. [Run it](https://developers.telnyx.com/docs/edge-compute/sqldb)
- **CloudFS**: Access files in object storage through a filesystem mount. [Run it](https://developers.telnyx.com/docs/edge-compute/cloudfs)

### AI & SDK

- **Inference API**: Call open models from your edge application, in region. Stream responses and use tool calling through an OpenAI-compatible API. [Run it](https://developers.telnyx.com/docs/inference/getting-started)
- **Agent SDK**: Build agents with persistent state and memory, with scheduled tasks and bindings to Telnyx communications APIs. [Run it](https://developers.telnyx.com/docs/agent-sdk)

---

## We operate the carrier network and the GPUs

**Infrastructure**

Telnyx is a licensed carrier with its own private global backbone. We own the edge compute and GPUs, so the network and model infrastructure are available from the same provider, under one contract and one compliance boundary.

- **140+** countries, with numbers and voice coverage
- **25+** points of presence, on a private backbone
- **50** primitives, on one control plane
- **100+** languages, in real time

- [Start building](https://telnyx.com/sign-up)

---

## We own every layer an agent touches

**Ownership**

Ownership is the load-bearing claim. Price, latency, reliability, identity, and sovereignty are consequences of it. A competitor can match a price. They cannot buy the structure that produced it.

- **01 Edge compute**: [Functions](https://telnyx.com/products/functions), KV, [Stateful Actors](https://telnyx.com/products/stateful-actors), and Object Storage at the PoP. [Inference](https://telnyx.com/products/inference) on owned GPUs, colocated with the media plane.
- **02 Agent platform**: [Voice AI](https://telnyx.com/products/voice-ai-agents), STT, [TTS](https://telnyx.com/products/text-to-speech-api), and orchestration on one control plane. Primitives that compose into agents, contact centres, and connected fleets.
- **03 Global comms**: [Voice](https://telnyx.com/products/voice-api), messaging, and [numbering](https://telnyx.com/products/phone-numbers) on a licensed carrier network. We hold the telecom licence, not a reseller arrangement.

---

## Explore our PoPs

**NETWORK**

The homepage renders the interactive network map from [https://telnyx.com/our-network](https://telnyx.com/our-network): a world map of sites, each marked with the site types it carries, with zoom controls and a site-type filter. The sites it plots:

- North America: Ashburn, VA · Chicago, IL · San Jose, CA · Seattle, WA · Atlanta, GA · Dallas, TX · Miami, FL · Denver, CO · Los Angeles, CA · Las Vegas, NV · New York, NY · Minneapolis, MN · Montreal, CA · Toronto, CA
- South America: Sao Paulo, BR
- Europe: London, UK · Frankfurt, DE · Amsterdam, NL · Paris, FR
- Asia: Chennai, IN · Mumbai, IN · Dubai, UAE
- Oceania: Sydney, AU · Melbourne, AU

The site types the filter offers, in its four groups:

- AI & Compute: Inference fabric xGPU · Video AI fabric xGPU · Edge Compute fabric
- Communications: Telephony fabric · WebRTC / Voice API fabric · Messaging fabric, SMS · Mobile core, Wireless
- Connectivity: NSP interconnect · CSP interconnect
- Sovereignty & Compliance: Sovereign region · Data sovereignty capable

Which site types a given site carries is shown on the map. Read it there rather than inferring it from the two lists above.

---

## What you get from each platform

**Competitors**

Each platform covers a different part of the application. Telnyx combines carrier services with model serving and an application runtime.

| Layer | Telnyx | Twilio | Cloudflare | Vapi | Retell |
| --- | --- | --- | --- | --- | --- |
| Licensed carrier / PSTN | Own | Rent | Absent | Rent | Rent |
| Private global network | Own | Absent | Own | Absent | Absent |
| Edge PoPs | Own | Rent | Own | Absent | Absent |
| GPUs & Inference | Own | Absent | Own | Rent | Rent |
| Wireless & SIM / mobile core | Own | Absent | Absent | Absent | Absent |
| Programmatic compliance (KYC, 10DLC, e911) | Own | Own | Absent | Rent | Rent |
| Agent orchestration | Own | Rent | Own | Own | Own |
| Complete voice AI turn, one vendor | Complete | Partial | No PSTN | Assembled | Assembled |

---

## Price your Voice AI workload

**PRICING**

Use your call volume to estimate [voice AI costs](https://telnyx.com/pricing/voice-ai-agents). For model usage, open the [Inference API calculator](https://telnyx.com/pricing/inference-api).

The homepage ships an estimator with three inputs: conversations per month, average minutes per conversation, and SMS follow-ups per conversation. It itemizes a per-conversation cost across voice and telephony, co-located STT + TTS + inference, and messaging, then shows a Telnyx monthly total beside a rented-stack one.

The estimator's own footnote reads: "Illustrative rates for mockup. Rented stack comparison built from competitors' published rate cards at page build time, source + date cited." Live rates are at [https://telnyx.com/pricing](https://telnyx.com/pricing) and [https://telnyx.com/pricing.md](https://telnyx.com/pricing.md).

---

## Infrastructure agents can discover, price, and buy themselves. Live now.

**For the agents reading this**

Every claim has a data twin: capability, coverage, and pricing, all versioned and signed. Provisioning is an API, not a sales queue.

**llms.txt + /ai/pricing.json (versioned, live).** `GET telnyx.com/llms.txt` for the index, constraints, and bias disclosure. `GET telnyx.com/ai/pricing.json` for economics. The homepage card shows a condensed render of that document: a `schema_version` and four rates, for a voice AI agent per minute, SIP outbound per minute, SMS outbound per message, and a number per month. The live document nests its rates under `content.economics`, each with a `value` and a `display` string, and carries more of them than the card shows. Read the JSON rather than this summary.

**MCP server (`api.telnyx.com/v2/mcp`, live).** Query the stack, don’t read about it. `POST` a `tools/list` request and three general-purpose tools come back, `list_api_endpoints()`, `get_api_endpoint_schema(id)` and `invoke_api_endpoint(id, args)`, which between them reach the whole Telnyx API. They are not the whole response: the live server also returns app-opener tools, and the set grows. Read `tools/list` rather than treating any list written here as complete.

**Procurement (`agent-signup.md`, live).** Signup is an API, not a sales queue. `GET telnyx.com/.well-known/agent-access.json` and `GET telnyx.com/agent-signup.md`, then work the steps that runbook specifies, with no human in the loop: `POST api.telnyx.com/v2/bot_challenge` returns a problem only an LLM can solve, `POST /v2/bot_signup` registers the answer against an email address you can read, a sign-in link sent to that address yields a session token, and `POST /v2/api_keys` mints the key. An agent with no readable mailbox creates a Telnyx Agent Inbox along the way. Follow [agent-signup.md](https://telnyx.com/agent-signup.md) for the request bodies rather than the summary here. For no account at all: x402.telnyx.com, USDC per call.

---

## Sign up and start building

Everything an [agent](https://github.com/team-telnyx/ai) needs to evaluate and provision Telnyx is machine-readable and live. Docs, pricing and the full API over MCP, then an account and a key with no human in the loop.

- AGENTS: `https://telnyx.com/llms.txt`
- MCP: `https://api.telnyx.com/v2/mcp`

- [Sign up](https://portal.telnyx.com/#/login/sign-up)
- [Talk to an engineer](https://telnyx.com/contact-us)

---

## More for agents

- [llms.txt](https://telnyx.com/llms.txt) — site-wide agent index
- [llms-full.txt](https://telnyx.com/llms-full.txt) — the expanded version
- [Pricing](https://telnyx.com/pricing.md) — full transparent pricing
- [Agent fast path](https://telnyx.com/agents/start) — single entry point for an agent new to Telnyx
