Infrastructure for real-time agents.

A licensed carrier that owns the network, the edge compute and the GPUs. Run inference, deploy agents and reach people and machines on one platform.

14,000+ INDUSTRY-LEADING COMPANIES choose telnyx

OpenAI - Artificial intelligence research leader using Telnyx communicationsCisco - Networking and telecommunications company using Telnyx servicesIBM - Global technology and consulting company partnering with TelnyxTalkdesk - Cloud contact center platform powered by TelnyxAmerican Red Cross - Humanitarian organization leveraging Telnyx communicationsMicrosoft - Technology corporation utilizing Telnyx infrastructureZillow - Real estate marketplace using Telnyx for customer communicationsOpenAI - Artificial intelligence research leader using Telnyx communicationsCisco - Networking and telecommunications company using Telnyx servicesIBM - Global technology and consulting company partnering with TelnyxTalkdesk - Cloud contact center platform powered by TelnyxAmerican Red Cross - Humanitarian organization leveraging Telnyx communicationsMicrosoft - Technology corporation utilizing Telnyx infrastructureZillow - Real estate marketplace using Telnyx for customer communications

Inference platform

Inference infrastructure that scales with you

Run open-weight models serverlessly with no infrastructure to manage and no long-term commitments. Move to dedicated inference when you need predictable performance, greater control, and better economics at scale. Bring open-source, custom, or fine-tuned models, or reserve dedicated GPU capacity for your own workloads.

  • Keep your OpenAI SDK

    Set the Telnyx base URL, API key, and model in your existing client.

  • Streaming and tool calling

    Stream responses as they are generated. Supported models can also call tools and return structured outputs.

  • Global inference footprint

    Run models closer to your users on Telnyx GPU infrastructure.

  • Serverless inference

    Run open-weight models on demand with usage-based pricing, no infrastructure to manage, and no long-term commitments.

  • Dedicated inference

    Run open-source, custom and fine-tuned models on GPU infrastructure built for high-throughput, low-latency inference.

  • GPU capacity

    Reserve dedicated GPU capacity for your own workloads, including custom inference stacks and fine-tuned models.

Our models

  • GLM-5.3-Flash

    Understands text and images, and delivers strong coding performance at a fraction of the compute

  • GLM-5.3

    Built for complex software engineering, from multi-step agent runs to spotting vulnerabilities in code

  • GLM-5.2

    Coding, reasoning, 1M context window

  • DeepSeek-V4.1-Flash

    Reads images and a million tokens of context, with reasoning you can dial up or down to control cost

  • DeepSeek-V4-Flash-0731

    Advanced coding, tool use, and long-horizon agentic workflows

  • telnyx-decision-flash

    Lowest cost and latency for high-volume decisions

  • telnyx/decision-pro

    Decisions that need long context, including inputs beyond Jev’s 32k per-decision limit

  • Kimi-K3

    State-of-the-art open-weight intelligence for coding, reasoning, and multimodal work

  • Kimi-K2.6

    Voice AI

  • MiniMax-M3-MXFP8

    Lowest cost at high intelligence

  • Qwen3.8-27B

    Multimodal coding and visual understanding across documents, diagrams, and video

  • Qwen3-Embedding-8B

    Multilingual embedding model for search and RAG, with 32K context and up to 4,096 dimensions.

GLM-5.3 Published benchmark

Workloads with the lowest median completion time

Read the benchmark
Median full-response latency · approximately 88.7k input / 1,200 output tokens
Telnyx
9.87 s
Together
14.67 s
Fireworks
20.61 s
Baseten
21.81 s

Published September 18, 2026. Full-response latency; 6–10 answered requests per provider and workload. Results are specific to this model and test, not an SLA.

COMPOSITION

50 primitives. One control plane. Watch the same blocks become five different products.

Start with one product. Inference, voice, storage, or wireless. They sit on the same infrastructure and appear on the same bill, so adding a channel never adds a vendor. Choose a workload to see what it composes, or pick a product directly.

Loading...

Runtime

Swap the cloud edge for the telco edge

Run your code in containers at Telnyx points of presence (PoPs). Persistent state carries context between sessions, while your application calls the communications and inference APIs it needs.

  • Compute/docs/edge-compute/overview

    Functions

    Deploy containerized applications with your own dependencies at Telnyx points of presence.

    Run it
  • Compute/docs/edge-compute/stateful-actors

    StatefulActor

    Each entity has a persistent instance whose state survives restarts. Calls run one at a time.

    Run it
  • Storage & Data/docs/cloud-storage/overview

    Object Storage

    S3-compatible storage for your application data, including recordings and model artifacts.

    Run it
  • Storage & Data/docs/edge-compute/kv

    KV

    Store session context with a TTL so agents can read it across channels.

    Run it
  • Storage & Data/docs/edge-compute/sqldb

    SQLDB

    Serverless SQLite for edge functions. Relational state without a database to run.

    Run it
  • Storage & Data/docs/edge-compute/cloudfs

    CloudFS

    Access files in object storage through a filesystem mount.

    Run it
  • AI & SDK/docs/inference/getting-started

    Inference API

    Call open models from your edge application, in region. Stream responses and use tool calling through an OpenAI-compatible API.

    Run it
  • AI & SDK/docs/agent-sdk

    Agent SDK

    Build agents with persistent state and memory, with scheduled tasks and bindings to Telnyx communications APIs.

    Run it

Infrastructure

We operate the carrier network and the GPUs

Telnyx is a licensed carrier with its own private global backbone. We own the edge compute and GPUs, so the network and model infrastructure are available from the same provider, under one contract and one compliance boundary.

Start building
Countries
140+
With numbers and voice coverage.
Points of presence
25+
On a private backbone.
Primitives
50
On one control plane.
Languages
100+
In real time.

Ownership

We own every layer an agent touches

Ownership is the load-bearing claim. Price, latency, reliability, identity, and sovereignty are consequences of it. A competitor can match a price. They cannot buy the structure that produced it.

Edge compute

Functions, KV, Stateful Actors, and Object Storage at the PoP. Inference on owned GPUs, colocated with the media plane.

Agent platform

Voice AI, STT, TTS, and orchestration on one control plane. Primitives that compose into agents, contact centres, and connected fleets.

Global comms

Voice, messaging, and numbering on a licensed carrier network. We hold the telecom licence, not a reseller arrangement.

NETWORK

Explore our PoPs

Competitors

What you get from each platform

Each platform covers a different part of the application. Telnyx combines carrier services with model serving and an application runtime.

LAYER
Telnyx logo
Twilio logo
Cloudflare logo
Vapi logo
Retell logo

Licensed carrier / PSTN

Own
Rent
Absent
Rent
Rent

Private global network

Own
Absent
Own
Absent
Absent

Edge PoPs

Own
Rent
Own
Absent
Absent

GPUs & Inference

Own
Absent
Own
Rent
Rent

Wireless & SIM / mobile core

Own
Absent
Absent
Absent
Absent

Programmatic compliance (KYC, 10DLC, e911)

Own
Own
Absent
Rent
Rent

Agent orchestration

Own
Rent
Own
Own
Own

Complete voice AI turn, one vendor

Complete
Partial
No PSTN
Assembled
Assembled

PRICING

Price your Voice AI workload

Use your call volume to estimate voice AI costs. For model usage, open the Inference API calculator.

Loading...

For the agents reading this

Infrastructure agents can discover, price, and buy themselves. Live now.

Every claim has a data twin: capability, coverage, and pricing, all versioned and signed. Provisioning is an API, not a sales queue.

  • llms.txt + /ai/pricing.jsonVersioned, live
    // canonical machine claims
    GET telnyx.com/llms.txt
      // index, constraints,
      // bias disclosure
    GET telnyx.com/ai/pricing.json
    {
     "schema_version": "1.0.0",
     "economics": {
      "voice_ai_agent_usd_per_min": 0.05,
      "sip_outbound_usd_per_min": 0.005,
      "sms_outbound_usd_per_msg": 0.004,
      "number_monthly_usd": 1
     }
    }
  • MCP serverapi.telnyx.com/v2/mcp, live
    // query the stack,
    // don't read about it
    POST api.telnyx.com/v2/mcp
      Accept: application/json,
        text/event-stream
    
    {"jsonrpc":"2.0","id":1,
     "method":"tools/list"}
    → list_api_endpoints()
    → get_api_endpoint_schema(id)
    → invoke_api_endpoint(id, args)
    // 3 tools reach the whole Telnyx API
  • Procurementagent-signup.md, live
    // signup is an API,
    // not a sales queue
    GET telnyx.com/.well-known/
        agent-access.json
    GET telnyx.com/agent-signup.md
    POST api.telnyx.com/v2/bot_challenge
      → a problem only an LLM can solve
      → account, inbox, API key
        — no human in the loop
    // no account at all:
    // x402.telnyx.com, USDC per call

Sign up and start building

Everything an agent needs to evaluate and provision Telnyx is machine-readable and live. Docs, pricing and the full API over MCP, then an account and a key with no human in the loop.

AGENTS: telnyx.com/llms.txt
MCP: api.telnyx.com/v2/mcp