In-region inference at one per-token rate.
Modal charges a 1.15x multiplier for broad regions and 1.75x for narrow regions on top of per-GPU-second compute. Its per-token endpoints offer no region selection at all. Telnyx runs open-weight models on owned GPUs in the US, EU, APAC and MENA at one per-token rate. In-region by default, no surcharges, no infra code.
14,000+ INDUSTRY-LEADING COMPANIES choose telnyx
Serverless inference runs on Telnyx-owned GPUs in the US, EU, APAC and MENA. In-region by architecture, not a premium tier.
Per-token Shared Endpoints have no region selection and route through us-west. Pinning compute to a region requires a Dedicated Endpoint at 1.15–1.75x base compute rates. Payloads over 2 MiB, logs and all stored data stay in the US.
Modal bills per-GPU-second with a 1.15–1.75x multiplier for any pinned region, and its per-token endpoints can't be pinned at all. Telnyx is per-token on owned GPUs, in-region by default, with cached input at up to 88% off and no GPU rental, plan tier or compute surcharge.
Both are OpenAI-compatible, so migration is a base URL change. What you leave behind: Modal's proxy tokens, per-model concurrency caps on shared endpoints, and the Python deployment you'd need for anything the shared pool doesn't serve. Point your existing client at Telnyx and run your first request today.
Python
from openai import OpenAI
client = OpenAI(
api_key="YOUR_TELNYX_API_KEY",
base_url="https://api.telnyx.com/v2/ai",
)
response = client.chat.completions.create(
model="moonshotai/Kimi-K3",
messages=[{"role": "user", "content": "Hello"}],
)Owned GPUs in the US, EU, APAC and MENA. No cloud markup, no regional multiplier.
From carrier network to co-located GPU compute, Telnyx owns every layer your agents need to run voice AI and inference in real time. No Frankenstack. No rented infrastructure. One control plane for inference, voice AI, and global communications. Configure once, deploy globally.
Pay as you go, no minimums. In-region inference at one per-token rate.
Both Telnyx and Modal expose OpenAI-compatible endpoints, so you can run them in parallel during migration. Point a percentage of traffic at the Telnyx base URL, validate results, then cut over. If your Modal workload is a Function or Dedicated Endpoint, the migration also removes the GPU-second bill and the region multiplier.