EDGE

Inference Pricing

OpenAI-compatible LLMs on owned GPUs. Per token, per model. Open-source models catching up with frontier at a fraction of the cost, with zero egress to storage and inference.

ADJUST SETTINGS
01 T
0100%
01 T
LLM
ESTIMATED COST
$368Per month
Input tokens150M × $0.665 / 1M
$99.75
Cached input tokens350M × $0.08 / 1M
$28.00
Output tokens60M × $4.00 / 1M
$240.00

Published rates, pay as you go, before volume tiers. Caching means large prompts cost far less than list rate; actual billing is per token at the rates in the table below, with no prompt-size multipliers and no minimums.

Rates

What's on the card.

Each row is a line on your invoice. Nothing is bundled into another primitive's rate and nothing turns on unless you call it.

Chat completions pricing (per 1M tokens)

Chat completions pricing (per 1M tokens)

ModelPrice
GLM-5.3 — Flagship frontier intelligence

Input: $1.400 / 1M tokens
Cached Input: $0.260 / 1M tokens
Output: $4.400 / 1M tokens

GLM-5.3 Flash — Fast, efficient inference

Input: $0.150 / 1M tokens
Cached Input: $0.030 / 1M tokens
Output: $0.500 / 1M tokens

DeepSeek V4.1 Flash — Low-latency multimodal inference

Input: $0.300 / 1M tokens
Cached Input: $0.006 / 1M tokens
Output: $1.200 / 1M tokens

Kimi K3 — Multimodal intelligence, 1M context

Input: $2.700 / 1M tokens
Cached Input: $0.270 / 1M tokens
Output: $13.500 / 1M tokens

Kimi K2.6 — Highest intelligence, voice AI

Input: $0.665 / 1M tokens
Cached Input: $0.080 / 1M tokens
Output: $4.000 / 1M tokens

GLM-5.2 — Next-gen efficient reasoning

Input: $1.000 / 1M tokens
Cached Input: $0.200 / 1M tokens
Output: $4.000 / 1M tokens

MiniMax-M3 — Cheapest while maintaining high intelligence

Input: $0.270 / 1M tokens
Cached Input: $0.080 / 1M tokens
Output: $1.100 / 1M tokens

DeepSeek V4 Flash — Ultra-fast, cost-efficient inference

Input: $0.130 / 1M tokens
Cached Input: $0.030 / 1M tokens
Output: $0.260 / 1M tokens

Qwen 3.8 27B — Lightweight open-weight reasoning

Input: $0.400 / 1M tokens
Cached Input: $0.050 / 1M tokens
Output: $3.000 / 1M tokens

Chat completions pricing (per 1M tokens)

ModelPrice
GLM-5.2 — Next-gen efficient reasoning

Input: $2.000 / 1M tokens
Cached Input: $0.400 / 1M tokens
Output: $8.000 / 1M tokens

Other services

ServicePrice
Embeddings (gte-large)

$0.0001 / 1K tokens

Speech to text

$0.003 / minute

AI-enabled storage and retrieval

$0.02 / GB / day

Machine readable: GET /v2/public/pricing. Rates may vary by destination and volume tier. Machine readable: GET /v2/public/pricing?primitive=inference.

VOLUME

Commit more, pay less per unit. Nothing is gated.

PAY AS YOU GO
$0

Standard rates. Card on file. Automatic tier discounts as usage crosses thresholds, no contract needed.

START BUILDING
COMMITTED
from $500/ mo

Discounted rates across everything you use, not one product. Named account manager, priority support, higher limits.

GET A VOLUME RATE
ENTERPRISE
from $5,000/ mo

Dedicated infrastructure and IPs, private interconnect, in region GPUs, unlimited rate limits.

CONTACT SALES

Live in five minutes. No sales gate in front of the API.

Card on file, provisioned by API. When you want a human, they run the network and pick up 24/7.