Standard rates. Card on file. Automatic tier discounts as usage crosses thresholds, no contract needed.
OpenAI-compatible LLMs on owned GPUs. Per token, per model. Open-source models catching up with frontier at a fraction of the cost, with zero egress to storage and inference.
Published rates, pay as you go, before volume tiers. Caching means large prompts cost far less than list rate; actual billing is per token at the rates in the table below, with no prompt-size multipliers and no minimums.
Each row is a line on your invoice. Nothing is bundled into another primitive's rate and nothing turns on unless you call it.
| Model | Price |
|---|---|
| GLM-5.3 — Flagship frontier intelligence | Input: $1.400 / 1M tokens |
| GLM-5.3 Flash — Fast, efficient inference | Input: $0.150 / 1M tokens |
| DeepSeek V4.1 Flash — Low-latency multimodal inference | Input: $0.300 / 1M tokens |
| Kimi K3 — Multimodal intelligence, 1M context | Input: $2.700 / 1M tokens |
| Kimi K2.6 — Highest intelligence, voice AI | Input: $0.665 / 1M tokens |
| GLM-5.2 — Next-gen efficient reasoning | Input: $1.000 / 1M tokens |
| MiniMax-M3 — Cheapest while maintaining high intelligence | Input: $0.270 / 1M tokens |
| DeepSeek V4 Flash — Ultra-fast, cost-efficient inference | Input: $0.130 / 1M tokens |
| Qwen 3.8 27B — Lightweight open-weight reasoning | Input: $0.400 / 1M tokens |
| Model | Price |
|---|---|
| GLM-5.2 — Next-gen efficient reasoning | Input: $2.000 / 1M tokens |
| Service | Price |
|---|---|
| Embeddings (gte-large) | $0.0001 / 1K tokens |
| Speech to text | $0.003 / minute |
| AI-enabled storage and retrieval | $0.02 / GB / day |
Machine readable: GET /v2/public/pricing. Rates may vary by destination and volume tier. Machine readable: GET /v2/public/pricing?primitive=inference.
Standard rates. Card on file. Automatic tier discounts as usage crosses thresholds, no contract needed.
START BUILDINGDiscounted rates across everything you use, not one product. Named account manager, priority support, higher limits.
GET A VOLUME RATEDedicated infrastructure and IPs, private interconnect, in region GPUs, unlimited rate limits.
CONTACT SALESSee the whole composition on private ai deployment ↗
Card on file, provisioned by API. When you want a human, they run the network and pick up 24/7.