Kimi K3 — Flagship multimodal intelligence, 1M context
Input: $2.700 / 1M tokens
Cached Input: $0.270 / 1M tokens
Output: $13.500 / 1M tokens
Kimi K2.6 — Highest intelligence, voice AI
Input: $0.665 / 1M tokens
Cached Input: $0.080 / 1M tokens
Output: $4.000 / 1M tokens
GLM-5.2 — Next-gen efficient reasoning
Input: $1.000 / 1M tokens
Cached Input: $0.200 / 1M tokens
Output: $4.000 / 1M tokens
GLM-5.3 — Next-generation frontier intelligence
Input: $1.400 / 1M tokens
Cached Input: $0.260 / 1M tokens
Output: $4.400 / 1M tokens
GLM-5.3 Flash — Fast, efficient inference, 50% launch pricing until September 9, 2026
Input: $0.150 $0.075* / 1M tokens
Cached Input: $0.030 $0.015* / 1M tokens
Output: $0.500 $0.250* / 1M tokens
MiniMax-M3 — Cheapest while maintaining high intelligence
Input: $0.270 / 1M tokens
Cached Input: $0.080 / 1M tokens
Output: $1.100 / 1M tokens
DeepSeek V4 Flash — Ultra-fast, cost-efficient inference
Input: $0.130 / 1M tokens
Cached Input: $0.030 / 1M tokens
Output: $0.260 / 1M tokens
Qwen 3.8 27B — Lightweight open-weight reasoning
Input: $0.400 / 1M tokens
Cached Input: $0.050 / 1M tokens
Output: $3.000 / 1M tokens