GLM-5.3 Flash, GLM-5.3, and Kimi K3 across 45 multi-turn agent conversations. The results show how to match a model to your workload.

Production agents rarely get a clean one-shot prompt. A user corrects the model, removes a requirement, adds a constraint, then asks for the final format. Every one of those turns is a chance to lose the thread, and a model that answers single prompts well tells you nothing about how it holds a conversation together.
So we tested what actually matters. We ran 45 complete four-turn conversations across three open-weight models, scored the final answers with deterministic checks, and had two independent blind judges grade every conversation with the model names hidden. All three models served through the same Telnyx Inference API, on the same client, at the same settings.
The result is not a single winner. Kimi K3 completed every conversation in the eyes of both judges and earned perfect correction scores. GLM-5.3 Flash delivered 90 percent completion with the fastest latency profile of the trio. GLM-5.3 matched Kimi K3 on objective checks while trading latency for depth. Which one fits depends on what your workload punishes hardest, and that is the thesis of this guide: there is no universal best model, only the right model for the loop you are running.
The testing loop is cheap for a structural reason. Open-weight models run on GPUs we own, which is why they serve at up to 75 percent less than proprietary frontier APIs, and every model sits behind one OpenAI-compatible API. Testing three candidates means changing one line in your client, not onboarding three vendors. That is what turns model evaluation from a procurement project into an afternoon.
Here is what single-turn benchmarks miss, how the test works, what the numbers showed, and how to choose.
A production multi-turn conversation has a shape. The first turn establishes the task. Middle turns add corrections, remove requirements, and introduce constraints that quietly contradict earlier instructions. A final turn demands a specific output format. Each step depends on the model tracking what still stands and what has been superseded.
Single-turn benchmarks measure how well a model answers one prompt at a time. They say little about whether a constraint from turn two survives to turn four, or whether a requirement the user explicitly removed resurfaces in the final answer. Those failures stay quiet because the response still reads fluently, which makes them expensive to catch in production. Support agents feel it first, since triage conversations are the workload behind AI contact centers where a dropped correction becomes a wrong answer to a real customer.
An eval for multi-turn agents should measure five things: multi-turn state tracking, following corrections and constraints, final-answer usability, task completion, and objective conformance to format rules. The method below scores all five, and it runs on any OpenAI-compatible endpoint.
The dataset is 15 synthetic scenarios across three workflow families. Five developer workflows cover cursor pagination, a zero-downtime database migration, webhook verification, JSON normalization, and distributed rate limiting. Five business workflows cover support triage, product-launch planning, policy FAQ creation, account planning, and incident communications. Five planning and editing workflows cover workshop scheduling, executive editing, candidate evaluation, dependency-aware planning, and invoice extraction.
Each scenario runs four user turns. The first establishes the task, later turns add corrections, remove requirements, and introduce new constraints, and the final turn demands a specific output format. The runner resends the complete message history on every request, so memory stays explicit in the context window rather than implicit in a framework.
Two layers score the work. Deterministic checks verify valid JSON, required fields, retained facts, forbidden superseded facts, and format constraints, 28 checks per model and 84 across the three models. Two blind judges, DeepSeek V4.1 Flash and MiniMax-M3, then read every conversation with candidate names stripped from the prompts. Each judge returns a binary completion decision plus five scores from 1 to 5 across task completion, correctness, multi-turn state tracking, following corrections and constraints, and final-answer usability.
The full run produced 45 complete four-turn conversations, 180 model turns, 84 deterministic checks, and 90 blind evaluations. Every model ran with reasoning_effort set to low, temperature 0, and a 4096 token ceiling, all through the same OpenAI-compatible inference API.
One detail matters for anyone building agents. The orchestrator itself was a Stateful Actor, which persisted progress across the run and exported turn-level outputs. It drove 45 conversations and 180 model calls with no HTTP error, no timeout, and no truncated response from any of the three models. Multi-turn orchestration with persisted state is exactly the workload Stateful Actors exist for, so the benchmark doubled as a platform test.

Judge completion is the binary question of whether the conversation finished the task, out of 30 decisions per model, two judges by 15 conversations. Objective checks grade only the final answers. All three models completed all 15 conversations and all 60 model turns without an HTTP error, a timeout, or a truncated response, so these numbers measure model behavior rather than infrastructure reliability.
The two judges agreed on the binary completion call for 43 of 45 conversations, a 95.6 percent agreement rate, and their five-dimensional scores differed by a mean of 0.16 points. That agreement matters because it suggests the rankings below are not an artifact of one judge's taste.

Kimi K3 earned the highest or tied-highest score in most judge and dimension combinations. Both judges gave it a perfect 5.00 out of 5 for following corrections, the dimension multi-turn workloads punish hardest. On final-answer usability, DeepSeek scored it 5.00 and MiniMax scored it 4.80.
Notice where the rankings separate. On objective checks the three models land within five points of each other, 22, 23, and 23 out of 28. The separation shows up in judged completion, where Kimi K3 went 30 for 30 and the GLM models dropped three conversations each. A judge marks a conversation incomplete when a correction goes unheeded or a removed fact resurfaces, which is exactly the failure mode a fixed checklist struggles to enumerate.
Latency is one selection criterion among several, not the headline. For multi-turn agents the relevant latency is the full conversation, since an agent's wall-clock time is the sum of its turns plus any retries.
Three stream events were measured. TTFT is the first non-empty model output, which may still be reasoning. TTFA is the time to first visible answer text. E2E is the moment the streamed response finishes. The distinction matters for reasoning models, where TTFT can be dominated by thinking tokens the user never sees.
| Model | p50 first answer | p95 first answer |
|---|---|---|
| GLM 5.3 Flash | 0.41s | 2.60s |
| GLM 5.3 | 0.52s | 2.86s |
| Kimi K3 | 0.49s | 2.20s |

GLM-5.3 Flash posts the strongest latency profile at every median and finishes a full conversation at a p95 of 31.27 seconds, under half the observed tail of the other two. Kimi K3 delivers the best p95 first answer at 2.20 seconds but the longest total conversation tail at 55.71 seconds, a signature of deeper reasoning mid-conversation. GLM-5.3 is the slowest at every median, the price of running the flagship.
Match the profile to the loop. Interactive agents feel the p50 first answer on every turn, and background pipelines feel total conversation time. The gap between first token and finished response is its own lesson, and an earlier TTFT versus E2E benchmark shows how often the fastest first token loses the race to a finished answer. For provider-level context on these model families, a dedicated latency study of GLM 5.3 compares Telnyx against other providers. A separate GLM 5.3 Flash benchmark tracks throughput and tail behavior the same way.
Start from what your workload punishes hardest. An agent that restarts a conversation after a failed turn pays in tokens and wall-clock time, so completion and correction-following dominate. An agent that runs high volumes of cheap turns pays per token, so economics dominate. An agent that works on genuinely hard problems pays in correctness, so reasoning depth dominates. Three models, three trade-off profiles.
Pick Kimi K3 when a failed turn is expensive. It was the only model both judges scored as completing all 15 conversations, and both judges rated its correction-following at a perfect 5.00. It also posted the best p95 first answer of the trio at 2.20 seconds. Moonshot positions it for long-horizon coding, knowledge work, and reasoning. Kimi K3 on Telnyx Inference carries a 1M token context window with configurable reasoning effort. The trade is cost and tail latency. Kimi K3 output tokens cost about 30 times more than GLM-5.3 Flash output, and it carried the longest p95 conversation time in the run. When a wrong turn means restarting a long workflow, paying for the model that never drops the thread is usually the cheap option.
Pick GLM-5.3 Flash when volume is the constraint. It posted the fastest p50 first answer at 0.41 seconds, the fastest p50 conversation at 8.35 seconds, and a p95 conversation tail under half of what the other two posted. Judge completion sat at 90 percent, and it trailed on objective checks by a single check out of 28. For an agent making many turns per task, retrying roughly one conversation in ten that a judge marks incomplete can cost less than paying a premium on every token of every turn. Z.ai designed it to reduce inference compute, which is where that latency profile comes from.
Pick GLM-5.3 when the task itself is hard. It matched Kimi K3 on objective checks at 23 of 28, held 90 percent judge completion, and Z.ai positions it for complex coding and long-horizon work. The cost is latency, since it was the slowest model at every median in this run. Output tokens run about 30 percent of Kimi K3's price, so it sits between the other two on cost. Choose it when correctness on a difficult problem is worth the wall-clock time and the budget sits between Flash and Kimi.
| If your workload punishes... | Pick | Evidence from this run |
|---|---|---|
| A dropped correction | Kimi K3 | 30/30 judge completion, 5.00/5 corrections from both judges |
| Per-token cost at volume | GLM-5.3 Flash | 90% completion, fastest medians, p95 conversation tail under 32s |
| Hard problems with room in the clock | GLM-5.3 | 23/28 objective checks, flagship coding and reasoning positioning |
Whichever you pick, the economics follow from the same structure. These are open-weight models on GPUs we own, which is how they serve at up to 75 percent less than proprietary frontier APIs. Multi-turn workloads amplify that advantage because token spend scales with turn count, and cached input pricing cuts it further since resent history is exactly what cache rates are built for. Kimi K3's cached input runs at a tenth of its uncached rate, and a four-turn conversation spends most of its input tokens on history the model has already seen by turn three.
Matching model to workload is the same logic behind our efficient frontier approach to model selection, where every model on the platform earns its place for a specific job. The open-weight model guide covers why weights you can take anywhere change the evaluation math. The model catalog lists every hosted option with context lengths and capabilities.
Read the numbers as a starting distribution, not a census.
The practical response is the one this guide ends with: run the harness on your own traffic before you commit.
The harness is small. A scenario list, a loop that resends full history each turn, a few deterministic checks, and a JSON export for the judges. The inference call itself is the OpenAI-compatible quickstart pattern with the model swapped per candidate.
import os from openai import OpenAI client = OpenAI( base_url="https://api.telnyx.com/v2/ai/openai", api_key=os.environ["TELNYX_API_KEY"], ) # One line to swap candidates: MODEL = "moonshotai/Kimi-K3" # or zai-org/GLM-5.3-Flash, zai-org/GLM-5.3 response = client.chat.completions.create( model=MODEL, messages=messages, # full history resent on every turn reasoning_effort="low", temperature=0, max_tokens=4096, ) answer = response.choices[0].message.content
Change the MODEL string and rerun. No new vendor, no new bill, no new client. That one-line swap is why the evaluation loop in this guide is cheap to repeat whenever a new model lands.
For orchestration, a Stateful Actor is the natural home for a multi-turn harness because it persists progress across the run, which is how this benchmark ran. A scheduled Telnyx Function works for shorter eval sweeps on every model release.
For agents that build themselves around the platform, llms.txt exposes the docs as a machine-readable index. ai/pricing.json serves model rates as structured data. Your harness can fetch current rates instead of hardcoding them.
Single-turn scores got you a shortlist. Multi-turn behavior picks the model. Create an account and run your 15 scenarios against all three. For production volume on any of them, talk to our team about rates.
Related articles
Inference Infrastructure: What to Run Where and Why It Matters

What is an inference engine? Types, uses, and vLLM

Open-Source Models Are Catching Up to Frontier

The best WhatsApp API providers in 2026

WhatsApp Business API Cost in 2026 After October 1

VoIP infrastructure explained: find the layer behind every bad call
