Last updated: September 28, 2026. Open model versions change monthly, so every pick below carries a release date and states whether its benchmark score is vendor-reported or independently measured.

Takeaways
Three popular rankings of open source LLMs name three different versions of Kimi and GLM, and none of them says which scores were measured independently. This guide fixes that. Each pick below lists its parameters, license, and release date, and it labels every benchmark by who ran it. It also covers the workload other rankings skip: agents that talk to people on a live call.
The best open source LLM for most teams in September 2026 is Kimi K3 for maximum capability, or DeepSeek-V4.1-Flash for the best capability per dollar. On the Vals Index, an independent evaluation suite, DeepSeek-V4.1-Flash scored 57.86% and Kimi K3 scored 57.81%, inside the margin of error of each other. DeepSeek-V4.1-Flash did it at $0.30 per test against $6.47 for Kimi K3.
The right pick still depends on the workload. The table below maps each model to the job it does best.
Best open source LLMs in 2026 by use case
| Model | Best for | Main trade-off |
|---|---|---|
| Kimi K3 | Highest-ceiling agentic coding | 2.8T parameters; too large to self-host for most teams |
| DeepSeek-V4.1-Flash | Best value for coding agents | Headline scores are vendor-reported; independent runs are lower |
| GLM-5.3 and GLM-5.3-Flash | Long-horizon agents | Flagship license requires a review for the largest hosts |
| Qwen3.8-27B | Dense model on one GPU | Smaller knowledge base than frontier MoE models |
| Gemma 4 26B A4B | Running locally | Weaker on agentic coding than larger picks |
| MiniMax M3 | Low-cost multimodal agents | License requires attribution and approval at scale |
| Kimi K2.6 | Real-time voice agents | Superseded by Kimi K3 on raw capability |
Six factors set the order: task performance on the benchmark closest to the workload, developer fit (published weights, vLLM and Ollama support, OpenAI-compatible APIs), hardware reality, license freedom, cost per task, and response speed.
Every score in this guide carries one of two labels. Vendor-reported means the model's own lab ran it, usually with its own test setup. Independent means a third party such as Vals AI ran it. The gap between the two can be large: DeepSeek reports 90.6 on Terminal-Bench 2.1, while Vals measured 74.53% across three full trials.
Which benchmarks still separate models? SWE-bench Verified is close to saturated. The coding benchmarks that still separate frontier open models are SWE-bench Pro, FrontierSWE, DeepSWE v1.1, and Terminal-Bench 2.1 and 3.0. Rank by the benchmark closest to the actual workload.
Moonshot AI's Kimi K3 is the largest open-weight model available. It is a 2.8 trillion parameter mixture-of-experts (MoE) model that routes each token to 16 of its 896 experts, activating about 104 billion parameters. It has a 1M-token context window and native text, image, and video input. Weights shipped on July 27, 2026.
Moonshot reports 88.3 on Terminal-Bench 2.1 and 81.2 on FrontierSWE (vendor-reported). Independently, Vals AI puts Kimi K3 first among open-weight models on Vibe Code Bench, 0.2 points ahead of DeepSeek-V4.1-Flash.
Choose Kimi K3 if the workload is long, autonomous coding or research where the best possible answer justifies slower, costlier runs.
DeepSeek released V4.1 Flash on September 10, 2026 under the MIT license. It has a 552B-parameter backbone, a 1M-token context window, up to 384K output tokens, and native image input.
DeepSeek reports 90.6 on Terminal-Bench 2.1, 74.2 on DeepSWE v1.1, and a 3471 Codeforces rating (vendor-reported). Vals measured 74.53% on Terminal-Bench 2.1 (independent), second among open-weight models. On Vibe Code Bench, Vals found it about 40 times cheaper per task than Kimi K3 and about five times faster.
Choose DeepSeek-V4.1-Flash if the workload runs many agent loops a day and cost per task decides the budget.
Z.ai launched GLM-5.3 by API on August 14, 2026, then released the weights on August 28 after a two-week safety hold. The flagship keeps the same roughly 753B-parameter base as GLM-5.2 with a 1M-token context. According to the model card, it lifted DeepSWE v1.1 from 46.2 to 66.9 (vendor-reported). Serving speed varies by provider: in Telnyx's own GLM-5.3 latency benchmarks against three other hosts, Telnyx had the lowest median end-to-end latency in five of six workloads.
GLM-5.3-Flash, released August 26 under the MIT license, is the more practical model for most teams. It is a 320B MoE with 18B active parameters and a hybrid linear-plus-sparse attention design that Z.ai says cuts KV cache 4.44 times compared with the flagship.
Choose GLM-5.3-Flash if an agent needs a 1M-token context at a low serving cost under a clean license.
Alibaba released Qwen3.8-27B on August 14, 2026 under Apache 2.0. It is a dense model with about 27.8B parameters, a 262K-token native context, and image and video input. It replaces Qwen3.6-27B, which some rankings still list as current.
Qwen reports 61.7% on SWE-bench Pro (vendor-reported). At 4-bit precision, the weights take roughly 14 GB, leaving room for the KV cache on a 24 GB consumer GPU.
Choose Qwen3.8-27B if the team wants a capable coding and document model it can run on one GPU with no license questions.
Google's Gemma 4 26B A4B is an MoE model with 25.2B total and 3.8B active parameters, a 256K-token context window, and text and image input. Gemma 4 moved to Apache 2.0, dropping the custom terms that covered Gemma 3. It supports over 140 languages.
Because only 3.8B parameters run per token, it generates text far faster than its total size suggests. It fits on laptops with 32 GB of unified memory and on a single 24 GB GPU once quantized.
Choose Gemma 4 26B A4B if data must stay on a local machine or an edge device.
MiniMax M3, released June 1, 2026, is a natively multimodal MoE with about 428B total and 23B active parameters. It accepts text, images, and video with a 1M-token context and targets agentic coding and tool use.
Choose MiniMax M3 if the goal is multimodal agent work at the lowest cost per token, and the license terms fit the product.
Moonshot released Kimi K2.6 in April 2026. It is a 1T-parameter MoE with 32B active parameters and a 256K-token context, under a Modified MIT license. Kimi K3 beats it on capability, but K2.6 keeps a lower time to first token, which matters more on a live call.
On Telnyx Inference, Kimi K2.6 is the default model verified for voice AI assistants. Reasoning is always disabled on voice calls, so the model answers directly instead of thinking first.
"I heavily encourage people to try one of the open source models first, the recommended Kimi K2.5, 2.6, because it's just… it's good, and it's super fast, you're not going to get anything that fast with OpenAI or Anthropic." James Whedbee, Telnyx |
Choose Kimi K2.6 if a caller is waiting on the line and every millisecond of silence counts.
Readers who want Kimi K3 or GLM-5.3 without building a multi-GPU cluster can call them through the OpenAI-compatible Telnyx Inference API by changing a base URL and model name.
Run Kimi K3 and 7 other open models on Telnyx InferenceOpenAI-compatible API, GPUs colocated with the telephony network, in-region serving. See pricing and available models
Start buildingAn open source LLM is a large language model released under terms that let anyone use, study, modify, and share it. The Open Source AI Definition from the Open Source Initiative (OSI) sets a high bar: it requires the model parameters, the complete code used to train and run the model, and detailed information about the training data.
Almost no leading model meets that bar. Projects such as OLMo and Pythia publish the full recipe, but the models at the top of the benchmarks, including Kimi, DeepSeek, GLM, Qwen, Gemma, and Llama, publish their weights and little else. That makes them open weight. Our guide to open weight models covers the distinction in more depth.
For a builder, the license matters more than the label. Open models fall into three license groups:
Closed models such as GPT and Claude publish nothing and are reachable only through the vendor's API.
The top of the open model rankings changed several times in two months. Each entry below is dated so the list stays useful as a reference.
Teams already in production need a migrate-or-stay answer, not a new ranking:
The best open source LLM for a given product comes down to three checks: the license, the hardware, and the cost of the path to production. A model that scores two points higher but carries a confusing commercial license may be the worse choice.
Before shipping, confirm each of these against the model's license file:
Match the model to the memory available:
A 1M-token context window multiplies KV-cache memory per request, so pick the smallest window that fits the task.
Self-hosting trades per-token fees for GPU capital and operations work. A vendor's own API is the fastest start but ties the product to one provider. A hosted OpenAI-compatible endpoint sits between the two: open weights, no GPUs to manage, and the option to switch models by changing one parameter.
Cached input changes the math. On Telnyx Inference pricing, cached input for Kimi K3 costs $0.27 per 1M tokens, one-tenth of the uncached rate. Log cache hits and misses separately to see the real bill.
"My feed is full of people proudly showing off their '10 Billion Token' trophies. What I see: 'I just paid OpenAI a premium for something I could've run with open-source models at 90% less cost.'" David Casem, CEO at Telnyx |
For most teams, a hybrid works best: small models locally for privacy-sensitive work, and frontier open models through a hosted endpoint.
For a real-time voice agent, the best open source LLM is the one with the lowest time to first token (TTFT) that still handles function calls reliably. Today that is Kimi K2.6, with fast MoE models such as GLM-5.3-Flash and DeepSeek-V4.1-Flash as options for text-side agent work.
A voice turn runs in order: speech-to-text (STT), then the LLM, then text-to-speech (TTS). The caller hears silence until the LLM's first token reaches TTS. Every millisecond of TTFT is dead air, so a model that answers fast beats one that is marginally more accurate but slower. Our breakdown of inference latency shows where those milliseconds go.
MoE models help because only a fraction of their parameters run per token. Gemma 4 26B A4B activates 3.8B, GLM-5.3-Flash 18B, and Kimi K2.6 32B.
Voice selection checklist: time to first token, streaming tokens per second, reliable function calling for transfers and lookups, structured output, and recovery when the first plan fails. No public benchmark measures these end to end, so test on real calls.
"The question is how much of the production system you want to own and debug yourself. Once you add streaming STT, an LLM, TTS, tool calls, observability, failover, latency tuning, transfers, voicemail handling, and carrier deliverability, you are no longer just wrapping an LLM around a call. You are building a real-time distributed voice system." James Whedbee, VP of Engineering at Telnyx |
Telnyx Voice AI Agents run STT, TTS, orchestration, and inference on one network, with Kimi K2.6 as the default voice model. Teams with their own model can also bring any OpenAI-compatible LLM to a Telnyx assistant.
To check how fast a model will answer before putting it on a call, the script below streams a short phone-style reply from Kimi K2.6 and times the first token. It needs Python 3.8 or later, pip install openai, and a Telnyx API key exported as TELNYX_API_KEY.
import os import time from openai import OpenAI client = OpenAI( api_key=os.environ["TELNYX_API_KEY"], base_url="https://api.telnyx.com/v2/ai/openai", ) start = time.perf_counter() stream = client.chat.completions.create( model="moonshotai/Kimi-K2.6", messages=[ {"role": "system", "content": "You are a phone agent. Reply in one short sentence."}, {"role": "user", "content": "Can I move my appointment to Friday?"}, ], max_tokens=80, stream=True, ) first_token = None reply = [] for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: if first_token is None: first_token = time.perf_counter() - start reply.append(chunk.choices[0].delta.content) print(f"Time to first token: {first_token:.3f}s") print("".join(reply))
Run it from the region where callers live. Swap the model value to compare candidates on the same prompt.
After picking a model, a builder depends on the ecosystem around it: where the weights live, which tools serve them, and how the community adapts them.
Hugging Face is the main distribution point for open weights. Each model card is the source of truth for license, parameter count, context length, and the lab's own benchmark scores. Read the license file itself, not a summary of it; some coverage still describes Kimi K3 as Modified MIT when it ships under its own license.
Because vLLM, SGLang, and hosted endpoints all speak the OpenAI API format, the same client code works everywhere. A team can prototype with Gemma 4 on a workstation and move to Kimi K3 on a hosted endpoint by changing a base URL and a model name.
Community practice fills in the rest: 4-bit quantization for consumer GPUs, LoRA fine-tuning on domain data, distillation from MIT-licensed outputs, and pairing a model with a vector database for retrieval over company documents.
Picking a model is half the work. The other half is serving it fast enough, close enough to users, and at a price that holds at scale. Telnyx Inference hosts Kimi K3, Kimi K2.6, GLM-5.3, GLM-5.3-Flash, DeepSeek-V4.1-Flash, MiniMax M3, and Qwen3.8-27B on GPUs Telnyx owns, colocated with its global telephony network. The API is OpenAI-compatible, inference runs in-region, and the same platform carries the call when the agent needs to talk.
Run the best open source LLMs on Telnyx InferenceOpenAI-compatible API, in-region serving, GPUs colocated with the telephony network.
Start building with Telnyx InferenceRelated articles
vLLM inference: how it works and when to stop self-hosting

AI DLP Test: Six Data Types, Six Blocks, $0 Charged

Open-Source Models Are Catching Up to Frontier
The 11 best conversational AI platforms in 2026
vLLM inference: how it works and when to stop self-hosting

Activate your eSIM manually with an SM-DP+ address
