Built on Mistral's sparse mixture-of-experts architecture that routes each token through 2 of 8 expert networks, this Nous Research fine-tune keeps only 12.9B of its 46.7B total parameters active per forward pass. The DPO alignment stage improved its MT-Bench score over the base Mixtral Instruct while preserving the model's efficiency advantage over dense 70B-class alternatives.
Discover the power and diversity of large language models available with Telnyx. Explore the options below to find the perfect model for your project.
Check out our helpful tools to help get you started.
Nous Hermes 2 Mixtral 8x7B DPO is a mixture-of-experts model from Nous Research, built on Mistral's Mixtral 8x7B architecture with DPO training. It features 56 billion total parameters with only a subset active per token, providing strong performance at efficient inference costs.
Nous Hermes 2 Mixtral 8x7B performs admirably on translation and complex topic understanding, and outperforms GPT-4 in certain areas like puzzles and roleplay. For general-purpose tasks, GPT-4 remains stronger overall. The model is available on multiple inference platforms as a free open-source alternative.
Nous Hermes 2 Mixtral 8x7B DPO scores 72.3% on MMLU, slightly above the base Mixtral 8x7B Instruct (70.6%) on the same sheet after DPO alignment. Its Arena ELO of 1,084 is about 30 points below the base Mixtral Instruct (1,114), a common pattern where DPO improves knowledge scores but can shift chat preference. It matches GPT-3.5 Turbo (70.0% MMLU) at the open-weight tier.
The cost of running the model with Telnyx Inference is $0.0003 per 1,000 tokens. To put this into perspective, analyzing 1,000,000 customer chats, assuming each chat is 1,000 tokens long, would cost $300.
| Organization | Model Name | Tasks | Languages Supported | Context Length | Parameters | Model Tier | License |
|---|---|---|---|---|---|---|---|
| Telnyx | DeepSeek-V4-Flash-0731 | text generation | English | 1,048,576 | 284.0B | unlisted | mit |
| Telnyx | DeepSeek-V4.1-Flash | text generation | English | 1,048,576 | 552.0B | unlisted | mit |
| gemma-2b-it | text generation | English | 8,192 | 2.5B | small | gemma | |
| gemma-4-26B-A4B-it | text to-text | Multilingual | 262,144 | 25.8B | small | apache-2.0 | |
| gemma-4-31B-it | text generation | Multilingual | 131,072 | 30.7B | unlisted | apache-2.0 | |
| meta-llama | Llama-3.3-70B-Instruct | text generation | Multilingual | 99,000 | 70.6B | large | llama3.3 |
| meta-llama | Meta-Llama-3.1-70B-Instruct | text generation | Multilingual | 99,000 | 70.6B | large | llama3.1 |
| meta-llama | Meta-Llama-3.1-8B-Instruct | text generation | Multilingual | 131,072 | 8.0B | small | llama3.1 |
| minimaxai | MiniMax-M2.7 | text generation | Multilingual | 200,000 | 0 | large | minimaxai |
| minimaxai | MiniMax-M3-MXFP8 | text generation | Multilingual | 1,000,000 | 428.0B | large | minimax-community |
| moonshotai | Kimi-K2.5 | text generation | Multilingual | 256,000 | 1.0T | large | modified-mit |
| moonshotai | Kimi-K2.6 | text generation | Multilingual | 262,144 | 1.0T | large | modified-mit |
| moonshotai | Kimi-K3 | text generation | English | 1,000,000 | 2.8M | large | MIT Modified |
| Qwen | Qwen3-235B-A22B | text generation | English | 32,768 | 235.1B | large | apache-2.0 |
| Telnyx | Qwen3.8-27B | text generation | English | 262,144 | 27.0B | unlisted | Apache 2.0 |
| zai-org | GLM-5.1-FP8 | text generation | Multilingual | 202,752 | 753.9B | large | mit |
| zai-org | GLM-5.2 | text generation | English | 1,000,000 | 753.9B | large | MIT |
| Telnyx | GLM-5.3 | text generation | English | 1,048,576 | 753.0B | unlisted | glm-5.3 |
| zai-org | GLM-5.3-Flash | text generation | English | 1,048,576 | 320.0B | unlisted | mit |
| anthropic | claude-haiku-4-5 | text generation | Multilingual | 400,000 | 0 | large | openai |
| gemini-2.5-flash | text generation | Multilingual | 400,000 | 0 | large | openai | |
| gemini-3.7-flash | text generation | Multilingual | 400,000 | 0 | large | openai | |
| groq | gpt-oss-120b | text generation | Multilingual | 400,000 | 0 | large | openai |
| openai | gpt-4.1 | text generation | Multilingual | 400,000 | 0 | large | openai |
| openai | gpt-4o | text generation | Multilingual | 400,000 | 0 | large | openai |
| openai | gpt-4o-mini | text generation | Multilingual | 400,000 | 0 | large | openai |
| openai | gpt-5 | text generation | Multilingual | 400,000 | 0 | large | openai |
| openai | gpt-5.1 | text generation | Multilingual | 400,000 | 0 | large | openai |
| openai | gpt-5.2 | text generation | Multilingual | 400,000 | 0 | large | openai |
| openai | gpt-5.4 | text generation | Multilingual | 400,000 | 0 | large | openai |
| openai | gpt-5.4-mini | text generation | Multilingual | 400,000 | 0 | large | openai |
| openai | gpt-5.6-luna | text generation | Multilingual | 400,000 | 0 | large | openai |
| openai | gpt-5.6-sol | text generation | Multilingual | 400,000 | 0 | large | openai |
| openai | gpt-live-1 | speech to-speech | Multilingual | 400,000 | 0 | large | openai |
The model excels at multilingual reasoning, content generation, and conversational AI with a 32K context window. Its DPO training gives it strong instruction-following capabilities and improved response quality compared to the base Mixtral model.
Yes, Nous Hermes 2 Mixtral 8x7B DPO is open-source and free to use. It is available on Hugging Face and through inference providers like Ollama and OpenRouter.