Skip to content

MiMo V2.6 Pro vs Kimi K3 vs Qwen 3.8: Best Open-Weight AI Model (2026)

Matt Li By Matt Li Co-Founder and Director 12 min read
TL;DR: MiMo-V2.6-Pro is the best open-weight model on the Artificial Analysis Intelligence Index, scoring 46 against 44 for Kimi K3 and 40 for the open Qwen3.8 2.4T, and its API costs $0.435 per million input tokens and $0.87 per million output tokens. Kimi K3 wins LMArena's human-preference vote and long-context retrieval, but its output tokens cost $15 per million. Qwen3.8 Max scores 45, ahead of Kimi K3, yet it is Alibaba's closed API model, not the open weights.

Xiaomi printed the bill on its own launch page: $3.47 million of reinforcement learning for the whole MiMo-V2.6 series, 30 RL steps and 156.4 billion training tokens. Moonshot's Kimi K3 is 2.8 trillion parameters, and its official weights fill 1,561 GB on Hugging Face. Alibaba's largest open Qwen3.8 is bigger again on disk, at 4.9 TB in BF16.

What decides the pick
  1. 1On Terminal-Bench 4.0, MiMo-V2.6-Pro passes 34.8% of tasks and Kimi K3 passes 12.6%, in Artificial Analysis's own runs.
  2. 2MiMo ships under plain MIT. Kimi K3 and the big Qwen3.8 need a separate deal once an API reseller passes $20 million and $50 million in yearly revenue.
  3. 3Qwen3.8 27B is the only model here that fits one 24 GB GPU: its 4-bit file is 16.5 GB, under Apache 2.0.
  4. 4DeepSeek V4.1 Flash is the fastest of the five at 217 tokens a second, five times MiMo's measured speed.

Which Open-Weight Model Is Best Right Now?

MiMo-V2.6-Pro, by 1.5 points on the index most buyers quote. The Artificial Analysis open-weights leaderboard put it at 46.3 on September 29, 2026, ahead of GLM-5.3 at 44.8 and Kimi K3 at 43.6. The open Qwen3.8 2.4T A95B sits fifth at 39.9, 0.4 above DeepSeek V4.1 Flash at 39.5.

Bar chart of the Artificial Analysis Intelligence Index v4.3 for open-weights models: MiMo-V2.6-Pro 46.3, GLM-5.3 44.8, Kimi K3 43.6, GLM-5.3-Flash 41.8, Qwen3.8 2.4T A95B 39.9, DeepSeek V4.1 Flash 39.5, Qwen3.8 27B 33.7.

The index is version 4.3, which added AutomationBench and moved Terminal-Bench to version 4.0. Scores quoted from version 4.1.1 or earlier sit on a different scale and do not line up with these. The best closed model, Claude Opus 5.5 at maximum effort, scores 57.6, so the open field still trails the frontier by about 11 points.

Qwen3.8 Max (0902) scores 45.4, which would place it second. Artificial Analysis and LMArena both list it as proprietary, because it is the hosted product built on the open 2.4T model with vision and a non-thinking mode added. So "MiMo beats Qwen3.8 Max" is true, but Max is not the open model.

What this comparison rests on. I could not run my own prompt set against these five models for this piece, so every score below comes from a named leaderboard or the maker's model card, with the date I read it. Where a maker's card and an independent board disagree, both numbers are printed.

Specs Side by Side

All five are mixture-of-experts models except Qwen3.8 27B, and the active parameter count matters more for running cost than the headline size. Kimi K3 wakes 104 billion parameters per token; MiMo wakes 42 billion.

ModelTotal / activeContextLicenceAPI, per 1M in / outAA index
MiMo-V2.6-Pro (Xiaomi)1.02T / 42B1MMIT$0.435 / $0.8746.3
Kimi K3 (Moonshot)2.8T / 104B1,048,576Kimi K3 License$3 / $1543.6
Qwen3.8 2.4T A95B (Alibaba)2.4T / 95B262K native, 1.01M extendedQwen3.8-Max License$2 / $639.9
Qwen3.8 27B (Alibaba)27B dense262K native, 1M extendedApache 2.0$0.50 / $333.7
DeepSeek V4.1 Flash552B / 8B to 16B1MMIT$0.30 / $1.20 (peak)39.5
GLM-5.3 (Z.ai)753B / not stated1MGLM-5.3 License$1.40 / $4.4044.8

Parameters and licences are from each official Hugging Face card; GLM-5.3's card gives no total, so the table uses its safetensors count. Prices are the Artificial Analysis list prices, which match OpenRouter for five of the six rows; Qwen3.8 27B is the exception.

Proportional circles of total parameters: Kimi K3 2,800 billion, Qwen3.8 2.4T 2,400 billion, MiMo-V2.6-Pro 1,020 billion, GLM-5.3 753 billion, DeepSeek V4.1 Flash 552 billion.

The five shipped inside ten weeks, which is why rankings from July already read as history:

Jul 16, 2026
Kimi K3 launches on Moonshot's API, with full weights promised by July 27
Aug 3 to 14, 2026
Qwen3.8 Max goes live in the cloud, then the 2.4T A95B weights (Aug 12) and the 27B (Aug 14)
Aug 18, 2026
GLM-5.3 from Z.ai, the same base as GLM-5.2 with new post-training
Sep 10, 2026
DeepSeek V4.1 Flash, with a KV cache about four times smaller than V4 Flash
Sep 21 to 22, 2026
MiMo-V2.6 Pro and Flash, MIT-licensed, with weights and a technical report on day one

MiMo-V2.6-Pro: Top Score, Lowest Cost per Task

Xiaomi's MiMo-V2.6 launch page dated September 22, 2026, showing a coding benchmark curve over 30 RL steps and a total cost of $3.47M.
$0.13 per index task384 experts, 8 active42 tokens a second

MiMo-V2.6-Pro is for teams that want the top-scoring open model at a budget price. Artificial Analysis spent $0.13 per task running it through the whole index, against $2.00 for Kimi K3 and $2.16 for the open Qwen3.8.

The MiMo-V2.6-Pro-RL model card describes 70 layers, 60 of them sliding-window attention with a 128-token window and 10 global. A sliding-window layer caches only its last 128 tokens, so most of the stack stays small at 1M tokens. The model reads text, images, video and audio, and a five-layer speculative decoder drafts seven tokens per pass.

Xiaomi's own table puts it close to Claude Opus 5 on agent work: 71.9 on DeepSWE v1.1 against 74.0, and 53.1 on AutomationBench against 50.3. The gap reopens on Terminal-Bench 4.0, 34.9 against 49.0.

"In an era when compute is brutally scarce, we still chose to dedicate a team of several dozen people to one goal over an extended period: scaling up RL."

Fuli Luo, head of Xiaomi MiMo, quoted by VentureBeat, September 2026
  • Best for: agentic coding, terminal work and automation where the per-task bill matters.
  • Where it struggles: speed. Artificial Analysis measures 42 tokens a second and about 51 seconds before the first answer token, since it reasons at length first.
  • Price: $0.435 in, $0.87 out per million tokens, cache hits $0.0036. An UltraSpeed tier costs ten times as much.

Kimi K3: Largest Model, Best Human Votes

Moonshot's blog page headed Kimi K3: Open Frontier Intelligence, with a Try Kimi K3 button.
2.8T parameters, 104B active16 of 896 experts per token88.7% on AA-LCR

Kimi K3 is for chat-heavy products and long-document work, where it beats MiMo. Its model card calls it "the world's first open 3T-class model", and at 2.8 trillion parameters it is the largest of the five.

Moonshot trained it with quantization built in, so the released weights are already MXFP4 with MXFP8 activations. The attention is new too: 69 Kimi Delta Attention layers and 24 gated MLA layers. Moonshot's launch blog recommends supernode deployments of 64 or more accelerators.

The card reports 93.5 on GPQA Diamond and 91.2 on BrowseComp. Terminal-Bench 4.0 is its weak spot: Artificial Analysis scored it at 12.6%, and DeepSeek's own card measured the same 12.6. Demand arrived before capacity. Foreign Policy reported that Moonshot suspended new subscriptions within 48 hours of launch.

  • Best for: long-context retrieval, research and browsing agents, English chat quality.
  • Where it struggles: hard terminal tasks, and price. Artificial Analysis had no speed measurement for it on September 29.
  • Price: $3 in, $15 out per million tokens, cache hits $0.30, from Moonshot or OpenRouter. OpenRouter's batch tier is $2.28 and $11.40.

Qwen 3.8: The Open Weights Trail the Max API

Hugging Face model card for Qwen3.8-2.4T-A95B, showing 2.4T params in BF16 and the qwen3.8-max licence tag.
5.5 points between open and Max512 experts, 10 routed + 1 shared27B dense sibling, Apache 2.0

Qwen 3.8 is for teams that want a choice of sizes from one family. The flagship open release, Qwen3.8-2.4T-A95B, scores 39.9 on Artificial Analysis, against 45.4 for the hosted Qwen3.8 Max built on it.

Alibaba's card spells out what the open weights leave out: vision input, the non-thinking mode and the 1M default context all belong to Max. The weights run 262,144 tokens as trained and stretch to 1,010,000. Each block stacks three Gated DeltaNet linear-attention layers under one gated attention layer, 92 layers in all.

The open Qwen most teams can run is the small one. Qwen3.8 27B scores 33.7 at its highest reasoning setting, ships under Apache 2.0 and adds a vision encoder. Our guide to every Qwen model covers the older sizes.

  • Best for: a local model you can fine-tune (27B), or one vendor for cloud and self-hosted.
  • Where it struggles: the open 2.4T scores 11.1% on Terminal-Bench 4.0 and loses vision.
  • Price: $2 in, $6 out for both the 2.4T and Max on Qwen Cloud; $0.50 and $3 for the 27B.

DeepSeek V4.1 Flash and GLM-5.3

GLM-5.3 is the model the MiMo headline skips. At 44.8 it sits second on the open board, and on Artificial Analysis's Terminal-Bench 4.0 it scores 41.9%, the best of any open model here. Z.ai says all of its gains over GLM-5.2 came from post-training on the same base. It streams at 88 tokens a second, twice MiMo's rate. Our GLM model guide covers the earlier releases, from GLM-4.5 on.

DeepSeek V4.1 Flash trades score for speed. Its model card lists a 552B backbone that activates 8B to 16B parameters, and Artificial Analysis clocks it at 217 tokens a second with a 1.1-second first token. The card also puts it at 74.2 on DeepSWE v1.1, level with Claude Opus 5.0's 74.0.

DeepSeek's pricing page halves the $0.30 and $1.20 rates outside peak hours, which are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays. Our DeepSeek statistics page tracks its usage.

Which Model Wins Which Task

Terminal work separates the three by the widest margin. Artificial Analysis runs all five through one harness, and MiMo passes close to three times as many Terminal-Bench 4.0 tasks as Kimi K3 or the open Qwen. On long-context retrieval (AA-LCR) the order flips, and Kimi K3 leads at 88.7.

Dot plot of four Artificial Analysis evals. Terminal-Bench 4.0: MiMo 34.8, Kimi K3 12.6, Qwen3.8 2.4T 11.1. Humanity's Last Exam: 49.4, 46.9, 42.4. SciCode: 60.9, 59.5, 54.1. AA-LCR: 86.3, 88.7, 80.3.
  • Agentic coding and terminal tasks: GLM-5.3 (41.9%) and MiMo (34.8%) on Terminal-Bench 4.0, well clear of Kimi K3 and Qwen.
  • Reasoning: MiMo leads Humanity's Last Exam at 49.4%, Kimi K3 is at 46.9% and Qwen 2.4T at 42.4%.
  • Factual recall: Kimi K3 gets 47.6% of AA-Omniscience questions right, MiMo 34.9%.
  • Chinese and English: no maker publishes a like-for-like Chinese score for all three. MiMo's card lists both languages; DeepSeek's base model scores 92.1 on C-Eval.

Kimi K3's card reports 77.8 on ProgramBench and MiMo's reports 26.5. The DeepSeek and Z.ai cards score Kimi K3 at 17.5 on the almost-solved measure, so Moonshot's figure is on another scale. Three cards agree on Kimi K3's 67.5 on DeepSWE v1.1, the one vendor number that repeats. For Kimi's coding history, see Kimi K2 for coding.

Where LMArena Disagrees

People voting blind on LMArena's text leaderboard put Kimi K3 first of the open models, at 1,488 and rank 16 overall, on 26,400 votes. MiMo-V2.6-Pro scores 1,480 on only 4,026 votes, so its confidence band runs from 1,470 to 1,489.

Range chart of LMArena text ratings with confidence bands: Kimi K3 1483 to 1493, MiMo-V2.6-Pro 1470 to 1489, GLM-5.3 1474 to 1485, Qwen3.8 Max 1474 to 1485, DeepSeek V4.1 Flash 1469 to 1484.

MiMo, GLM-5.3 and Qwen3.8 Max overlap from 1,474 to 1,485, so LMArena cannot separate them yet. The two boards measure different things. Artificial Analysis grades finished tasks, while LMArena records which chat answer a person preferred. For a chatbot, weigh the arena; for an agent that runs code, weigh the index.

Price per Million Tokens and Where to Buy

Kimi K3's output tokens cost 17 times MiMo's, and the gap widens once you count the tokens each model spends. A per-token price hides how long a model thinks, so Artificial Analysis also reports what the full index run cost per task.

Scatter chart of index score against cost per index task: MiMo-V2.6-Pro 46.3 at $0.13, GLM-5.3 44.8 at $2.01, Kimi K3 43.6 at $2.00, GLM-5.3-Flash 41.8 at $0.25, Qwen3.8 2.4T 39.9 at $2.16, DeepSeek V4.1 Flash 39.5 at $0.27, Qwen3.8 27B 33.7 at $1.01.
ModelInput / output per 1MCache hit per 1MWhere to buy
MiMo-V2.6-Pro$0.435 / $0.87$0.0036Xiaomi MiMo API, OpenRouter
Kimi K3$3 / $15$0.30platform.kimi.ai, OpenRouter
Qwen3.8 2.4T / Max$2 / $6$0.25Qwen Cloud, OpenRouter
DeepSeek V4.1 Flash$0.30 / $1.20 (half off-peak)$0.006DeepSeek API, OpenRouter
GLM-5.3$1.40 / $4.40$0.26Z.ai, OpenRouter

Prices were read on September 29, 2026 and move often; OpenRouter lists cheaper batch tiers for Kimi K3, DeepSeek and GLM. Our US vs China AI model cost comparison sets these against GPT and Claude.

Hardware to Self-Host

Only one of these runs on a workstation. Everything above 500 GB needs a multi-GPU server before quantization, and Qwen3.8 2.4T in BF16 is almost 4.9 TB.

Column chart of official weights size on Hugging Face in GB: Qwen3.8 27B 56, DeepSeek V4.1 Flash 510, MiMo-V2.6-Pro 573, GLM-5.3 756, Kimi K3 1,561, Qwen3.8 2.4T 4,892.

Community quantizations cut those numbers but not to desktop size. The smallest files published by Unsloth and other quantizers:

16.5 GB
Qwen3.8 27B at 4-bit, one 24 GB GPU
236 GB
MiMo-V2.6-Pro at 2 bits per weight
217 GB
GLM-5.3 at 1-bit (IQ1_S)
466 GB
Kimi K3 at 1-bit (UD-Q1_0)
Source: file sizes on unsloth/Qwen3.8-27B-GGUF, AesSedai/MiMo-V2.6-Pro-RL-GGUF, unsloth/GLM-5.3-GGUF and unsloth/Kimi-K3-GGUF on Hugging Face, September 29, 2026.

For production, follow the makers' recipes. Xiaomi's vLLM command serves MiMo-V2.6-Pro with tensor parallelism across 8 GPUs, and its SGLang recipe spreads the model over two nodes. Moonshot points Kimi K3 at 64-accelerator supernodes. If you plan to adapt one of these, our list of platforms to fine-tune open-source LLMs compares hosted fine-tuning services.

Licences: What Commercial Use Allows

All five allow commercial use, but the custom licences switch on at different sizes of business. Moonshot's Kimi K3 License and Alibaba's Qwen3.8-Max License both exempt internal use that exposes nothing to third parties.

Standard licences
  • MiMo-V2.6-Pro: MIT
  • DeepSeek V4.1 Flash: MIT
  • Qwen3.8 27B: Apache 2.0
Custom licences
  • Kimi K3: a model-as-a-service business over $20M in 12 months needs a separate agreement
  • Qwen3.8 2.4T: same trigger at $50M, and it also covers AI coding and office assistants
  • GLM-5.3: a Z.ai security review, only for model-as-a-service firms above $10 billion

Kimi K3 and Qwen3.8 also require their names on screen in products above 100 million monthly users or $20 million in monthly revenue. LMArena lists GLM-5.3 as MIT, but Z.ai's repo carries its own GLM-5.3 License, so read the file, not the label.

Verdict by Use Case

If you needPickWhy
The best open model for agents and codingMiMo-V2.6-ProTop AA index score, $0.13 per task, MIT
Hard terminal and DevOps tasksGLM-5.341.9% on Terminal-Bench 4.0, 88 tokens a second
Chat quality and long documentsKimi K3LMArena open leader, 88.7 on AA-LCR
Speed at high volumeDeepSeek V4.1 Flash217 tokens a second, off-peak rates halved
A model on one GPUQwen3.8 27B16.5 GB at 4-bit, Apache 2.0
One vendor for cloud and self-hostingQwen 3.8Max API plus open 2.4T and 27B weights

For a wider view of the Chinese labs behind these five, see our roundup of Chinese open-source LLM leaders and the full Kimi model history, which runs from Kimi K1.5 to K3.

Hire Engineers Who Can Deploy These Models

Serving a 1-trillion-parameter model on your own GPUs, or wiring MiMo into an agent harness, takes engineers who have done it before. We place vetted LLM developers from Asia, and our LLM developer cost guide shows what a hire costs. Tell us what you are building and we will send a shortlist within 48 hours.

Frequently Asked Questions

What is the Artificial Analysis Intelligence Index?

A composite of 10 evaluations that Artificial Analysis runs itself, including Terminal-Bench 4.0, Humanity's Last Exam, SciCode, GDPval-AA and AA-LCR. Version 4.3 is current, and scores are not comparable across versions.

How does MiMo-V2.6-Flash compare with the Pro?

Flash has 310B total and 15B active parameters, costs $0.14 in and $0.28 out per million tokens, and scores 37.9 on the index against 46.3 for Pro. It is also MIT-licensed.

Is there a small MiMo model for a laptop?

Xiaomi released MiMo-V2.6-Distill-Qwen-9B alongside the series, a 9B model distilled onto a Qwen base, with GGUF builds from the community on Hugging Face.

Hire LLM engineers.

Pre-vetted senior engineers from Asia at 50 to 70% below US hiring costs, with first profiles in 24 hours.

Hire LLM engineers Apply as talent →
Matt Li

Written by

Matt Li is a tech-driven entrepreneur with deep expertise in global talent strategy, digital experience optimization, e-commerce, and Web3 innovation. He is the Co-Founder of Second Talent, a US-based company that connects businesses with top-tier tech professionals worldwide. Since launching the company in 2024, Matt has led its growth by leveraging technology to streamline remote hiring and scale distributed teams. With a background spanning product, operations, and innovation, Matt brings a cross-disciplinary perspective to the evolving digital economy. His work sits at the intersection of global talent, emerging technology, and scalable digital transformation.

More posts by Matt Li →

Loading available times…