TL;DR: MiMo-V2.6-Pro is the best open-weight model on the Artificial Analysis Intelligence Index, scoring 46 against 44 for Kimi K3 and 40 for the open Qwen3.8 2.4T, and its API costs $0.435 per million input tokens and $0.87 per million output tokens. Kimi K3 wins LMArena's human-preference vote and long-context retrieval, but its output tokens cost $15 per million. Qwen3.8 Max scores 45, ahead of Kimi K3, yet it is Alibaba's closed API model, not the open weights.
Xiaomi printed the bill on its own launch page: $3.47 million of reinforcement learning for the whole MiMo-V2.6 series, 30 RL steps and 156.4 billion training tokens. Moonshot's Kimi K3 is 2.8 trillion parameters, and its official weights fill 1,561 GB on Hugging Face. Alibaba's largest open Qwen3.8 is bigger again on disk, at 4.9 TB in BF16.
- 1On Terminal-Bench 4.0, MiMo-V2.6-Pro passes 34.8% of tasks and Kimi K3 passes 12.6%, in Artificial Analysis's own runs.
- 2MiMo ships under plain MIT. Kimi K3 and the big Qwen3.8 need a separate deal once an API reseller passes $20 million and $50 million in yearly revenue.
- 3Qwen3.8 27B is the only model here that fits one 24 GB GPU: its 4-bit file is 16.5 GB, under Apache 2.0.
- 4DeepSeek V4.1 Flash is the fastest of the five at 217 tokens a second, five times MiMo's measured speed.
Which Open-Weight Model Is Best Right Now?
MiMo-V2.6-Pro, by 1.5 points on the index most buyers quote. The Artificial Analysis open-weights leaderboard put it at 46.3 on September 29, 2026, ahead of GLM-5.3 at 44.8 and Kimi K3 at 43.6. The open Qwen3.8 2.4T A95B sits fifth at 39.9, 0.4 above DeepSeek V4.1 Flash at 39.5.

The index is version 4.3, which added AutomationBench and moved Terminal-Bench to version 4.0. Scores quoted from version 4.1.1 or earlier sit on a different scale and do not line up with these. The best closed model, Claude Opus 5.5 at maximum effort, scores 57.6, so the open field still trails the frontier by about 11 points.
Qwen3.8 Max (0902) scores 45.4, which would place it second. Artificial Analysis and LMArena both list it as proprietary, because it is the hosted product built on the open 2.4T model with vision and a non-thinking mode added. So "MiMo beats Qwen3.8 Max" is true, but Max is not the open model.
Specs Side by Side
All five are mixture-of-experts models except Qwen3.8 27B, and the active parameter count matters more for running cost than the headline size. Kimi K3 wakes 104 billion parameters per token; MiMo wakes 42 billion.
| Model | Total / active | Context | Licence | API, per 1M in / out | AA index |
|---|---|---|---|---|---|
| MiMo-V2.6-Pro (Xiaomi) | 1.02T / 42B | 1M | MIT | $0.435 / $0.87 | 46.3 |
| Kimi K3 (Moonshot) | 2.8T / 104B | 1,048,576 | Kimi K3 License | $3 / $15 | 43.6 |
| Qwen3.8 2.4T A95B (Alibaba) | 2.4T / 95B | 262K native, 1.01M extended | Qwen3.8-Max License | $2 / $6 | 39.9 |
| Qwen3.8 27B (Alibaba) | 27B dense | 262K native, 1M extended | Apache 2.0 | $0.50 / $3 | 33.7 |
| DeepSeek V4.1 Flash | 552B / 8B to 16B | 1M | MIT | $0.30 / $1.20 (peak) | 39.5 |
| GLM-5.3 (Z.ai) | 753B / not stated | 1M | GLM-5.3 License | $1.40 / $4.40 | 44.8 |
Parameters and licences are from each official Hugging Face card; GLM-5.3's card gives no total, so the table uses its safetensors count. Prices are the Artificial Analysis list prices, which match OpenRouter for five of the six rows; Qwen3.8 27B is the exception.

The five shipped inside ten weeks, which is why rankings from July already read as history:
MiMo-V2.6-Pro: Top Score, Lowest Cost per Task

MiMo-V2.6-Pro is for teams that want the top-scoring open model at a budget price. Artificial Analysis spent $0.13 per task running it through the whole index, against $2.00 for Kimi K3 and $2.16 for the open Qwen3.8.
The MiMo-V2.6-Pro-RL model card describes 70 layers, 60 of them sliding-window attention with a 128-token window and 10 global. A sliding-window layer caches only its last 128 tokens, so most of the stack stays small at 1M tokens. The model reads text, images, video and audio, and a five-layer speculative decoder drafts seven tokens per pass.
Xiaomi's own table puts it close to Claude Opus 5 on agent work: 71.9 on DeepSWE v1.1 against 74.0, and 53.1 on AutomationBench against 50.3. The gap reopens on Terminal-Bench 4.0, 34.9 against 49.0.
"In an era when compute is brutally scarce, we still chose to dedicate a team of several dozen people to one goal over an extended period: scaling up RL."
Fuli Luo, head of Xiaomi MiMo, quoted by VentureBeat, September 2026
- Best for: agentic coding, terminal work and automation where the per-task bill matters.
- Where it struggles: speed. Artificial Analysis measures 42 tokens a second and about 51 seconds before the first answer token, since it reasons at length first.
- Price: $0.435 in, $0.87 out per million tokens, cache hits $0.0036. An UltraSpeed tier costs ten times as much.
Kimi K3: Largest Model, Best Human Votes

Kimi K3 is for chat-heavy products and long-document work, where it beats MiMo. Its model card calls it "the world's first open 3T-class model", and at 2.8 trillion parameters it is the largest of the five.
Moonshot trained it with quantization built in, so the released weights are already MXFP4 with MXFP8 activations. The attention is new too: 69 Kimi Delta Attention layers and 24 gated MLA layers. Moonshot's launch blog recommends supernode deployments of 64 or more accelerators.
The card reports 93.5 on GPQA Diamond and 91.2 on BrowseComp. Terminal-Bench 4.0 is its weak spot: Artificial Analysis scored it at 12.6%, and DeepSeek's own card measured the same 12.6. Demand arrived before capacity. Foreign Policy reported that Moonshot suspended new subscriptions within 48 hours of launch.
- Best for: long-context retrieval, research and browsing agents, English chat quality.
- Where it struggles: hard terminal tasks, and price. Artificial Analysis had no speed measurement for it on September 29.
- Price: $3 in, $15 out per million tokens, cache hits $0.30, from Moonshot or OpenRouter. OpenRouter's batch tier is $2.28 and $11.40.
Qwen 3.8: The Open Weights Trail the Max API

Qwen 3.8 is for teams that want a choice of sizes from one family. The flagship open release, Qwen3.8-2.4T-A95B, scores 39.9 on Artificial Analysis, against 45.4 for the hosted Qwen3.8 Max built on it.
Alibaba's card spells out what the open weights leave out: vision input, the non-thinking mode and the 1M default context all belong to Max. The weights run 262,144 tokens as trained and stretch to 1,010,000. Each block stacks three Gated DeltaNet linear-attention layers under one gated attention layer, 92 layers in all.
The open Qwen most teams can run is the small one. Qwen3.8 27B scores 33.7 at its highest reasoning setting, ships under Apache 2.0 and adds a vision encoder. Our guide to every Qwen model covers the older sizes.
- Best for: a local model you can fine-tune (27B), or one vendor for cloud and self-hosted.
- Where it struggles: the open 2.4T scores 11.1% on Terminal-Bench 4.0 and loses vision.
- Price: $2 in, $6 out for both the 2.4T and Max on Qwen Cloud; $0.50 and $3 for the 27B.
DeepSeek V4.1 Flash and GLM-5.3
GLM-5.3 is the model the MiMo headline skips. At 44.8 it sits second on the open board, and on Artificial Analysis's Terminal-Bench 4.0 it scores 41.9%, the best of any open model here. Z.ai says all of its gains over GLM-5.2 came from post-training on the same base. It streams at 88 tokens a second, twice MiMo's rate. Our GLM model guide covers the earlier releases, from GLM-4.5 on.
DeepSeek V4.1 Flash trades score for speed. Its model card lists a 552B backbone that activates 8B to 16B parameters, and Artificial Analysis clocks it at 217 tokens a second with a 1.1-second first token. The card also puts it at 74.2 on DeepSWE v1.1, level with Claude Opus 5.0's 74.0.
DeepSeek's pricing page halves the $0.30 and $1.20 rates outside peak hours, which are 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays. Our DeepSeek statistics page tracks its usage.
Which Model Wins Which Task
Terminal work separates the three by the widest margin. Artificial Analysis runs all five through one harness, and MiMo passes close to three times as many Terminal-Bench 4.0 tasks as Kimi K3 or the open Qwen. On long-context retrieval (AA-LCR) the order flips, and Kimi K3 leads at 88.7.

- Agentic coding and terminal tasks: GLM-5.3 (41.9%) and MiMo (34.8%) on Terminal-Bench 4.0, well clear of Kimi K3 and Qwen.
- Reasoning: MiMo leads Humanity's Last Exam at 49.4%, Kimi K3 is at 46.9% and Qwen 2.4T at 42.4%.
- Factual recall: Kimi K3 gets 47.6% of AA-Omniscience questions right, MiMo 34.9%.
- Chinese and English: no maker publishes a like-for-like Chinese score for all three. MiMo's card lists both languages; DeepSeek's base model scores 92.1 on C-Eval.
Kimi K3's card reports 77.8 on ProgramBench and MiMo's reports 26.5. The DeepSeek and Z.ai cards score Kimi K3 at 17.5 on the almost-solved measure, so Moonshot's figure is on another scale. Three cards agree on Kimi K3's 67.5 on DeepSWE v1.1, the one vendor number that repeats. For Kimi's coding history, see Kimi K2 for coding.
Where LMArena Disagrees
People voting blind on LMArena's text leaderboard put Kimi K3 first of the open models, at 1,488 and rank 16 overall, on 26,400 votes. MiMo-V2.6-Pro scores 1,480 on only 4,026 votes, so its confidence band runs from 1,470 to 1,489.

MiMo, GLM-5.3 and Qwen3.8 Max overlap from 1,474 to 1,485, so LMArena cannot separate them yet. The two boards measure different things. Artificial Analysis grades finished tasks, while LMArena records which chat answer a person preferred. For a chatbot, weigh the arena; for an agent that runs code, weigh the index.
Price per Million Tokens and Where to Buy
Kimi K3's output tokens cost 17 times MiMo's, and the gap widens once you count the tokens each model spends. A per-token price hides how long a model thinks, so Artificial Analysis also reports what the full index run cost per task.

| Model | Input / output per 1M | Cache hit per 1M | Where to buy |
|---|---|---|---|
| MiMo-V2.6-Pro | $0.435 / $0.87 | $0.0036 | Xiaomi MiMo API, OpenRouter |
| Kimi K3 | $3 / $15 | $0.30 | platform.kimi.ai, OpenRouter |
| Qwen3.8 2.4T / Max | $2 / $6 | $0.25 | Qwen Cloud, OpenRouter |
| DeepSeek V4.1 Flash | $0.30 / $1.20 (half off-peak) | $0.006 | DeepSeek API, OpenRouter |
| GLM-5.3 | $1.40 / $4.40 | $0.26 | Z.ai, OpenRouter |
Prices were read on September 29, 2026 and move often; OpenRouter lists cheaper batch tiers for Kimi K3, DeepSeek and GLM. Our US vs China AI model cost comparison sets these against GPT and Claude.
Hardware to Self-Host
Only one of these runs on a workstation. Everything above 500 GB needs a multi-GPU server before quantization, and Qwen3.8 2.4T in BF16 is almost 4.9 TB.

Community quantizations cut those numbers but not to desktop size. The smallest files published by Unsloth and other quantizers:
For production, follow the makers' recipes. Xiaomi's vLLM command serves MiMo-V2.6-Pro with tensor parallelism across 8 GPUs, and its SGLang recipe spreads the model over two nodes. Moonshot points Kimi K3 at 64-accelerator supernodes. If you plan to adapt one of these, our list of platforms to fine-tune open-source LLMs compares hosted fine-tuning services.
Licences: What Commercial Use Allows
All five allow commercial use, but the custom licences switch on at different sizes of business. Moonshot's Kimi K3 License and Alibaba's Qwen3.8-Max License both exempt internal use that exposes nothing to third parties.
- MiMo-V2.6-Pro: MIT
- DeepSeek V4.1 Flash: MIT
- Qwen3.8 27B: Apache 2.0
- Kimi K3: a model-as-a-service business over $20M in 12 months needs a separate agreement
- Qwen3.8 2.4T: same trigger at $50M, and it also covers AI coding and office assistants
- GLM-5.3: a Z.ai security review, only for model-as-a-service firms above $10 billion
Kimi K3 and Qwen3.8 also require their names on screen in products above 100 million monthly users or $20 million in monthly revenue. LMArena lists GLM-5.3 as MIT, but Z.ai's repo carries its own GLM-5.3 License, so read the file, not the label.
Verdict by Use Case
| If you need | Pick | Why |
|---|---|---|
| The best open model for agents and coding | MiMo-V2.6-Pro | Top AA index score, $0.13 per task, MIT |
| Hard terminal and DevOps tasks | GLM-5.3 | 41.9% on Terminal-Bench 4.0, 88 tokens a second |
| Chat quality and long documents | Kimi K3 | LMArena open leader, 88.7 on AA-LCR |
| Speed at high volume | DeepSeek V4.1 Flash | 217 tokens a second, off-peak rates halved |
| A model on one GPU | Qwen3.8 27B | 16.5 GB at 4-bit, Apache 2.0 |
| One vendor for cloud and self-hosting | Qwen 3.8 | Max API plus open 2.4T and 27B weights |
For a wider view of the Chinese labs behind these five, see our roundup of Chinese open-source LLM leaders and the full Kimi model history, which runs from Kimi K1.5 to K3.
Hire Engineers Who Can Deploy These Models
Serving a 1-trillion-parameter model on your own GPUs, or wiring MiMo into an agent harness, takes engineers who have done it before. We place vetted LLM developers from Asia, and our LLM developer cost guide shows what a hire costs. Tell us what you are building and we will send a shortlist within 48 hours.
Frequently Asked Questions
What is the Artificial Analysis Intelligence Index?
A composite of 10 evaluations that Artificial Analysis runs itself, including Terminal-Bench 4.0, Humanity's Last Exam, SciCode, GDPval-AA and AA-LCR. Version 4.3 is current, and scores are not comparable across versions.
How does MiMo-V2.6-Flash compare with the Pro?
Flash has 310B total and 15B active parameters, costs $0.14 in and $0.28 out per million tokens, and scores 37.9 on the index against 46.3 for Pro. It is also MIT-licensed.
Is there a small MiMo model for a laptop?
Xiaomi released MiMo-V2.6-Distill-Qwen-9B alongside the series, a 9B model distilled onto a Qwen base, with GGUF builds from the community on Hugging Face.





