The Open-Weight League
The weekly form guide to open-weight models worth running on your own infrastructure.
INT Intelligence · COD Coding · SPD Speed t/s · OPEN Openness / 18 · PTS Blend 40·25·15·20
Season 2026/27 · Matchday 03| # | Club | Nation | INTIntelligence Index · reasoning, knowledge & agentic ability · AA Index v4.1 | CODCoding Index · software engineering & code generation benchmarks | SPDOutput Speed · median generation throughput, tokens per second | OPENOpenness Index · weights & transparency, scored out of 18 | PTSLeague Points · blended score · INT 40 · COD 25 · SPD 15 · OPEN 20 |
|---|---|---|---|---|---|---|---|
| 1 | Kimi K3Moonshot AI | CHN | 57 | 76 | 36 | 7 | –74 |
| 2 | GLM-5.2Z.AI · Zhipu AI | CHN | 51 | 69 | 111 | 8 | –72 |
| 3 | Nemotron 3 UltraNVIDIA | USA | 38 | 49 | 193 | 15 | –67 |
| 4 | DeepSeek V4 ProDeepSeek | CHN | 44 | 59 | 65 | 9 | –63 |
| 5 | MiniMax M3MiniMax | CHN | 44 | 59 | 69 | 6 | ▲160 |
| 6 | MiMo V2.5 ProXiaomi | CHN | 42 | 60 | 50 | 7 | ▲259 |
| 7 | SFStep 3.7 FlashStepFun | CHN | 30 | — | 393 | 7 | ▲258 |
| 8 | InklingThinking Machines | USA | 41 | 52 | 85 | 7 | ▼157 |
| 9 | NXNex-N2-ProNex AGI | CHN | 41 | — | 127 | 7 | ▲155 |
| 10 | Nemotron 3 SuperNVIDIA | USA | 25 | — | 165 | 15 | ▲154 |
| 11 | MiMo V2.5Xiaomi | CHN | 37 | — | 91 | 7 | ▲250 |
| 12 | Kimi K2.7 CodeMoonshot AI | CHN | 42 | — | 50 | 5 | –49 |
| 13 | Qwen3.6 27BAlibaba | CHN | 37 | — | 55 | 7 | ▲248 |
| 14 | Qwen3.6 35B A3BAlibaba | CHN | 32 | — | 141 | 7 | –47 |
| 15 | Qwen3.5 122B A10BAlibaba | CHN | 32 | — | 136 | 7 | ▲247 |
| 16 | HNHyperNova 60BMultiverse Computing | ESP | 18 | — | 375 | 7 | –46 |
| 17 | DeepSeek V4 FlashDeepSeek | CHN | 29 | — | 107 | 9 | ▼1246 |
| 18 | Mistral Medium 3.5Mistral AI | FRA | 30 | 47 | 66 | 6 | –46 |
| 19 | Qwen3.5 397B A17BAlibaba | CHN | 34 | — | 69 | 7 | –46 |
| 20 | RGRing-2.6-1TInclusionAI · Ant Group | CHN | 31 | — | 122 | 7 | –46 |
Key — what the columns mean
INT
Intelligence Index. Reasoning, knowledge & agentic ability · AA Index v4.1
COD
Coding Index. Software engineering & code generation benchmarks
SPD
Output Speed. Median generation throughput · tokens per second
OPEN
Openness Index. Weights & transparency, scored out of 18
PTS
League Points. Blended score · INT 40 · COD 25 · SPD 15 · OPEN 20
Just dropped — not yet scored
DeepSeek V4 Flash 0731 · INT 50 · COD 69 · released today, API-only — speed & weights pending. Eligible for a provisional seat (†) the moment Artificial Analysis benchmarks a provider. Until then the V4 Flash club is scored on its open checkpoint below.
Methodology
Every underlying score on this table comes from Artificial Analysis, an independent AI benchmarking organisation. We do not run our own evaluations and we do not adjust theirs.
INT is the Artificial Analysis Intelligence Index (v4.1): reasoning, knowledge and agentic ability. COD is their Coding Index across software-engineering and code-generation benchmarks. SPD is median output speed in tokens per second. OPEN is their Openness Index — weights availability, licence and transparency — scored out of 18.
PTS is the only number we compute. INT, COD and SPD are scored relative to the matchday’s best (best = 100); OPEN is scored against its fixed 18-point maximum. The four are then blended INT 40% · COD 25% · SPD 15% · OPEN 20% and rounded to whole points — ties are ordered by the unrounded blend. The method is stated here precisely so the table can be checked, argued with, or rebuilt by anyone.
A dash means Artificial Analysis has not published that score yet. A club missing one column is still seated: that column’s weight is redistributed pro-rata across its published scores. A club missing two or more isn’t seated until the data exists.
† marks a provisional seat. A newly released model may be seated before its weights are public when its lab has publicly committed the weights or has an established open-weights record, and Artificial Analysis has published every column except openness. Three matchdays maximum — the weights land or the seat lapses. This rule seated Kimi K3 at launch; its weights shipped on schedule and the mark came off.
Valarian builds no models and holds no stake in any lab on this table. ACRA runs whichever club you pick — so we have no reason to favour one. The table exists because our customers keep asking the same question: which open-weight models are actually good now?
Valarie’s squad
Match-fit today: validated in Valarie, uploaded privately into ACRA, and served from infrastructure you control.
DeepSeek V4 Pro
DeepSeek · CHN
live in ACRADeploy your stack
gpt-oss-120b
OpenAI · USA
live in ACRADeploy your stack
Gemma 4 31B
Google · USA
live in ACRADeploy your stack
Any open-weight club is loadable · Control Plane → Private AI → upload