Best open source LLM rankings

150 open-weights models you can self-host or fine-tune, ranked on the same LLM Index Score as everything else — 7 of them carry a computed score. Weights are downloadable for every model listed here.

150 models
#ModelLLM Index ScoreoursReasoningCodingAgentic CodingMathematicsData AnalysisLanguageIFPrice/M inContext
1Z.ai: GLM 5.2
z-ai
73.8478.679.751.989.873.776.262.3$0.9521.0M
2DeepSeek: DeepSeek V4 Pro
deepseek
72.2682.770.042.690.774.578.162.4$0.4351.0M
3MoonshotAI: Kimi K2.6
moonshotai
71.0679.478.646.984.365.175.164.4$0.684262K
4MoonshotAI: Kimi K2.7 Code
moonshotai
69.2682.874.045.779.662.777.956.3$0.820262K
5MiniMax: MiniMax M3
minimax
67.6374.568.240.777.076.276.857.5$0.3001.0M
6DeepSeek: DeepSeek V4 Flash
deepseek
65.6270.669.237.679.768.070.163.1$0.0941.0M
7Qwen: Qwen3.6 27B
qwen
64.9170.371.839.379.970.463.353.2$0.450262K
Meituan: LongCat 2.0
meituan
not evaluated$0.3001.0M
Thinking Machines: Inkling
thinkingmachines
not evaluated$1.001.0M
Tencent: Hy3
tencent
not evaluated$0.140262K
Poolside: Laguna XS 2.1
poolside
not evaluated$0.060262K
Poolside: Laguna XS 2.1 (free)
poolside
not evaluatedFree262K
Nex AGI: Nex-N2-Mini
nex-agi
not evaluated$0.025262K
Cohere: North Mini Code (free)
cohere
not evaluatedFree256K
Nex AGI: Nex-N2-Pro
nex-agi
not evaluated$0.250262K
NVIDIA: Nemotron 3.5 Content Safety (free)
nvidia
not evaluatedFree128K
NVIDIA: Nemotron 3 Ultra
nvidia
not evaluated$0.6001M
NVIDIA: Nemotron 3 Ultra (free)
nvidia
not evaluatedFree1M
StepFun: Step 3.7 Flash
stepfun
not evaluated$0.200256K
IBM: Granite 4.1 8B
ibm-granite
not evaluated$0.050131K
NVIDIA: Nemotron 3 Nano Omni (free)
nvidia
not evaluatedFree256K
Poolside: Laguna M.1
poolside
not evaluated$0.200262K
Poolside: Laguna M.1 (free)
poolside
not evaluatedFree262K
Qwen: Qwen3.6 35B A3B
qwen
not evaluated$0.140262K
Tencent: Hy3 preview
tencent
not evaluated$0.063262K
Xiaomi: MiMo-V2.5-Pro
xiaomi
not evaluated$0.4351.0M
Xiaomi: MiMo-V2.5
xiaomi
not evaluated$0.1401.0M
Z.ai: GLM 5.1
z-ai
not evaluated$0.966203K
Google: Gemma 4 26B A4B
google
not evaluated$0.070262K
Google: Gemma 4 26B A4B (free)
google
not evaluatedFree262K
Google: Gemma 4 31B
google
not evaluated$0.120262K
Google: Gemma 4 31B (free)
google
not evaluatedFree262K
Arcee AI: Trinity Large Thinking
arcee-ai
not evaluated$0.250262K
Reka Edge
rekaai
not evaluated$0.10016K
MiniMax: MiniMax M2.7
minimax
not evaluated$0.250205K
Mistral: Mistral Small 4
mistralai
not evaluated$0.150262K
NVIDIA: Nemotron 3 Super
nvidia
not evaluated$0.0801M
NVIDIA: Nemotron 3 Super (free)
nvidia
not evaluatedFree1M
Qwen: Qwen3.5-9B
qwen
not evaluated$0.100262K
Qwen: Qwen3.5-35B-A3B
qwen
not evaluated$0.140262K
Qwen: Qwen3.5-27B
qwen
not evaluated$0.260262K
Qwen: Qwen3.5-122B-A10B
qwen
not evaluated$0.260262K
Qwen: Qwen3.5 397B A17B
qwen
not evaluated$0.390262K
MiniMax: MiniMax M2.5
minimax
not evaluated$0.150205K
Z.ai: GLM 5
z-ai
not evaluated$0.950205K
Qwen: Qwen3 Coder Next
qwen
not evaluated$0.110262K
StepFun: Step 3.5 Flash
stepfun
not evaluated$0.100262K
MoonshotAI: Kimi K2.5
moonshotai
not evaluated$0.570262K
Z.ai: GLM 4.7 Flash
z-ai
not evaluated$0.061200K
MiniMax: MiniMax M2.1
minimax
not evaluated$0.300205K
Z.ai: GLM 4.7
z-ai
not evaluated$0.400203K
NVIDIA: Nemotron 3 Nano 30B A3B
nvidia
not evaluated$0.050262K
NVIDIA: Nemotron 3 Nano 30B A3B (free)
nvidia
not evaluatedFree256K
Mistral: Devstral 2 2512
mistralai
not evaluated$0.400262K
Z.ai: GLM 4.6V
z-ai
not evaluated$0.300131K
Mistral: Ministral 3 14B 2512
mistralai
not evaluated$0.200262K
Mistral: Ministral 3 8B 2512
mistralai
not evaluated$0.150262K
Mistral: Ministral 3 3B 2512
mistralai
not evaluated$0.100131K
DeepSeek: DeepSeek V3.2
deepseek
not evaluated$0.269164K
AllenAI: Olmo 3 32B Think
allenai
not evaluated$0.15066K
MoonshotAI: Kimi K2 Thinking
moonshotai
not evaluated$0.600262K
Mistral: Voxtral Small 24B 2507
mistralai
not evaluated$0.10032K
OpenAI: gpt-oss-safeguard-20b
openai
not evaluated$0.075131K
NVIDIA: Nemotron Nano 12B 2 VL (free)
nvidia
not evaluatedFree128K
MiniMax: MiniMax M2
minimax
not evaluated$0.300205K
Qwen: Qwen3 VL 32B Instruct
qwen
not evaluated$0.104262K
IBM: Granite 4.0 Micro
ibm-granite
not evaluated$0.017131K
Qwen: Qwen3 VL 8B Thinking
qwen
not evaluated$0.117256K
Qwen: Qwen3 VL 8B Instruct
qwen
not evaluated$0.117256K
Qwen: Qwen3 VL 30B A3B Thinking
qwen
not evaluated$0.130131K
Qwen: Qwen3 VL 30B A3B Instruct
qwen
not evaluated$0.130262K
Z.ai: GLM 4.6
z-ai
not evaluated$0.500203K
DeepSeek: DeepSeek V3.2 Exp
deepseek
not evaluated$0.270164K
TheDrummer: Cydonia 24B V4.1
thedrummer
not evaluated$0.300131K
Qwen: Qwen3 VL 235B A22B Thinking
qwen
not evaluated$0.260131K
Qwen: Qwen3 VL 235B A22B Instruct
qwen
not evaluated$0.210131K
DeepSeek: DeepSeek V3.1 Terminus
deepseek
not evaluated$0.270131K
Qwen: Qwen3 Next 80B A3B Thinking
qwen
not evaluated$0.098262K
Qwen: Qwen3 Next 80B A3B Instruct
qwen
not evaluated$0.098262K
NVIDIA: Nemotron Nano 9B V2 (free)
nvidia
not evaluatedFree128K
MoonshotAI: Kimi K2 0905
moonshotai
not evaluated$0.600262K
Qwen: Qwen3 30B A3B Thinking 2507
qwen
not evaluated$0.130131K
Nous: Hermes 4 70B
nousresearch
not evaluated$0.130131K
Nous: Hermes 4 405B
nousresearch
not evaluated$1.00131K
DeepSeek: DeepSeek V3.1
deepseek
not evaluated$0.250164K
Z.ai: GLM 4.5V
z-ai
not evaluated$0.60066K
AI21: Jamba Large 1.7
ai21
not evaluated$2.00256K
OpenAI: gpt-oss-120b
openai
not evaluated$0.037131K
OpenAI: gpt-oss-20b
openai
not evaluated$0.030131K
OpenAI: gpt-oss-20b (free)
openai
not evaluatedFree131K
Qwen: Qwen3 Coder 30B A3B Instruct
qwen
not evaluated$0.070160K
Qwen: Qwen3 30B A3B Instruct 2507
qwen
not evaluated$0.100262K
Z.ai: GLM 4.5
z-ai
not evaluated$0.600131K
Z.ai: GLM 4.5 Air
z-ai
not evaluated$0.130131K
Qwen: Qwen3 235B A22B Thinking 2507
qwen
not evaluated$0.300262K
Qwen: Qwen3 Coder 480B A35B
qwen
not evaluated$0.3001.0M
ByteDance: UI-TARS 7B
bytedance
not evaluated$0.100128K
Qwen: Qwen3 235B A22B Instruct 2507
qwen
not evaluated$0.090262K
MoonshotAI: Kimi K2 0711
moonshotai
not evaluated$0.570131K
Venice: Uncensored
cognitivecomputations
not evaluated$0.200128K
Tencent: Hunyuan A13B Instruct
tencent
not evaluated$0.140131K
Baidu: ERNIE 4.5 VL 424B A47B
baidu
not evaluated$0.420131K
Mistral: Mistral Small 3.2 24B
mistralai
not evaluated$0.100131K
DeepSeek: R1 0528
deepseek
not evaluated$0.500164K
Google: Gemma 3n 4B
google
not evaluated$0.06033K
Meta: Llama Guard 4 12B
meta-llama
not evaluated$0.180164K
Qwen: Qwen3 30B A3B
qwen
not evaluated$0.130131K
Qwen: Qwen3 8B
qwen
not evaluated$0.117131K
Qwen: Qwen3 14B
qwen
not evaluated$0.120132K
Qwen: Qwen3 32B
qwen
not evaluated$0.080131K
Qwen: Qwen3 235B A22B
qwen
not evaluated$0.455131K
Meta: Llama 4 Maverick
meta-llama
not evaluated$0.2001.0M
Meta: Llama 4 Scout
meta-llama
not evaluated$0.10010M
DeepSeek: DeepSeek V3 0324
deepseek
not evaluated$0.270164K
Mistral: Mistral Small 3.1 24B
mistralai
not evaluated$0.351128K
Google: Gemma 3 4B
google
not evaluated$0.050131K
Google: Gemma 3 12B
google
not evaluated$0.050131K
Cohere: Command A
cohere
not evaluated$2.50256K
Reka Flash 3
rekaai
not evaluated$0.10066K
Google: Gemma 3 27B
google
not evaluated$0.100131K
TheDrummer: Skyfall 36B V2
thedrummer
not evaluated$0.55033K
Qwen: Qwen2.5 VL 72B Instruct
qwen
not evaluated$0.800131K
Mistral: Mistral Small 3
mistralai
not evaluated$0.05033K
DeepSeek: R1 Distill Llama 70B
deepseek
not evaluated$0.800128K
DeepSeek: R1
deepseek
not evaluated$0.700164K
MiniMax: MiniMax-01
minimax
not evaluated$0.2001.0M
Microsoft: Phi 4
microsoft
not evaluated$0.07016K
DeepSeek: DeepSeek V3
deepseek
not evaluated$0.200131K
Sao10K: Llama 3.3 Euryale 70B
sao10k
not evaluated$0.650131K
Meta: Llama 3.3 70B Instruct
meta-llama
not evaluated$0.130131K
Qwen2.5 Coder 32B Instruct
qwen
not evaluated$0.660128K
TheDrummer: UnslopNemo 12B
thedrummer
not evaluated$0.40033K
Magnum v4 72B
anthracite-org
not evaluated$3.0033K
Qwen: Qwen2.5 7B Instruct
qwen
not evaluated$0.040131K
TheDrummer: Rocinante 12B
thedrummer
not evaluated$0.25066K
Meta: Llama 3.2 1B Instruct
meta-llama
not evaluated$0.027131K
Meta: Llama 3.2 3B Instruct
meta-llama
not evaluated$0.051131K
Qwen2.5 72B Instruct
qwen
not evaluated$0.360131K
Sao10K: Llama 3.1 Euryale 70B v2.2
sao10k
not evaluated$0.850131K
Nous: Hermes 3 70B Instruct
nousresearch
not evaluated$0.700131K
Nous: Hermes 3 405B Instruct
nousresearch
not evaluated$1.00131K
Sao10K: Llama 3 8B Lunaris
sao10k
not evaluated$0.0408K
Meta: Llama 3.1 70B Instruct
meta-llama
not evaluated$0.400131K
Meta: Llama 3.1 8B Instruct
meta-llama
not evaluated$0.050131K
Mistral: Mistral Nemo
mistralai
not evaluated$0.019131K
Google: Gemma 2 27B
google
not evaluated$0.6508K
Mistral: Mixtral 8x22B Instruct
mistralai
not evaluated$2.0066K
WizardLM-2 8x22B
microsoft
not evaluated$0.62066K
ReMM SLERP 13B
undi95
not evaluated$0.4506K
MythoMax 13B
gryphe
not evaluated$0.0604K

What "open weights" means here

A model is listed here when its weights are published and downloadable — you can run it on your own hardware, fine-tune it, and keep your prompts off a third party's servers. That is a narrower claim than "open source": most of these models publish weights without publishing training data or training code, and several carry licences with commercial restrictions. Check the licence on the model card before you build a business on one.

Why the ranking is the same score

Open-weight models are scored on exactly the same LiveBench tasks and the same published weights as proprietary ones — there is no separate, easier scale. That is deliberate: an open-model ranking is only useful if it tells you what you give up, if anything, by self-hosting. Where an open model reaches the top of the overall index, that is a like-for-like result, not a handicap category.

The price column still applies: most open-weight models here are also served by hosted providers, and the figure shown is that hosted price. If you self-host, your real cost is your own compute instead, which is usually cheaper at sustained volume and more expensive at low or spiky volume. Context length, by contrast, is a property of the model and holds wherever you run it.

Models with no score are absent from the current benchmark snapshot. They keep their factual data — price, context, modality — and make no capability claim, rather than being quietly ranked last. See methodology.