AI model rankings

All 338 indexed models, sortable by score, price and context. 26 carry a computed LLM Index Score from the 2026_06_25 benchmark snapshot; the rest show facts only and make no score claim.

338 models
#ModelLLM Index ScoreoursReasoningCodingAgentic CodingMathematicsData AnalysisLanguageIFPrice/M inContext
1OpenAI: GPT-5.5
openai
80.4989.782.252.195.981.687.470.7$5.001.1M
2Anthropic: Claude Fable 5
anthropic
79.8287.782.550.795.778.789.572.0$10.001M
3Anthropic: Claude Opus 4.8
anthropic
79.6789.779.356.195.378.381.472.4$5.001M
4OpenAI: GPT-5.4
openai
78.6488.177.553.894.279.382.670.2$2.501.1M
5Anthropic: Claude Opus 4.7
anthropic
77.4887.282.150.792.878.377.966.7$5.001M
6Google: Gemini 3.1 Pro Preview
google
76.9684.076.545.491.078.585.479.1$2.001.0M
7Anthropic: Claude Sonnet 5
anthropic
76.0888.780.751.192.971.775.063.9$2.001M
8OpenAI: GPT-5.2
openai
75.4583.276.150.393.278.279.861.8$1.75400K
9Anthropic: Claude Opus 4.6
anthropic
75.3588.778.249.089.369.983.363.3$5.001M
10OpenAI: GPT-5.2-Codex
openai
74.5577.783.649.488.878.273.766.5$1.75400K
11Google: Gemini 3.5 Flash
google
74.4682.078.249.088.264.984.675.6$1.501.0M
12Anthropic: Claude Sonnet 4.6
anthropic
73.9184.879.342.687.078.076.163.2$3.001M
13Z.ai: GLM 5.2
z-ai
73.8478.679.751.989.873.776.262.3$0.9521.0M
14Qwen: Qwen3.7 Max
qwen
73.2783.374.243.685.371.879.774.0$1.481M
15Anthropic: Claude Opus 4.5
anthropic
73.0380.179.739.790.474.481.362.5$5.00200K
16DeepSeek: DeepSeek V4 Pro
deepseek
72.2682.770.042.690.774.578.162.4$0.4351.0M
17MoonshotAI: Kimi K2.6
moonshotai
71.0679.478.646.984.365.175.164.4$0.684262K
18OpenAI: GPT-5.4 Nano
openai
70.6381.170.846.891.067.662.567.2$0.200400K
19Qwen: Qwen3.6 Plus
qwen
69.4775.878.241.483.769.975.058.3$0.3251M
20MoonshotAI: Kimi K2.7 Code
moonshotai
69.2682.874.045.779.662.777.956.3$0.820262K
21xAI: Grok Build 0.1
x-ai
68.1176.465.445.878.470.872.565.2$1.00256K
22MiniMax: MiniMax M3
minimax
67.6374.568.240.777.076.276.857.5$0.3001.0M
23OpenAI: GPT-5.4 Mini
openai
66.7271.371.641.778.570.871.059.8$0.750400K
24DeepSeek: DeepSeek V4 Flash
deepseek
65.6270.669.237.679.768.070.163.1$0.0941.0M
25Qwen: Qwen3.6 27B
qwen
64.9170.371.839.379.970.463.353.2$0.450262K
26xAI: Grok 4.3
x-ai
62.0870.869.918.584.355.873.662.8$1.251M
Meituan: LongCat 2.0
meituan
not evaluated$0.3001.0M
Thinking Machines: Inkling
thinkingmachines
not evaluated$1.001.0M
Auto Router (Beta)
openrouter
not evaluated$-1000000.0002M
MoonshotAI: Kimi K3
moonshotai
not evaluated$3.001.0M
Meta: Muse Spark 1.1
meta
not evaluated$1.251.0M
Kwaipilot: KAT-Coder-Air V2.5
kwaipilot
not evaluated$0.150256K
Kwaipilot: KAT-Coder-Pro V2.5
kwaipilot
not evaluated$0.740256K
OpenAI: GPT-5.6 Luna Pro
openai
not evaluated$1.001.1M
OpenAI: GPT-5.6 Luna
openai
not evaluated$1.001.1M
OpenAI: GPT-5.6 Terra Pro
openai
not evaluated$2.501.1M
OpenAI: GPT-5.6 Terra
openai
not evaluated$2.501.1M
OpenAI: GPT-5.6 Sol Pro
openai
not evaluated$5.001.1M
OpenAI: GPT-5.6 Sol
openai
not evaluated$5.001.1M
xAI: Grok 4.5
x-ai
not evaluated$2.00500K
xAI: Grok Latest
~x-ai
not evaluated$2.00500K
AionLabs: Aion-3.0-Mini
aion-labs
not evaluated$0.700131K
AionLabs: Aion-3.0
aion-labs
not evaluated$3.00131K
Tencent: Hy3
tencent
not evaluated$0.140262K
Poolside: Laguna XS 2.1
poolside
not evaluated$0.060262K
Poolside: Laguna XS 2.1 (free)
poolside
not evaluatedFree262K
Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
google
not evaluated$0.25066K
Nex AGI: Nex-N2-Mini
nex-agi
not evaluated$0.025262K
Sakana: Fugu Ultra
sakana
not evaluated$5.001M
Google: Nano Banana 2 (Gemini 3.1 Flash Image)
google
not evaluated$0.500131K
Google: Nano Banana Pro (Gemini 3 Pro Image)
google
not evaluated$2.0066K
Cohere: North Mini Code (free)
cohere
not evaluatedFree256K
OpenRouter: Fusion
openrouter
not evaluated$-1000000.0001M
Anthropic: Claude Fable Latest
~anthropic
not evaluated$10.001M
Nex AGI: Nex-N2-Pro
nex-agi
not evaluated$0.250262K
NVIDIA: Nemotron 3.5 Content Safety (free)
nvidia
not evaluatedFree128K
NVIDIA: Nemotron 3 Ultra
nvidia
not evaluated$0.6001M
NVIDIA: Nemotron 3 Ultra (free)
nvidia
not evaluatedFree1M
Qwen: Qwen3.7 Plus
qwen
not evaluated$0.3201M
StepFun: Step 3.7 Flash
stepfun
not evaluated$0.200256K
Anthropic: Claude Opus 4.8 (Fast)
anthropic
not evaluated$10.001M
Anthropic: Claude Opus 4.7 (Fast)
anthropic
not evaluated$30.001M
Perceptron: Perceptron Mk1
perceptron
not evaluated$0.15033K
inclusionAI: Ring-2.6-1T
inclusionai
not evaluated$0.075262K
Google: Gemini 3.1 Flash Lite
google
not evaluated$0.2501.0M
OpenAI: GPT Chat Latest
openai
not evaluated$5.00400K
IBM: Granite 4.1 8B
ibm-granite
not evaluated$0.050131K
Mistral: Mistral Medium 3.5
mistralai
not evaluated$1.50262K
NVIDIA: Nemotron 3 Nano Omni (free)
nvidia
not evaluatedFree256K
Poolside: Laguna M.1
poolside
not evaluated$0.200262K
Poolside: Laguna M.1 (free)
poolside
not evaluatedFree262K
Anthropic Claude Haiku Latest
~anthropic
not evaluated$1.00200K
OpenAI GPT Mini Latest
~openai
not evaluated$0.750400K
Google Gemini Pro Latest
~google
not evaluated$2.001.0M
MoonshotAI Kimi Latest
~moonshotai
not evaluated$3.001.0M
Google Gemini Flash Latest
~google
not evaluated$1.501.0M
Anthropic Claude Sonnet Latest
~anthropic
not evaluated$2.001M
OpenAI GPT Latest
~openai
not evaluated$5.001.1M
Qwen: Qwen3.5 Plus 2026-04-20
qwen
not evaluated$0.3001M
Qwen: Qwen3.6 Flash
qwen
not evaluated$0.1881M
Qwen: Qwen3.6 35B A3B
qwen
not evaluated$0.140262K
Qwen: Qwen3.6 Max Preview
qwen
not evaluated$1.04262K
OpenAI: GPT-5.5 Pro
openai
not evaluated$30.001.1M
inclusionAI: Ling-2.6-1T
inclusionai
not evaluated$0.075262K
Tencent: Hy3 preview
tencent
not evaluated$0.063262K
Xiaomi: MiMo-V2.5-Pro
xiaomi
not evaluated$0.4351.0M
Xiaomi: MiMo-V2.5
xiaomi
not evaluated$0.1401.0M
OpenAI: GPT-5.4 Image 2
openai
not evaluated$8.00272K
inclusionAI: Ling-2.6-flash
inclusionai
not evaluated$0.010262K
Anthropic: Claude Opus Latest
~anthropic
not evaluated$5.001M
Pareto Code Router
openrouter
not evaluated$-1000000.0002M
Z.ai: GLM 5.1
z-ai
not evaluated$0.966203K
Google: Gemma 4 26B A4B
google
not evaluated$0.070262K
Google: Gemma 4 26B A4B (free)
google
not evaluatedFree262K
Google: Gemma 4 31B
google
not evaluated$0.120262K
Google: Gemma 4 31B (free)
google
not evaluatedFree262K
Z.ai: GLM 5V Turbo
z-ai
not evaluated$1.20203K
Arcee AI: Trinity Large Thinking
arcee-ai
not evaluated$0.250262K
xAI: Grok 4.20 Multi-Agent
x-ai
not evaluated$1.252M
xAI: Grok 4.20
x-ai
not evaluated$1.252M
Google: Lyria 3 Pro Preview
google
not evaluatedFree1.0M
Google: Lyria 3 Clip Preview
google
not evaluatedFree1.0M
Kwaipilot: KAT-Coder-Pro V2
kwaipilot
not evaluated$0.300256K
Reka Edge
rekaai
not evaluated$0.10016K
MiniMax: MiniMax M2.7
minimax
not evaluated$0.250205K
Mistral: Mistral Small 4
mistralai
not evaluated$0.150262K
Z.ai: GLM 5 Turbo
z-ai
not evaluated$1.20203K
NVIDIA: Nemotron 3 Super
nvidia
not evaluated$0.0801M
NVIDIA: Nemotron 3 Super (free)
nvidia
not evaluatedFree1M
ByteDance Seed: Seed-2.0-Lite
bytedance-seed
not evaluated$0.250262K
Qwen: Qwen3.5-9B
qwen
not evaluated$0.100262K
OpenAI: GPT-5.4 Pro
openai
not evaluated$30.001.1M
Inception: Mercury 2
inception
not evaluated$0.250128K
OpenAI: GPT-5.3 Chat
openai
not evaluated$1.75128K
Google: Gemini 3.1 Flash Lite Preview
google
not evaluated$0.2501.0M
ByteDance Seed: Seed-2.0-Mini
bytedance-seed
not evaluated$0.100262K
Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)
google
not evaluated$0.500131K
Qwen: Qwen3.5-35B-A3B
qwen
not evaluated$0.140262K
Qwen: Qwen3.5-27B
qwen
not evaluated$0.260262K
Qwen: Qwen3.5-122B-A10B
qwen
not evaluated$0.260262K
Qwen: Qwen3.5-Flash
qwen
not evaluated$0.0651M
Google: Gemini 3.1 Pro Preview Custom Tools
google
not evaluated$2.001.0M
OpenAI: GPT-5.3-Codex
openai
not evaluated$1.75400K
AionLabs: Aion-2.0
aion-labs
not evaluated$0.800131K
Qwen: Qwen3.5 Plus 2026-02-15
qwen
not evaluated$0.2601M
Qwen: Qwen3.5 397B A17B
qwen
not evaluated$0.390262K
MiniMax: MiniMax M2.5
minimax
not evaluated$0.150205K
Z.ai: GLM 5
z-ai
not evaluated$0.950205K
Qwen: Qwen3 Max Thinking
qwen
not evaluated$0.780262K
Qwen: Qwen3 Coder Next
qwen
not evaluated$0.110262K
Free Models Router
openrouter
not evaluatedFree200K
StepFun: Step 3.5 Flash
stepfun
not evaluated$0.100262K
MoonshotAI: Kimi K2.5
moonshotai
not evaluated$0.570262K
Upstage: Solar Pro 3
upstage
not evaluated$0.150128K
MiniMax: MiniMax M2-her
minimax
not evaluated$0.30066K
Writer: Palmyra X5
writer
not evaluated$0.6001.0M
OpenAI: GPT Audio
openai
not evaluated$2.50128K
OpenAI: GPT Audio Mini
openai
not evaluated$0.600128K
Z.ai: GLM 4.7 Flash
z-ai
not evaluated$0.061200K
ByteDance Seed: Seed 1.6 Flash
bytedance-seed
not evaluated$0.075262K
ByteDance Seed: Seed 1.6
bytedance-seed
not evaluated$0.250262K
MiniMax: MiniMax M2.1
minimax
not evaluated$0.300205K
Z.ai: GLM 4.7
z-ai
not evaluated$0.400203K
Google: Gemini 3 Flash Preview
google
not evaluated$0.5001.0M
NVIDIA: Nemotron 3 Nano 30B A3B
nvidia
not evaluated$0.050262K
NVIDIA: Nemotron 3 Nano 30B A3B (free)
nvidia
not evaluatedFree256K
OpenAI: GPT-5.2 Chat
openai
not evaluated$1.75128K
OpenAI: GPT-5.2 Pro
openai
not evaluated$21.00400K
Mistral: Devstral 2 2512
mistralai
not evaluated$0.400262K
Relace: Relace Search
relace
not evaluated$1.00256K
Z.ai: GLM 4.6V
z-ai
not evaluated$0.300131K
Body Builder (beta)
openrouter
not evaluated$-1000000.000128K
OpenAI: GPT-5.1-Codex-Max
openai
not evaluated$1.25400K
Amazon: Nova 2 Lite
amazon
not evaluated$0.3001M
Mistral: Ministral 3 14B 2512
mistralai
not evaluated$0.200262K
Mistral: Ministral 3 8B 2512
mistralai
not evaluated$0.150262K
Mistral: Ministral 3 3B 2512
mistralai
not evaluated$0.100131K
Mistral: Mistral Large 3 2512
mistralai
not evaluated$0.500262K
DeepSeek: DeepSeek V3.2
deepseek
not evaluated$0.269164K
AllenAI: Olmo 3 32B Think
allenai
not evaluated$0.15066K
Google: Nano Banana Pro (Gemini 3 Pro Image Preview)
google
not evaluated$2.0066K
Deep Cogito: Cogito v2.1 671B
deepcogito
not evaluated$1.25128K
OpenAI: GPT-5.1
openai
not evaluated$1.25400K
OpenAI: GPT-5.1 Chat
openai
not evaluated$1.25128K
OpenAI: GPT-5.1-Codex
openai
not evaluated$1.25400K
OpenAI: GPT-5.1-Codex-Mini
openai
not evaluated$0.250400K
MoonshotAI: Kimi K2 Thinking
moonshotai
not evaluated$0.600262K
Amazon: Nova Premier 1.0
amazon
not evaluated$2.501M
Perplexity: Sonar Pro Search
perplexity
not evaluated$3.00200K
Mistral: Voxtral Small 24B 2507
mistralai
not evaluated$0.10032K
OpenAI: gpt-oss-safeguard-20b
openai
not evaluated$0.075131K
NVIDIA: Nemotron Nano 12B 2 VL (free)
nvidia
not evaluatedFree128K
MiniMax: MiniMax M2
minimax
not evaluated$0.300205K
Qwen: Qwen3 VL 32B Instruct
qwen
not evaluated$0.104262K
IBM: Granite 4.0 Micro
ibm-granite
not evaluated$0.017131K
OpenAI: GPT-5 Image Mini
openai
not evaluated$2.50400K
Anthropic: Claude Haiku 4.5
anthropic
not evaluated$1.00200K
Qwen: Qwen3 VL 8B Thinking
qwen
not evaluated$0.117256K
Qwen: Qwen3 VL 8B Instruct
qwen
not evaluated$0.117256K
OpenAI: GPT-5 Image
openai
not evaluated$10.00400K
OpenAI: o3 Deep Research
openai
not evaluated$10.00200K
OpenAI: o4 Mini Deep Research
openai
not evaluated$2.00200K
Google: Nano Banana (Gemini 2.5 Flash Image)
google
not evaluated$0.30033K
Qwen: Qwen3 VL 30B A3B Thinking
qwen
not evaluated$0.130131K
Qwen: Qwen3 VL 30B A3B Instruct
qwen
not evaluated$0.130262K
OpenAI: GPT-5 Pro
openai
not evaluated$15.00400K
Z.ai: GLM 4.6
z-ai
not evaluated$0.500203K
Anthropic: Claude Sonnet 4.5
anthropic
not evaluated$3.001M
DeepSeek: DeepSeek V3.2 Exp
deepseek
not evaluated$0.270164K
TheDrummer: Cydonia 24B V4.1
thedrummer
not evaluated$0.300131K
Relace: Relace Apply 3
relace
not evaluated$0.850256K
Qwen: Qwen3 VL 235B A22B Thinking
qwen
not evaluated$0.260131K
Qwen: Qwen3 VL 235B A22B Instruct
qwen
not evaluated$0.210131K
Qwen: Qwen3 Max
qwen
not evaluated$0.780262K
Qwen: Qwen3 Coder Plus
qwen
not evaluated$0.6501M
OpenAI: GPT-5 Codex
openai
not evaluated$1.25400K
DeepSeek: DeepSeek V3.1 Terminus
deepseek
not evaluated$0.270131K
Qwen: Qwen3 Coder Flash
qwen
not evaluated$0.1951M
Qwen: Qwen3 Next 80B A3B Thinking
qwen
not evaluated$0.098262K
Qwen: Qwen3 Next 80B A3B Instruct
qwen
not evaluated$0.098262K
Qwen: Qwen Plus 0728
qwen
not evaluated$0.2601M
Qwen: Qwen Plus 0728 (thinking)
qwen
not evaluated$0.2601M
NVIDIA: Nemotron Nano 9B V2 (free)
nvidia
not evaluatedFree128K
MoonshotAI: Kimi K2 0905
moonshotai
not evaluated$0.600262K
Qwen: Qwen3 30B A3B Thinking 2507
qwen
not evaluated$0.130131K
Nous: Hermes 4 70B
nousresearch
not evaluated$0.130131K
Nous: Hermes 4 405B
nousresearch
not evaluated$1.00131K
DeepSeek: DeepSeek V3.1
deepseek
not evaluated$0.250164K
Mistral: Mistral Medium 3.1
mistralai
not evaluated$0.400131K
Z.ai: GLM 4.5V
z-ai
not evaluated$0.60066K
AI21: Jamba Large 1.7
ai21
not evaluated$2.00256K
OpenAI: GPT-5 Chat
openai
not evaluated$1.25128K
OpenAI: GPT-5
openai
not evaluated$1.25400K
OpenAI: GPT-5 Mini
openai
not evaluated$0.250400K
OpenAI: GPT-5 Nano
openai
not evaluated$0.050400K
OpenAI: gpt-oss-120b
openai
not evaluated$0.037131K
OpenAI: gpt-oss-20b
openai
not evaluated$0.030131K
OpenAI: gpt-oss-20b (free)
openai
not evaluatedFree131K
Anthropic: Claude Opus 4.1
anthropic
not evaluated$15.00200K
Mistral: Codestral 2508
mistralai
not evaluated$0.300256K
Qwen: Qwen3 Coder 30B A3B Instruct
qwen
not evaluated$0.070160K
Qwen: Qwen3 30B A3B Instruct 2507
qwen
not evaluated$0.100262K
Z.ai: GLM 4.5
z-ai
not evaluated$0.600131K
Z.ai: GLM 4.5 Air
z-ai
not evaluated$0.130131K
Qwen: Qwen3 235B A22B Thinking 2507
qwen
not evaluated$0.300262K
Qwen: Qwen3 Coder 480B A35B
qwen
not evaluated$0.3001.0M
ByteDance: UI-TARS 7B
bytedance
not evaluated$0.100128K
Google: Gemini 2.5 Flash Lite
google
not evaluated$0.1001.0M
Qwen: Qwen3 235B A22B Instruct 2507
qwen
not evaluated$0.090262K
MoonshotAI: Kimi K2 0711
moonshotai
not evaluated$0.570131K
Venice: Uncensored
cognitivecomputations
not evaluated$0.200128K
Tencent: Hunyuan A13B Instruct
tencent
not evaluated$0.140131K
Morph: Morph V3 Large
morph
not evaluated$0.900262K
Morph: Morph V3 Fast
morph
not evaluated$0.80082K
Baidu: ERNIE 4.5 VL 424B A47B
baidu
not evaluated$0.420131K
Mistral: Mistral Small 3.2 24B
mistralai
not evaluated$0.100131K
MiniMax: MiniMax M1
minimax
not evaluated$0.5501M
Google: Gemini 2.5 Flash
google
not evaluated$0.3001.0M
Google: Gemini 2.5 Pro
google
not evaluated$1.251.0M
OpenAI: o3 Pro
openai
not evaluated$20.00200K
Google: Gemini 2.5 Pro Preview 06-05
google
not evaluated$1.251.0M
DeepSeek: R1 0528
deepseek
not evaluated$0.500164K
Anthropic: Claude Opus 4
anthropic
not evaluated$15.00200K
Anthropic: Claude Sonnet 4
anthropic
not evaluated$3.001M
Google: Gemma 3n 4B
google
not evaluated$0.06033K
Mistral: Mistral Medium 3
mistralai
not evaluated$0.400131K
Google: Gemini 2.5 Pro Preview 05-06
google
not evaluated$1.251.0M
Arcee AI: Virtuoso Large
arcee-ai
not evaluated$0.750131K
Meta: Llama Guard 4 12B
meta-llama
not evaluated$0.180164K
Qwen: Qwen3 30B A3B
qwen
not evaluated$0.130131K
Qwen: Qwen3 8B
qwen
not evaluated$0.117131K
Qwen: Qwen3 14B
qwen
not evaluated$0.120132K
Qwen: Qwen3 32B
qwen
not evaluated$0.080131K
Qwen: Qwen3 235B A22B
qwen
not evaluated$0.455131K
OpenAI: o4 Mini High
openai
not evaluated$1.10200K
OpenAI: o3
openai
not evaluated$2.00200K
OpenAI: o4 Mini
openai
not evaluated$1.10200K
OpenAI: GPT-4.1
openai
not evaluated$2.001.0M
OpenAI: GPT-4.1 Mini
openai
not evaluated$0.4001.0M
OpenAI: GPT-4.1 Nano
openai
not evaluated$0.1001.0M
Meta: Llama 4 Maverick
meta-llama
not evaluated$0.2001.0M
Meta: Llama 4 Scout
meta-llama
not evaluated$0.10010M
DeepSeek: DeepSeek V3 0324
deepseek
not evaluated$0.270164K
OpenAI: o1-pro
openai
not evaluated$150.00200K
Mistral: Mistral Small 3.1 24B
mistralai
not evaluated$0.351128K
Google: Gemma 3 4B
google
not evaluated$0.050131K
Google: Gemma 3 12B
google
not evaluated$0.050131K
Cohere: Command A
cohere
not evaluated$2.50256K
OpenAI: GPT-4o-mini Search Preview
openai
not evaluated$0.150128K
OpenAI: GPT-4o Search Preview
openai
not evaluated$2.50128K
Reka Flash 3
rekaai
not evaluated$0.10066K
Google: Gemma 3 27B
google
not evaluated$0.100131K
TheDrummer: Skyfall 36B V2
thedrummer
not evaluated$0.55033K
Perplexity: Sonar Reasoning Pro
perplexity
not evaluated$2.00128K
Perplexity: Sonar Pro
perplexity
not evaluated$3.00200K
Perplexity: Sonar Deep Research
perplexity
not evaluated$2.00128K
Mistral: Saba
mistralai
not evaluated$0.20033K
OpenAI: o3 Mini High
openai
not evaluated$1.10200K
AionLabs: Aion-RP 1.0 (8B)
aion-labs
not evaluated$0.80033K
Qwen: Qwen2.5 VL 72B Instruct
qwen
not evaluated$0.800131K
Qwen: Qwen-Plus
qwen
not evaluated$0.2601M
OpenAI: o3 Mini
openai
not evaluated$1.10200K
Mistral: Mistral Small 3
mistralai
not evaluated$0.05033K
Perplexity: Sonar
perplexity
not evaluated$1.00127K
DeepSeek: R1 Distill Llama 70B
deepseek
not evaluated$0.800128K
DeepSeek: R1
deepseek
not evaluated$0.700164K
MiniMax: MiniMax-01
minimax
not evaluated$0.2001.0M
Microsoft: Phi 4
microsoft
not evaluated$0.07016K
DeepSeek: DeepSeek V3
deepseek
not evaluated$0.200131K
Sao10K: Llama 3.3 Euryale 70B
sao10k
not evaluated$0.650131K
OpenAI: o1
openai
not evaluated$15.00200K
Cohere: Command R7B (12-2024)
cohere
not evaluated$0.037128K
Meta: Llama 3.3 70B Instruct
meta-llama
not evaluated$0.130131K
Amazon: Nova Lite 1.0
amazon
not evaluated$0.060300K
Amazon: Nova Micro 1.0
amazon
not evaluated$0.035128K
Amazon: Nova Pro 1.0
amazon
not evaluated$0.800300K
OpenAI: GPT-4o (2024-11-20)
openai
not evaluated$2.50128K
Mistral Large 2407
mistralai
not evaluated$2.00131K
Qwen2.5 Coder 32B Instruct
qwen
not evaluated$0.660128K
TheDrummer: UnslopNemo 12B
thedrummer
not evaluated$0.40033K
Magnum v4 72B
anthracite-org
not evaluated$3.0033K
Qwen: Qwen2.5 7B Instruct
qwen
not evaluated$0.040131K
Inflection: Inflection 3 Pi
inflection
not evaluated$2.508K
Inflection: Inflection 3 Productivity
inflection
not evaluated$2.508K
TheDrummer: Rocinante 12B
thedrummer
not evaluated$0.25066K
Meta: Llama 3.2 1B Instruct
meta-llama
not evaluated$0.027131K
Meta: Llama 3.2 3B Instruct
meta-llama
not evaluated$0.051131K
Qwen2.5 72B Instruct
qwen
not evaluated$0.360131K
Cohere: Command R (08-2024)
cohere
not evaluated$0.150128K
Cohere: Command R+ (08-2024)
cohere
not evaluated$2.50128K
Sao10K: Llama 3.1 Euryale 70B v2.2
sao10k
not evaluated$0.850131K
Nous: Hermes 3 70B Instruct
nousresearch
not evaluated$0.700131K
Nous: Hermes 3 405B Instruct
nousresearch
not evaluated$1.00131K
Sao10K: Llama 3 8B Lunaris
sao10k
not evaluated$0.0408K
OpenAI: GPT-4o (2024-08-06)
openai
not evaluated$2.50128K
Meta: Llama 3.1 70B Instruct
meta-llama
not evaluated$0.400131K
Meta: Llama 3.1 8B Instruct
meta-llama
not evaluated$0.050131K
Mistral: Mistral Nemo
mistralai
not evaluated$0.019131K
OpenAI: GPT-4o-mini
openai
not evaluated$0.150128K
OpenAI: GPT-4o-mini (2024-07-18)
openai
not evaluated$0.150128K
Google: Gemma 2 27B
google
not evaluated$0.6508K
OpenAI: GPT-4o
openai
not evaluated$2.50128K
OpenAI: GPT-4o (2024-05-13)
openai
not evaluated$5.00128K
Mistral: Mixtral 8x22B Instruct
mistralai
not evaluated$2.0066K
WizardLM-2 8x22B
microsoft
not evaluated$0.62066K
OpenAI: GPT-4 Turbo
openai
not evaluated$10.00128K
Anthropic: Claude 3 Haiku
anthropic
not evaluated$0.250200K
Mistral Large
mistralai
not evaluated$2.00128K
OpenAI: GPT-3.5 Turbo (older v0613)
openai
not evaluated$1.004K
OpenAI: GPT-4 Turbo Preview
openai
not evaluated$10.00128K
Auto Router
openrouter
not evaluated$-1000000.0002M
OpenAI: GPT-3.5 Turbo Instruct
openai
not evaluated$1.504K
OpenAI: GPT-3.5 Turbo 16k
openai
not evaluated$3.0016K
Mancer: Weaver (alpha)
mancer
not evaluated$0.5008K
ReMM SLERP 13B
undi95
not evaluated$0.4506K
MythoMax 13B
gryphe
not evaluated$0.0604K
OpenAI: GPT-3.5 Turbo
openai
not evaluated$0.50016K
OpenAI: GPT-4
openai
not evaluated$30.008K

How to sort this table

The default order is by LLM Index Score, which weights seven capabilities into one figure. That is the right starting point and the wrong finishing point: your workload almost certainly does not weight those capabilities the way the overall score does. Sort by an individual category column when the job is narrow — code generation, long-context retrieval, quantitative work — and the ordering will change, sometimes substantially.

Sorting by price is the other view worth taking. Capability and cost are set independently, and the difference between the first and fifth model is frequently a rounding error next to the difference in what they charge. For sustained or high-volume work, the model a few places down the table is often the correct answer.

Why some rows have no score

Of 338 indexed models, 26 carry a computed score. The rest are absent from the current benchmark snapshot — usually because they shipped after it was taken. Those rows keep their factual data and are labelled explicitly rather than being silently sorted to the bottom, which would read as "measured and found wanting" when the truth is "not measured yet".

We do not fill those gaps with estimates. Interpolating a score from a model's family, its parameter count or its vendor's track record would produce a number that looks exactly like a measurement and is not one. When the next snapshot lands, those rows gain real scores.

What the columns mean

The seven category columns are the components of the overall score, each a plain mean of its raw LiveBench tasks on a 0–100 scale. Price is per million input tokens from live provider data; output tokens are priced separately and are usually the larger share of a real bill, so check the model page before budgeting. Context is the maximum window accepted in one request, covering prompt and response together. Neither price nor context feeds the score — they sit beside it because a ranking without cost is only half a decision. Full derivation is on the methodology page.