LLM Leaderboard — compare 338 AI models by capability, price & context

An independent index of 338 models. The LLM Index Score is computed by us from LiveBench's public raw task scores using published weights — every component is shown, so anyone can recompute it. Benchmark snapshot 2026_06_25.

tops the index
OpenAI: GPT-5.5
80.49 LLM Index Score
leads on reasoning
Anthropic: Claude Opus 4.8
89.7 reasoning
wins at coding
OpenAI: GPT-5.2-Codex
83.6 coding
cheapest in the top 10
OpenAI: GPT-5.2
$1.75 /M tok
longest context
Meta: Llama 4 Scout
10M tokens
best open-weights
Z.ai: GLM 5.2
73.84 LLM Index Score

Performance index

26 of 338 models carry a computed score. Models not in the benchmark snapshot show their facts and make no score claim.

25 models
#ModelLLM Index ScoreoursReasoningCodingAgentic CodingMathematicsData AnalysisLanguageIFPrice/M inContext
1OpenAI: GPT-5.5
openai
80.4989.782.252.195.981.687.470.7$5.001.1M
2Anthropic: Claude Fable 5
anthropic
79.8287.782.550.795.778.789.572.0$10.001M
3Anthropic: Claude Opus 4.8
anthropic
79.6789.779.356.195.378.381.472.4$5.001M
4OpenAI: GPT-5.4
openai
78.6488.177.553.894.279.382.670.2$2.501.1M
5Anthropic: Claude Opus 4.7
anthropic
77.4887.282.150.792.878.377.966.7$5.001M
6Google: Gemini 3.1 Pro Preview
google
76.9684.076.545.491.078.585.479.1$2.001.0M
7Anthropic: Claude Sonnet 5
anthropic
76.0888.780.751.192.971.775.063.9$2.001M
8OpenAI: GPT-5.2
openai
75.4583.276.150.393.278.279.861.8$1.75400K
9Anthropic: Claude Opus 4.6
anthropic
75.3588.778.249.089.369.983.363.3$5.001M
10OpenAI: GPT-5.2-Codex
openai
74.5577.783.649.488.878.273.766.5$1.75400K
11Google: Gemini 3.5 Flash
google
74.4682.078.249.088.264.984.675.6$1.501.0M
12Anthropic: Claude Sonnet 4.6
anthropic
73.9184.879.342.687.078.076.163.2$3.001M
13Z.ai: GLM 5.2
z-ai
73.8478.679.751.989.873.776.262.3$0.9521.0M
14Qwen: Qwen3.7 Max
qwen
73.2783.374.243.685.371.879.774.0$1.481M
15Anthropic: Claude Opus 4.5
anthropic
73.0380.179.739.790.474.481.362.5$5.00200K
16DeepSeek: DeepSeek V4 Pro
deepseek
72.2682.770.042.690.774.578.162.4$0.4351.0M
17MoonshotAI: Kimi K2.6
moonshotai
71.0679.478.646.984.365.175.164.4$0.684262K
18OpenAI: GPT-5.4 Nano
openai
70.6381.170.846.891.067.662.567.2$0.200400K
19Qwen: Qwen3.6 Plus
qwen
69.4775.878.241.483.769.975.058.3$0.3251M
20MoonshotAI: Kimi K2.7 Code
moonshotai
69.2682.874.045.779.662.777.956.3$0.820262K
21xAI: Grok Build 0.1
x-ai
68.1176.465.445.878.470.872.565.2$1.00256K
22MiniMax: MiniMax M3
minimax
67.6374.568.240.777.076.276.857.5$0.3001.0M
23OpenAI: GPT-5.4 Mini
openai
66.7271.371.641.778.570.871.059.8$0.750400K
24DeepSeek: DeepSeek V4 Flash
deepseek
65.6270.669.237.679.768.070.163.1$0.0941.0M
25Qwen: Qwen3.6 27B
qwen
64.9170.371.839.379.970.463.353.2$0.450262K

Best AI for each capability

How this index is built

Two layers, kept deliberately separate. The factual layer — price per token, context window, modality, launch date — comes from live provider data covering all 338 models, refreshed when the index rebuilds. The capability layer comes from LiveBench, an independent benchmark that publishes dated snapshots and resists contamination by rotating its questions. We compute the LLM Index Score from LiveBench's raw per-task results using weights we publish, and we show every component on every model page.

Nothing here is hand-tuned. No model gets a nudge, no score is adjusted after the fact, and where a third party publishes its own index we display it attributed to them rather than folding it into ours. When a model is missing benchmark data we say so and show no score, instead of guessing one — a partial score is not comparable to a complete one, and an estimate presented as a measurement is worse than no number at all.

Why component transparency matters

Every model leaderboard compresses many abilities into one number, and that compression is where the information goes missing. Two models with the same overall score can be completely different hires: one steady across the board, the other carried by a single very strong category while trailing everywhere else. If the leaderboard only shows you the total, you cannot tell them apart — and you will pick the wrong one for a narrow workload.

So the breakdown is not a footnote here, it is the point. Each model page prints all seven category scores, their weights, their contribution to the total, and the raw task scores underneath. You can add the contributions up yourself and land on the published score. A ranking you can check is worth more than a ranking you have to trust.

FAQ

Which AI model ranks first right now?

OpenAI: GPT-5.5 leads the index with an LLM Index Score of 80.49, computed from the LiveBench 2026_06_25 snapshot. The full ranking, including every category component, is on the leaderboard.

How many models does this leaderboard cover?

338 models carry full factual data — price, context window, modality and launch date. Of those, 26 also carry a computed LLM Index Score, because scoring requires complete benchmark coverage in the current snapshot.

Can I reproduce the LLM Index Score myself?

Yes, and that is the design goal. The inputs are public LiveBench task scores, the category weights are published on the methodology page, and every per-category and per-task value is printed on the model pages. Average the tasks in each category, apply the weights, and you get the same figure.

Why do some models show no score?

A model is scored only when every task in every category is present in the snapshot. Models released after the snapshot, or with partial results, keep their factual data and make no capability claim. We would rather show an honest gap than an estimate dressed up as a measurement.

Is the top-ranked model the right choice for me?

Often not. The overall score weights seven capabilities, and your workload probably does not. If you only care about code, or maths, or long-context retrieval, rank by that category instead — and check price, because capability and cost move independently.

How current is the data?

Factual data was retrieved 2026-07-21, and capability scores come from the LiveBench 2026_06_25 snapshot. Benchmarks are dated on purpose: they describe a model as measured on that date, not as it may behave after a provider-side update.