LLM Leaderboard — compare 338 AI models by capability, price & context
An independent index of 338 models. The LLM Index Score is computed by us from LiveBench's public raw task scores using published weights — every component is shown, so anyone can recompute it. Benchmark snapshot 2026_06_25.
Performance index
26 of 338 models carry a computed score. Models not in the benchmark snapshot show their facts and make no score claim.
| # | Model | LLM Index Scoreours | Reasoning | Coding | Agentic Coding | Mathematics | Data Analysis | Language | IF | Price/M in | Context |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | OpenAI: GPT-5.5 openai | 80.49 | 89.7 | 82.2 | 52.1 | 95.9 | 81.6 | 87.4 | 70.7 | $5.00 | 1.1M |
| 2 | Anthropic: Claude Fable 5 anthropic | 79.82 | 87.7 | 82.5 | 50.7 | 95.7 | 78.7 | 89.5 | 72.0 | $10.00 | 1M |
| 3 | Anthropic: Claude Opus 4.8 anthropic | 79.67 | 89.7 | 79.3 | 56.1 | 95.3 | 78.3 | 81.4 | 72.4 | $5.00 | 1M |
| 4 | OpenAI: GPT-5.4 openai | 78.64 | 88.1 | 77.5 | 53.8 | 94.2 | 79.3 | 82.6 | 70.2 | $2.50 | 1.1M |
| 5 | Anthropic: Claude Opus 4.7 anthropic | 77.48 | 87.2 | 82.1 | 50.7 | 92.8 | 78.3 | 77.9 | 66.7 | $5.00 | 1M |
| 6 | Google: Gemini 3.1 Pro Preview google | 76.96 | 84.0 | 76.5 | 45.4 | 91.0 | 78.5 | 85.4 | 79.1 | $2.00 | 1.0M |
| 7 | Anthropic: Claude Sonnet 5 anthropic | 76.08 | 88.7 | 80.7 | 51.1 | 92.9 | 71.7 | 75.0 | 63.9 | $2.00 | 1M |
| 8 | OpenAI: GPT-5.2 openai | 75.45 | 83.2 | 76.1 | 50.3 | 93.2 | 78.2 | 79.8 | 61.8 | $1.75 | 400K |
| 9 | Anthropic: Claude Opus 4.6 anthropic | 75.35 | 88.7 | 78.2 | 49.0 | 89.3 | 69.9 | 83.3 | 63.3 | $5.00 | 1M |
| 10 | OpenAI: GPT-5.2-Codex openai | 74.55 | 77.7 | 83.6 | 49.4 | 88.8 | 78.2 | 73.7 | 66.5 | $1.75 | 400K |
| 11 | Google: Gemini 3.5 Flash google | 74.46 | 82.0 | 78.2 | 49.0 | 88.2 | 64.9 | 84.6 | 75.6 | $1.50 | 1.0M |
| 12 | Anthropic: Claude Sonnet 4.6 anthropic | 73.91 | 84.8 | 79.3 | 42.6 | 87.0 | 78.0 | 76.1 | 63.2 | $3.00 | 1M |
| 13 | Z.ai: GLM 5.2 z-ai | 73.84 | 78.6 | 79.7 | 51.9 | 89.8 | 73.7 | 76.2 | 62.3 | $0.952 | 1.0M |
| 14 | Qwen: Qwen3.7 Max qwen | 73.27 | 83.3 | 74.2 | 43.6 | 85.3 | 71.8 | 79.7 | 74.0 | $1.48 | 1M |
| 15 | Anthropic: Claude Opus 4.5 anthropic | 73.03 | 80.1 | 79.7 | 39.7 | 90.4 | 74.4 | 81.3 | 62.5 | $5.00 | 200K |
| 16 | DeepSeek: DeepSeek V4 Pro deepseek | 72.26 | 82.7 | 70.0 | 42.6 | 90.7 | 74.5 | 78.1 | 62.4 | $0.435 | 1.0M |
| 17 | MoonshotAI: Kimi K2.6 moonshotai | 71.06 | 79.4 | 78.6 | 46.9 | 84.3 | 65.1 | 75.1 | 64.4 | $0.684 | 262K |
| 18 | OpenAI: GPT-5.4 Nano openai | 70.63 | 81.1 | 70.8 | 46.8 | 91.0 | 67.6 | 62.5 | 67.2 | $0.200 | 400K |
| 19 | Qwen: Qwen3.6 Plus qwen | 69.47 | 75.8 | 78.2 | 41.4 | 83.7 | 69.9 | 75.0 | 58.3 | $0.325 | 1M |
| 20 | MoonshotAI: Kimi K2.7 Code moonshotai | 69.26 | 82.8 | 74.0 | 45.7 | 79.6 | 62.7 | 77.9 | 56.3 | $0.820 | 262K |
| 21 | xAI: Grok Build 0.1 x-ai | 68.11 | 76.4 | 65.4 | 45.8 | 78.4 | 70.8 | 72.5 | 65.2 | $1.00 | 256K |
| 22 | MiniMax: MiniMax M3 minimax | 67.63 | 74.5 | 68.2 | 40.7 | 77.0 | 76.2 | 76.8 | 57.5 | $0.300 | 1.0M |
| 23 | OpenAI: GPT-5.4 Mini openai | 66.72 | 71.3 | 71.6 | 41.7 | 78.5 | 70.8 | 71.0 | 59.8 | $0.750 | 400K |
| 24 | DeepSeek: DeepSeek V4 Flash deepseek | 65.62 | 70.6 | 69.2 | 37.6 | 79.7 | 68.0 | 70.1 | 63.1 | $0.094 | 1.0M |
| 25 | Qwen: Qwen3.6 27B qwen | 64.91 | 70.3 | 71.8 | 39.3 | 79.9 | 70.4 | 63.3 | 53.2 | $0.450 | 262K |
Best AI for each capability
How this index is built
Two layers, kept deliberately separate. The factual layer — price per token, context window, modality, launch date — comes from live provider data covering all 338 models, refreshed when the index rebuilds. The capability layer comes from LiveBench, an independent benchmark that publishes dated snapshots and resists contamination by rotating its questions. We compute the LLM Index Score from LiveBench's raw per-task results using weights we publish, and we show every component on every model page.
Nothing here is hand-tuned. No model gets a nudge, no score is adjusted after the fact, and where a third party publishes its own index we display it attributed to them rather than folding it into ours. When a model is missing benchmark data we say so and show no score, instead of guessing one — a partial score is not comparable to a complete one, and an estimate presented as a measurement is worse than no number at all.
Why component transparency matters
Every model leaderboard compresses many abilities into one number, and that compression is where the information goes missing. Two models with the same overall score can be completely different hires: one steady across the board, the other carried by a single very strong category while trailing everywhere else. If the leaderboard only shows you the total, you cannot tell them apart — and you will pick the wrong one for a narrow workload.
So the breakdown is not a footnote here, it is the point. Each model page prints all seven category scores, their weights, their contribution to the total, and the raw task scores underneath. You can add the contributions up yourself and land on the published score. A ranking you can check is worth more than a ranking you have to trust.
FAQ
Which AI model ranks first right now?
OpenAI: GPT-5.5 leads the index with an LLM Index Score of 80.49, computed from the LiveBench 2026_06_25 snapshot. The full ranking, including every category component, is on the leaderboard.
How many models does this leaderboard cover?
338 models carry full factual data — price, context window, modality and launch date. Of those, 26 also carry a computed LLM Index Score, because scoring requires complete benchmark coverage in the current snapshot.
Can I reproduce the LLM Index Score myself?
Yes, and that is the design goal. The inputs are public LiveBench task scores, the category weights are published on the methodology page, and every per-category and per-task value is printed on the model pages. Average the tasks in each category, apply the weights, and you get the same figure.
Why do some models show no score?
A model is scored only when every task in every category is present in the snapshot. Models released after the snapshot, or with partial results, keep their factual data and make no capability claim. We would rather show an honest gap than an estimate dressed up as a measurement.
Is the top-ranked model the right choice for me?
Often not. The overall score weights seven capabilities, and your workload probably does not. If you only care about code, or maths, or long-context retrieval, rank by that category instead — and check price, because capability and cost move independently.
How current is the data?
Factual data was retrieved 2026-07-21, and capability scores come from the LiveBench 2026_06_25 snapshot. Benchmarks are dated on purpose: they describe a model as measured on that date, not as it may behave after a provider-side update.