About LLM Index

LLM Index is an independent index of AI models. We publish one thing: comparable, sourced data on what each model can do, what it costs, and how much context it takes — refreshed as new models ship.

Why we exist

Choosing a model means reconciling scattered specs, stale benchmark claims and pricing pages that contradict each other. The existing leaderboards solve part of that — they rank models — but they hand you a single composite number and ask you to trust it. You cannot see what went into it, you cannot tell a broad performer from a lopsided one, and you certainly cannot recompute it. That opacity is fine right up until the number decides how you spend real money.

So we publish the weights, show every component on every model page, and name every source. If you disagree with how we weight coding against reasoning, you can see exactly what we did and reweight it yourself from the same public inputs. A ranking you can audit is more useful than a ranking you have to believe.

Where the data comes from

Two layers, kept separate on purpose. Factual data — price per input and output token, context window, modality, weight availability, launch date — comes from live provider data rather than marketing pages, and covers every model in the index. Capability data comes from LiveBench, an independent benchmark that publishes dated snapshots and rotates its questions to resist contamination, which matters because a benchmark models have memorised measures memory, not ability.

The LLM Index Score is ours: LiveBench supplies the raw per-task results and the category grouping, we supply the weights and publish them. The category means are plain averages, the total is a weighted sum, and no step is renormalised or adjusted for an individual model. Everything is recomputable by hand from public data.

What we do not do

We do not run private evaluations and present them as objective truth. We do not fold another company's proprietary index into our score and call it ours — where a third party publishes a score we show it attributed to them, clearly separated from our own. We do not publish a score for a model whose benchmark data is incomplete, because a partial score is not comparable to a complete one and quietly ranking such a model last would be a lie of omission.

We also do not pretend benchmarks are the whole story. Scores say nothing about latency, rate limits, regional availability, safety behaviour, or how a provider holds up under sustained load. Those things decide whether a model is usable in production, and none of them appear on this site. Treat what we publish as the measurable half of the decision, not the decision.

How we think about being wrong

An index is a claim about reality, and claims about reality go stale. Provider pricing changes without announcement. Models get updated behind stable names, so today's measurement can describe yesterday's model. Benchmarks themselves age as their questions leak into training data. None of that is avoidable, but all of it is disclosable — so every figure here carries the date it was retrieved and the snapshot it came from, and you can judge for yourself how much weight an older number deserves.

The failure mode we most want to avoid is the confident gap-fill: inventing a plausible number where a real one is missing. It is easy to interpolate a score from a model's family or parameter count, it looks exactly like a measurement, and it is not one. We would rather show an obvious hole than a smooth surface with something made up underneath.

Who this is for

Engineers picking a model for a product, teams estimating what a workload will cost before committing to it, and anyone who has tried to compare two models and found the answer scattered across a dozen pages that disagree. If you want a single confident recommendation with no working shown, this is the wrong site — there are plenty of those, and they are easier to read. If you want to see the arithmetic and reach your own conclusion, that is exactly what this is built for.

How the index stays current

The factual layer rebuilds against live provider data, so new models, price changes and context-window bumps land without waiting on anyone to write a post about them. The capability layer moves at the pace of the benchmark: LiveBench publishes dated snapshots, and when a new one lands we recompute every score against it and update the snapshot date shown throughout the site. Between snapshots, newly released models appear immediately with full factual data and no score — visible, comparable on price and context, and honestly marked as unmeasured.

Corrections

If a number here is wrong, we want to know — a data index that resists correction is just an opinion with tables. Reach us at hello@llm-index.com with the page and the discrepancy, and see the methodology page for exactly how any score on this site is derived.