Best AI for Agentic Coding
26 models ranked by their Agentic Coding score from the LiveBench 2026_06_25 snapshot. Agentic Coding carries 15% of the overall LLM Index Score. Prices are per million input tokens.
| # | Model | Agentic Coding | Overall | Price | Context |
|---|---|---|---|---|---|
| 1 | Anthropic: Claude Opus 4.8 anthropic | 56.1 | 79.67 | $5.00 | 1M |
| 2 | OpenAI: GPT-5.4 openai | 53.8 | 78.64 | $2.50 | 1.1M |
| 3 | OpenAI: GPT-5.5 openai | 52.1 | 80.49 | $5.00 | 1.1M |
| 4 | Z.ai: GLM 5.2 z-ai | 51.9 | 73.84 | $0.952 | 1.0M |
| 5 | Anthropic: Claude Sonnet 5 anthropic | 51.1 | 76.08 | $2.00 | 1M |
| 6 | Anthropic: Claude Fable 5 anthropic | 50.7 | 79.82 | $10.00 | 1M |
| 7 | Anthropic: Claude Opus 4.7 anthropic | 50.7 | 77.48 | $5.00 | 1M |
| 8 | OpenAI: GPT-5.2 openai | 50.3 | 75.45 | $1.75 | 400K |
| 9 | OpenAI: GPT-5.2-Codex openai | 49.4 | 74.55 | $1.75 | 400K |
| 10 | Anthropic: Claude Opus 4.6 anthropic | 49.0 | 75.35 | $5.00 | 1M |
| 11 | Google: Gemini 3.5 Flash google | 49.0 | 74.46 | $1.50 | 1.0M |
| 12 | MoonshotAI: Kimi K2.6 moonshotai | 46.9 | 71.06 | $0.684 | 262K |
| 13 | OpenAI: GPT-5.4 Nano openai | 46.8 | 70.63 | $0.200 | 400K |
| 14 | xAI: Grok Build 0.1 x-ai | 45.8 | 68.11 | $1.00 | 256K |
| 15 | MoonshotAI: Kimi K2.7 Code moonshotai | 45.7 | 69.26 | $0.820 | 262K |
| 16 | Google: Gemini 3.1 Pro Preview google | 45.4 | 76.96 | $2.00 | 1.0M |
| 17 | Qwen: Qwen3.7 Max qwen | 43.6 | 73.27 | $1.48 | 1M |
| 18 | Anthropic: Claude Sonnet 4.6 anthropic | 42.6 | 73.91 | $3.00 | 1M |
| 19 | DeepSeek: DeepSeek V4 Pro deepseek | 42.6 | 72.26 | $0.435 | 1.0M |
| 20 | OpenAI: GPT-5.4 Mini openai | 41.7 | 66.72 | $0.750 | 400K |
| 21 | Qwen: Qwen3.6 Plus qwen | 41.4 | 69.47 | $0.325 | 1M |
| 22 | MiniMax: MiniMax M3 minimax | 40.7 | 67.63 | $0.300 | 1.0M |
| 23 | Anthropic: Claude Opus 4.5 anthropic | 39.7 | 73.03 | $5.00 | 200K |
| 24 | Qwen: Qwen3.6 27B qwen | 39.3 | 64.91 | $0.450 | 262K |
| 25 | DeepSeek: DeepSeek V4 Flash deepseek | 37.6 | 65.62 | $0.094 | 1.0M |
| 26 | xAI: Grok 4.3 x-ai | 18.5 | 62.08 | $1.25 | 1M |
What this ranking actually measures
The Agentic Coding score is not a vibe or an editorial opinion. It is the mean of 3 specific LiveBench tasks, each scored 0–100 and run against every model in the snapshot under the same conditions:
javascripttypescriptpython
Because a category score is a plain mean, a model can rank highly here while being uneven underneath — a strong average may hide one weak task. Every model page lists all 3 raw task scores separately, so you can check whether a lead is broad or carried by a single result. That matters when your workload leans on one specific ability rather than the category as a whole.
This category contributes 15% of the overall LLM Index Score, so a model at the top of this table is not automatically the best model overall — and a model that wins overall may sit mid-table here. If agentic coding is the job you are hiring a model for, rank by this column rather than by the overall score.
Reading the price and context columns
Price is per million input tokens, taken from live provider pricing rather than a marketing page, and it moves independently of capability — the top model on this table is frequently not the cheapest, and the gap between rank 1 and rank 3 is often far smaller than the gap in cost. Context is the maximum window the model accepts; a large window matters for long-document and repository-scale work, and is close to irrelevant for short prompts. Neither column feeds the score. They are shown alongside it because a ranking without cost is only half a decision.
Scores come from the LiveBench 2026_06_25 snapshot. Models released after that snapshot appear in the index with full factual data but no score — we do not estimate a score for a model we have no measurements for. See methodology for the weights and the exact formula.
FAQ
Which AI model is best for agentic coding?
Anthropic: Claude Opus 4.8 leads on Agentic Coding with a score of 56.1 in the LiveBench 2026_06_25 snapshot, ahead of OpenAI: GPT-5.4. That is a measurement from a fixed set of tasks, not an editorial pick.
How is the Agentic Coding ranking calculated?
Each model's Agentic Coding score is the mean of its raw LiveBench tasks in that category (javascript, typescript, python), each scored 0-100. That category mean then contributes 15% of the overall LLM Index Score. Nothing is hand-adjusted per model.
Is the highest-scoring model here also the best overall?
Not necessarily. Agentic Coding is only 15% of the overall score, so a model can top this table and rank lower overall, or win overall while sitting mid-table here. Rank by this column when agentic coding is the specific job you need done.
Why do some models show no score?
A model is scored only when every task in every category is present in the snapshot. Models released after the snapshot, or missing any task, keep their factual data (price, context, modality) and make no capability claim. We do not estimate a score from partial results, because a partial score is not comparable to a complete one.
Does a higher score justify a higher price?
That is your call, and it is why price sits next to the score. Capability and cost move independently: the gap between the first and third model on this table is often small, while the price gap between them can be several times over. For high-volume work the cheaper model is frequently the correct choice.