Head-to-head

Kimi K3 vs Fable 5

The largest open-weight model ever announced — 2.8 trillion parameters — against Anthropic's new Mythos-class ceiling. On the one scorecard that rates both the same way they finish within three points, and coding is a dead heat. But Fable 5 ships today at 3.3x the price, while Kimi K3's weights are an IOU: downloadable only from July 27.

The short version

This is the first pairing here where the closed model isn't Opus — Fable 5 is Anthropic's Mythos-class tier, positioned a notch above Opus 4.8 — and the surprise is how close the open challenger runs. On Artificial Analysis, the only evaluator that scores both under one methodology, Fable 5 leads the Intelligence Index 59.9 to 57.1 and the Agentic Index 52.8 to 50.1 — roughly three points each — while the Coding Index is a genuine dead heat, 76.5 to 76.2. So Fable is ahead, but by a low-single-digit margin, and not everywhere. Set that against the price sheet and the timing. Fable 5 lists at $10 / $50 per million tokens against Kimi K3's $3 / $15 — about 3.3x — and it is a shipped, hosted, multimodal model you can call right now. Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model (16 of 896 experts active) that Moonshot calls the first open-source model in the 3-trillion class — but its weights do not release until July 27, 2026, so today it is an API-only value play, not a model you can yet download and own. The question is not who is smarter — Fable, narrowly — but whether a three-point edge, file input and a managed frontier ceiling justify more than triple the token price, and whether Kimi's not-yet-open weights are a promise you want to build on.

Benchmark head-to-head

Only one evaluator scores both of these models under a single published methodology: Artificial Analysis, whose Intelligence Index v4.1 runs the same nine evaluations (GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR) against every model on its board, then reports a Coding Index and an Agentic Index alongside. Those three are the rows below — Kimi K3 and Fable 5 as captured in our 2026-07-21 snapshot (Kimi K3's Intelligence Index re-confirmed live on the Artificial Analysis model page at 57, matching the snapshot's 57.1 within rounding). We do not pair Moonshot's own headline numbers against Anthropic's, because the two are produced on different harnesses and are not directly comparable.

BenchmarkKimi K3Fable 5Margin
Intelligence Indexcomposite, 9 evals57.159.9Fable +2.8
Coding Indexcoding76.276.5≈ even (Fable +0.3)
Agentic Indextool use / agents50.152.8Fable +2.7

Artificial Analysis reports these for the max-effort configuration of each model, and the Intelligence Index is a rolled-up composite — a single number over nine evaluations — so it hides where within the suite each model pulls ahead or falls back. Read the three indices as a directional cross-vendor snapshot, not a task-by-task verdict. Both are reasoning models scored with thinking enabled; both, in fact, keep thinking on by default.

Two things are deliberately left off. First, we do not pair vendor suites: Anthropic reports Fable 5 as state-of-the-art on nearly all of its own software-engineering and knowledge-work evals, while third-party coverage has flagged Kimi K3 topping Fable 5 on a front-end coding arena — exactly the harness-dependent split that is the reason to show only a single evaluator that runs both the same way. The Artificial Analysis Coding Index, which does run both the same way, already puts them level. Second, there is no LLM Index Score row here: our own LiveBench-based composite covers Fable 5 (79.82) but Kimi K3 is absent from the snapshot we score, so a like-for-like native number does not exist for this pair — we show what a single source produced for both, and nothing invented to fill the gap.

The two models, side by side

Kimi K3Fable 5
WeightsOpen — full weights due July 27, 2026 (not yet downloadable)Closed — hosted only
VendorMoonshot AIAnthropic
Tier / positioningMoonshot's flagship — 'first open-source model in the 3T class'Mythos-class — a tier above Opus 4.8
ReleasedJuly 16, 2026June 9, 2026
API model IDkimi-k3claude-fable-5
Price (input / output)$3 / $15 per MTok$10 / $50 per MTok
Cached input$0.30 per MTok$1 per MTok
Context / max output1M in / up to 1M out (default 131K)1M in / 128K out
Parameters2.8T total / 16 of 896 experts active (MoE)Not disclosed
Knowledge cutoffNot publishedNot published
TokenizerMoonshot (counts differ from Claude)Claude (counts differ from Moonshot)
Reasoning controlAlways-on thinking; effort low → max (default max)Always-on thinking; effort low → max (default medium)
Input modalitiesText, image (video on the Kimi platform)Text, image, file
Safety routingNoneSome sensitive queries auto-answered by Opus 4.8 (<5% of sessions)
AvailabilityKimi API (OpenAI-compatible), kimi.com, Kimi app; third-party hostsClaude API, Amazon Bedrock, Google Cloud, Microsoft Foundry
Self-hostNot yet — weights due July 27, 2026, then downloadableNo — proprietary, hosted only

What the gap costs

Take a steady load — 10M input and 2M output tokens a day, the shape of a sustained agentic-coding workload — against each model's first-party API price:

Kimi K3Fable 5
Per day (10M in / 2M out)$60$200
Per 30-day month$1,800$6,000
Vs Fable 5 each month−$4,200

Prices are Moonshot's own API rate for Kimi K3 ($3 / $15) and Anthropic's for Fable 5 ($10 / $50) — about 3.3x on both input and output. The token counts assume an identical split on both, a simplification: the two use different tokenizers, so the same text splits into different numbers of tokens on each and the real gap moves a few percent either way. Both discount cache hits — Kimi K3 to $0.30 per 1M input, Fable 5 to $1. Once Kimi K3's weights ship on July 27 you can self-host and drop the per-token bill entirely, trading it for compute you run yourself — but at 2.8 trillion parameters that is a serious hardware commitment, not a laptop job.

Which should you pick?

Reach for Kimi K3 when…

  • Cost dominates at scale: at $3 / $15 it is about a third of Fable 5's per-token price for a capability gap that, on the comparable index, runs under three points and disappears entirely on coding.
  • You are betting on the open-weights roadmap — from July 27, 2026 you can download a 2.8-trillion-parameter model, self-host it, run it offline and pin a frozen version, which no hosted-only model can offer.
  • The work is text or image reasoning and coding where a dead-heat Coding Index (76.2 vs 76.5) means you are giving up nothing measurable on the task most teams care about most.
  • You want an OpenAI-compatible API and the option to leave hosted endpoints later — the flexibility of the largest open model in its class, at the lowest token price of the two.

Step up to Fable 5 when…

  • You want the measured ceiling and it has to be available today: Fable 5 leads the Artificial Analysis Intelligence and Agentic indices by roughly three points and is a shipped, hosted Mythos-class model, a tier Anthropic positions above Opus 4.8.
  • The job needs file input — Fable 5 accepts text, images and files, where Kimi K3's document workflows are thinner outside Moonshot's own platform.
  • You want a managed frontier model across Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry, with Anthropic's safety routing, rather than waiting on a July 27 weights drop and standing up 2.8T-parameter inference yourself.
  • Per-token cost is a rounding error against the value of the task, so a three-point edge and the reliability of a fully hosted model are worth more than three-times-lower tokens.

Before you commit

  • The open-weights advantage is a dated promise, not a present fact. Kimi K3's weights are announced for July 27, 2026 and are not downloadable before then, so today it is an API-only model like Fable 5 — the 'own it, self-host it' case only unlocks once the release lands, and on hardware that can hold 2.8T parameters.
  • The only cross-vendor benchmark that scores both — Artificial Analysis — puts Fable 5 ahead on intelligence and agentic by under three points and level on coding, but each is a rolled-up composite. A task that leans hard on one specific capability can diverge from what the index implies.
  • The two use different tokenizers, so cost and context comparisons are approximate: '1M tokens' of your text is not the same amount of text on each model.
  • Fable 5 ships a safety layer that routes some sensitive cybersecurity and biology queries to Opus 4.8 instead — on average under 5% of sessions. For most workloads it is invisible, but it means a small slice of Fable traffic is answered by a different model.
  • Input modalities differ at the edges: both take text and images, but only Fable 5 accepts file input, while Kimi K3 adds native video on Moonshot's own platform. Match the modality to the workload before the benchmark.

The full read

This is the first comparison on the site where the closed model is not Opus 4.8 — it is Fable 5, the safe-for-general-use member of Anthropic's Mythos class, which the company positions a step above the Opus line. The opponent is not a scaled-down open clone either: Kimi K3 is a 2.8-trillion-parameter Mixture-of-Experts model, activating 16 of 896 experts per token, that Moonshot calls the first open-source model in the three-trillion-parameter class. So the framing is not 'frontier versus budget.' It is a shipped, hosted premium ceiling against the largest open-weight model anyone has announced — where the open one is cheaper and, on the one comparable scorecard, barely behind.

On the measurable side the read is close, not a rout. Artificial Analysis — the only evaluator that runs both under one methodology — gives Fable 5 the Intelligence Index 59.9 to 57.1 and the Agentic Index 52.8 to 50.1, each about three points, and calls the Coding Index a draw at 76.5 to 76.2. Directionally Fable is ahead; practically, on coding, the two are interchangeable, and even the wider gaps are low single digits on a rolled-up composite. That matters because coding and agentic work is exactly where most teams spend their tokens, and it is where the open challenger holds level.

Then the two columns pull apart on everything a benchmark cannot score, and timing is the sharpest of them. Kimi K3's headline advantage — open weights you can download, self-host, run offline and freeze — is real but not yet available: Moonshot has committed to releasing the full weights on July 27, 2026, and until then Kimi K3 is an API-only model exactly like Fable 5. Even after the drop, holding 2.8 trillion parameters in memory is a data-centre undertaking, not a workstation one. So the open-weights case is a roadmap bet, strong in principle, that pays off later and only for teams with the hardware to cash it. Priced through Moonshot's own API it is $3 / $15 per million tokens against Fable 5's $10 / $50 — about a third of the cost, today, before any self-hosting enters the picture.

It is worth being explicit about why the benchmark table stops at the three Artificial Analysis indices. Anthropic reports Fable 5 as state-of-the-art on nearly all of its own software-engineering and knowledge-work evaluations; separately, third-party coverage has flagged Kimi K3 coming out ahead of Fable 5 on a front-end coding arena. Both can be true at once, because different harnesses, effort settings and scaffolds move the score — which is the whole reason we do not set one vendor's self-run numbers beside the other's. The one place both are measured identically, the Artificial Analysis Coding Index, already tells the honest version of that story: on coding, this is a tie.

The practical rule follows from price and timing rather than a winner. If your workload is cost-sensitive coding or reasoning, if a dead-heat Coding Index means you sacrifice nothing measurable, and if the open-weights roadmap is something you intend to use once July 27 lands, Kimi K3 is the default — the largest open model in its class at the lowest token price of the two. If you need the measured ceiling available right now, file input, a fully managed frontier model across Bedrock, Google Cloud and Microsoft Foundry, and you would rather not stand up trillion-parameter inference yourself, Fable 5 is the escalation and its three-point edge is worth the premium. These are vendor-reported and third-party figures, not our own hands-on runs, so read them as a well-sourced starting point and confirm on the workload you actually have — and, for Kimi K3, on whether the weights you are counting on have actually shipped.

Full index pages

FAQ

Is Kimi K3 or Fable 5 the better model?

On the one benchmark that scores both under a single methodology — Artificial Analysis — Fable 5 is ahead, but narrowly. It leads the Intelligence Index 59.9 to 57.1 and the Agentic Index 52.8 to 50.1, roughly three points each, and the Coding Index is a dead heat at 76.5 to 76.2. So Fable is the stronger model on the comparable numbers, but not on coding and not by much — and Kimi K3 costs about a third as much, which makes 'better' a question of whether a three-point edge is worth more than triple the token price.

Can I download and run Kimi K3 today?

Not yet. Moonshot has announced Kimi K3 as an open-weight model — it calls it the first open-source model in the three-trillion-parameter class — but the full weights are scheduled to release on July 27, 2026, and were not downloadable before then. Until the drop, Kimi K3 is an API-only model just like Fable 5. And even once the weights ship, self-hosting a 2.8-trillion-parameter model is a data-centre-scale job, so the 'own it, run it offline' advantage is real but gated behind both a release date and serious hardware.

Which is cheaper, and by how much?

Kimi K3, by about 3.3x. Through Moonshot's own API it lists at $3 per million input tokens and $15 output, against Fable 5's $10 / $50. On a sustained 10M-input / 2M-output day that is roughly $60 versus $200, or $1,800 versus $6,000 a month. Cache hits discount both — Kimi K3 to $0.30 per 1M input, Fable 5 to $1 — and once Kimi K3's weights ship you could self-host and remove the per-token bill entirely. Because the two use different tokenizers, treat the token-for-token figure as close rather than exact.

Is Fable 5 more capable than Opus 4.8?

Anthropic positions Fable 5 as a Mythos-class model — a tier above the generally available Opus line — and says its capabilities exceed those of any model it had previously made broadly available, with the lead growing on longer, more complex tasks. It ships with a safety layer that routes some sensitive cybersecurity and biology queries to Opus 4.8 instead, on average in under 5% of sessions. So for most work Fable 5 is Anthropic's strongest hosted option, sitting above Opus 4.8 rather than beside it.

Why isn't there an LLM Index Score comparison on this page?

Because we only have the number for one of the two. Our LLM Index Score is a composite built from the LiveBench task set, and Fable 5 carries one (79.82) while Kimi K3 is absent from the snapshot we score, so there is no like-for-like native figure to place beside it. Rather than estimate or half-report, we leave the section out and rely on the Artificial Analysis indices, which do measure both under one methodology. Kimi K3 will get an LLM Index Score once it appears in a snapshot we run.

Spec and reasoning behaviour for Kimi K3 from Moonshot's Kimi platform documentation, with pricing from Moonshot's API rate (via OpenRouter); for Fable 5 from Anthropic's Claude Fable 5 & Mythos 5 announcement and OpenRouter's model page. Cross-vendor benchmark figures from Artificial Analysis (snapshot 2026-07-21; Kimi K3's Intelligence Index re-confirmed live at 57). Kimi K3's open weights are announced for release July 27, 2026 and were not yet downloadable at publication. Last verified 2026-07-24. Vendor-reported and third-party data, not our own hands-on runs.