Ornith 1.0 vs Sonnet 5
An MIT-licensed, self-hosted coding specialist against a hosted general-purpose flagship. Ornith 1.0 is a four-size open family — 9B to 397B — that learns to write its own agent scaffolds and runs entirely on your hardware for no per-token fee; Sonnet 5 is a closed, metered generalist with a 1M-token window and image plus file input. For the first time on this site there is no scoreboard that rates both the same way, so this is a deployment-and-fit decision, not a ranking.
This pairing breaks the pattern of every earlier comparison here in two ways. First, the axis is specialist-versus-generalist, not open-versus-closed on a like-for-like flagship: Ornith 1.0 (DeepReinforce, released June 25, 2026) is a coding-only open family — 9B Dense, 31B Dense, 35B MoE and a 397B MoE flagship — post-trained on Gemma 4 and Qwen 3.5, shipped weights-first under an MIT license on Hugging Face, whose defining trick is self-scaffolding: it learns during training to generate its own task-specific agent harnesses rather than run inside a human-written one. Sonnet 5 is Anthropic's closed, hosted general-purpose model, priced at $2 / $10 per million tokens (introductory, rising to $3 / $15 on September 1, 2026), with a 1M-token context window and text, image and file input. Second — and this is the honest crux — there is no third-party evaluator that scores both under one methodology: Ornith 1.0 is not on Artificial Analysis (the cross-vendor index we lean on elsewhere) and not in the LiveBench snapshot behind our own LLM Index Score, so unlike our other pages we refuse to publish a head-to-head benchmark ranking. What can be compared is concrete and numeric: deployment (download-and-own vs rent-by-the-token), price ($0 per token, self-host compute, vs $2–3 / $10–15), context (Ornith's 256K vs Sonnet 5's 1M), and reach (a text-first coding specialist vs a 1M-context multimodal generalist). If you want to own a coding model, size it to your hardware, run it offline and pay no token bill, Ornith 1.0; if you want a hosted generalist with four times the context, image and file input, and no GPUs to operate, Sonnet 5.
Benchmark head-to-head
Every other comparison on this site opens with a benchmark table, because a single evaluator scored both models the same way. This one cannot. Ornith 1.0 is absent from Artificial Analysis — the cross-vendor Intelligence/Coding/Agentic index we use elsewhere lists no Ornith entry — and it is absent from the LiveBench snapshot behind our own LLM Index Score, so there is no source that measures Ornith and Sonnet 5 under one methodology. We therefore do not publish a winner-and-margin ranking for this pair. What follows is each vendor's OWN self-reported figure on the coding benchmarks they both happen to name — DeepReinforce's for the Ornith-1.0-397B flagship, Anthropic's for Sonnet 5 from its System Card. These come off different harnesses, effort settings and scaffolds, so they are shown side by side for reference only, marked as a tie in every row on purpose: read them as two separate self-reports, never as a head-to-head.
| Benchmark | Ornith 1.0 | Sonnet 5 | Margin |
|---|---|---|---|
| SWE-bench Verifiedself-reported | 82.4 | 85.2% | Different harnesses — not comparable |
| SWE-bench Proself-reported | 62.2 | 63.2% | Different harnesses — not comparable |
| SWE-bench Multilingualself-reported | 78.9 | 78.3% | Different harnesses — not comparable |
| Terminal-Bench 2.1self-reported | 77.5 | 80.4% | Different harnesses — not comparable |
Do not read the closeness of these numbers as a verdict. The Ornith column is DeepReinforce's own reporting for the 397B MoE flagship (its Terminal-Bench figure is the Terminus-2 harness; it also reports 78.2 under a Claude Code harness — the harness alone moves the score); the Sonnet 5 column is Anthropic's own System Card, run under adaptive thinking at max effort. Neither party ran the other's model, and no independent lab ran both, so a two-point difference here carries none of the same-source precision the tables on our other comparison pages do. They are included only so each model's self-reported standing is visible; they settle nothing between the two.
There is deliberately no benchmark ranking and no LLM Index Score section on this page. Our LLM Index Score is a composite over the LiveBench task set, and while Sonnet 5 carries one (76.08), Ornith 1.0 is not in the snapshot we score, so a like-for-like native figure does not exist — and we do not estimate one. Likewise, Artificial Analysis publishes Intelligence, Coding and Agentic indices for Sonnet 5 (53.4 / 71.5 / 46.7 in our 2026-07-21 snapshot) but has no Ornith 1.0 entry to place beside them. Rather than pair a scored model against an unscored one, or cobble DeepReinforce's self-run coding numbers against Anthropic's, we show what each vendor reports for itself and rest the actual decision on the axes that are genuinely comparable — deployment, license, price, context and reach — below.
The two models, side by side
| Ornith 1.0 | Sonnet 5 | |
|---|---|---|
| Weights | Open — MIT license, downloadable (weights-first) | Closed — hosted only |
| Vendor | DeepReinforce | Anthropic |
| Kind | Coding specialist — a four-size family | General-purpose — single hosted model |
| Model sizes | 9B Dense · 31B Dense · 35B MoE · 397B MoE (flagship compared) | One tier (claude-sonnet-5) |
| Released | June 25, 2026 | June 30, 2026 |
| Base models | Post-trained on Gemma 4 + Qwen 3.5 | Anthropic proprietary |
| Price (input / output) | $0 per token — self-host only (no hosted API) | $2 / $10 intro → $3 / $15 (Sep 1, 2026) per MTok |
| Cached input | n/a (self-host) | $0.20 per MTok |
| Context / max output | 256K in (262,144 tokens) | 1M in / 128K out |
| Parameters | 397B total (MoE); active count not published | Not disclosed |
| Knowledge cutoff | Not published | Jan 2026 |
| Tokenizer | Gemma 4 / Qwen 3.5 base (differs from Claude) | Claude (differs from Ornith) |
| Agent control | Self-scaffolding — writes its own task-specific agent harness | Adaptive thinking (configurable effort) |
| Fast mode | n/a — throughput is whatever you provision | No (Opus-tier only) |
| Input modalities | Text, image (per HF model cards); coding-focused | Text, image, file |
| Availability | Hugging Face (self-host via vLLM / SGLang / Ollama / LM Studio) | Claude API, Bedrock, Google Cloud, Microsoft Foundry; default on Free/Pro & Claude Code |
| Self-host | Yes — download weights, run offline | No — closed, hosted only |
What the gap costs
There is no per-token price to compare, because Ornith 1.0 has no hosted API — you run the weights yourself. So the honest cost picture is a metered bill on one side and a compute bill on the other. Take Sonnet 5's first-party rate against a steady 10M-input / 2M-output day, the shape of a sustained agentic-coding workload:
| Ornith 1.0 | Sonnet 5 | |
|---|---|---|
| Per day (10M in / 2M out) | $0 per token (self-host) | $40 intro · $60 standard |
| Per 30-day month | $0 per token (self-host) | $1,200 intro · $1,800 standard |
| What you pay instead | GPUs / inference compute you run | Nothing but the token bill |
Sonnet 5's figures use Anthropic's own API price ($2 / $10 introductory through August 31, 2026; $3 / $15 from September 1); cache reads bill at $0.20 per 1M input. Ornith 1.0's per-token cost is genuinely zero — it is MIT-licensed and self-hosted — but that is not free: you provision and run the hardware, and at 397B parameters (MoE) that is a data-centre-class commitment, while the 9B and 31B Dense members are small enough for a single accelerator. The trade is a predictable metered bill you never operate against a fixed compute footprint you own outright. Because the two use different tokenizers, any token-for-token intuition is approximate regardless.
Which should you pick?
Reach for Ornith 1.0 when…
- The work is agentic coding specifically, and you want a model tuned for it — self-scaffolding lets it generate its own task harnesses, and DeepReinforce reports strong coding results down to the 9B and 31B sizes, not just the 397B flagship.
- You need to own the model: MIT-licensed weights you can download, self-host, run air-gapped or offline, keep your code in-house, and pin a frozen version — none of which a hosted model can offer.
- Token cost at scale is the constraint, or unpredictable: self-hosting removes the per-token bill entirely in exchange for compute you run, and the four-size family lets you fit the model to the hardware you actually have.
- You want no vendor lock-in and full commercial freedom — an MIT coding model you can embed, modify and ship without a metered API in the loop.
Step up to Sonnet 5 when…
- You need a generalist, not just a coder — chat, reasoning, mixed knowledge work and long documents — where a coding-specialist family is the wrong tool regardless of its benchmark self-reports.
- Context or modality decides it: Sonnet 5's 1M-token window is four times Ornith's 256K, and it takes image and file input, so whole-repo context and document workflows land on Sonnet 5.
- You do not want to operate inference — no GPUs, no serving stack, no uptime to own — and would rather rent a hosted, managed model across Claude API, Bedrock, Google Cloud and Microsoft Foundry.
- You value a measured, independently benchmarked model: Sonnet 5 carries an Artificial Analysis profile and our own LLM Index Score (76.08), whereas Ornith 1.0 has neither, so Sonnet 5's standing is the one you can actually verify against third parties.
Before you commit
- No source scores both models the same way. Ornith 1.0 is absent from Artificial Analysis and from our LiveBench snapshot, so there is no independent head-to-head — the coding-benchmark rows above are each vendor's self-report on different harnesses and are not comparable. Treat any capability claim here as directional at best.
- Ornith 1.0 is a coding specialist, not a general model. It is post-trained specifically for agentic software work; it is not positioned for open-ended chat, long-document reasoning or the breadth Sonnet 5 covers. Match it to coding, not to a generalist role.
- Context is not equal: Ornith 1.0's window is 256K (262,144 tokens) against Sonnet 5's 1M — roughly a quarter. Whole-monorepo or very-long-context work favours Sonnet 5 outright.
- Input reach differs. Sonnet 5 accepts text, images and files; Ornith 1.0's model cards expose text and image but it is built and marketed as text-first coding, so document/file workflows are a Sonnet 5 strength.
- Owning Ornith 1.0 means operating it. The MIT weights remove the per-token bill and any vendor lock-in, but you take on GPU provisioning, inference tuning and uptime — a real engineering cost that a hosted model like Sonnet 5 absorbs for you.
The full read
Every earlier comparison on this site has been, at bottom, a like-for-like: two general-purpose models, one benchmark that scored them the same way, and a margin to argue over. This one is neither. Ornith 1.0 is not a scaled general model — it is a coding specialist, a family DeepReinforce released weights-first on June 25, 2026 in four sizes (9B Dense, 31B Dense, 35B MoE and a 397B MoE flagship), post-trained on Gemma 4 and Qwen 3.5, all under an MIT license. Its defining idea is self-scaffolding: instead of running inside a fixed human-written agent harness, it learns during training to produce its own task-specific scaffolds alongside its solutions, jointly optimising the two. Sonnet 5, by contrast, is Anthropic's closed, hosted, general-purpose model — a metered API you call, not a checkpoint you download. So the comparison is specialist-you-run against generalist-you-rent, and the two barely occupy the same category.
The honest headline is what is missing. On our other pages, a third party — Artificial Analysis, or the LiveBench set behind our own LLM Index Score — scored both models under one methodology, and the whole exercise was reading the margin. Here there is no such source. Ornith 1.0 has no Artificial Analysis entry and is not in the LiveBench snapshot we score, so no independent evaluator has measured it and Sonnet 5 side by side. We could have pasted DeepReinforce's self-reported coding numbers next to Anthropic's System Card figures and called it a benchmark table — SWE-bench Verified reads 82.4 for the Ornith flagship and 85.2 for Sonnet 5, the rest are within a point or two — but those come off different harnesses (DeepReinforce even reports two different Terminal-Bench numbers for its own model depending on the scaffold), and pairing them would manufacture a precision that does not exist. So we mark every row a tie on purpose and rank nothing. The decision has to rest on what is actually comparable.
What is comparable is concrete. Deployment: Ornith 1.0 is downloadable, MIT-licensed weights you host yourself — offline, air-gapped, frozen at a version, with your code never leaving your infrastructure — while Sonnet 5 is hosted-only, reached through the Claude API, Bedrock, Google Cloud or Microsoft Foundry. Price: Ornith has no hosted API and therefore no per-token cost at all, trading the bill for compute you provision; Sonnet 5 lists at $2 / $10 per million tokens introductory, $3 / $15 from September 1, 2026, which on a sustained 10M-input / 2M-output day is $40–$60. Context: Ornith's window is 256K tokens against Sonnet 5's 1M — a real, sourced, four-to-one gap that whole-repo work will feel. Reach: Ornith is a text-first coding model (its cards expose image input from the Gemma 4 base, but it is built for code), while Sonnet 5 is a multimodal generalist that also takes files. None of these is a benchmark point, and all of them are decision-grade.
The four-size family is itself part of the argument, and it is where the specialist framing pays off. Because Ornith ships from 9B to 397B, you fit the model to the hardware and latency budget you actually have: the 9B and 31B Dense members run on a single accelerator for edge or low-latency use, while the 397B MoE flagship is a data-centre commitment reserved for maximum capability. Self-scaffolding is what lets DeepReinforce claim useful agentic-coding behaviour that far down the size curve — the model brings its own harness rather than depending on one you build. Sonnet 5 offers none of that surface: it is one hosted tier, and you tune it with an effort dial, not by choosing a checkpoint. If your constraint is 'run a capable coder on this specific box, offline, forever,' that is a shape only the open family can take.
So the practical rule is not 'which is better' — with no shared benchmark, that question has no honest answer here — but 'which shape fits.' Choose Ornith 1.0 when the job is agentic coding, when owning the model matters (self-host, offline, data-in-house, a frozen version, zero token bill), and when you can operate inference yourself; the MIT license and the size ladder are the whole point. Choose Sonnet 5 when you need a generalist rather than a coder, when the 1M context or image-and-file input decides it, or when you would simply rather rent a hosted, independently benchmarked model than stand up a serving stack. These are vendor self-reported specs and figures, not our own hands-on runs, and for Ornith 1.0 there is no third-party score to lean on at all — so treat this as a well-sourced framing of the trade, and confirm on your own coding workload and, for Ornith, on the hardware you would actually run it on.
Full index pages
FAQ
Is Ornith 1.0 or Sonnet 5 the better model?
There is no honest single answer, because no source scores both the same way — Ornith 1.0 is not on Artificial Analysis or in the LiveBench snapshot behind our LLM Index Score. They are also different kinds of model: Ornith 1.0 is an MIT-licensed, self-hosted coding specialist (a four-size open family), while Sonnet 5 is a closed, hosted general-purpose model. So it is a fit decision, not a ranking — pick Ornith to own and self-host a coding model, pick Sonnet 5 for a hosted generalist with more context and multimodal input.
Can I self-host Ornith 1.0 but not Sonnet 5?
Yes. Ornith 1.0 ships weights-first under an MIT license on Hugging Face in four sizes (9B Dense, 31B Dense, 35B MoE, 397B MoE), so you can download it and run it offline, air-gapped, with vLLM, SGLang, Ollama or LM Studio, and pin a frozen version. Sonnet 5 is closed and hosted-only — available through the Claude API, Bedrock, Google Cloud and Microsoft Foundry, but never as downloadable weights. Owning Ornith also means operating it: you provision the GPUs and run the inference yourself.
Which is cheaper to run?
Ornith 1.0 has no per-token price at all — it is MIT-licensed and self-hosted, so there is no API bill. Sonnet 5 lists at $2 / $10 per million tokens introductory (rising to $3 / $15 on September 1, 2026), which on a sustained 10M-input / 2M-output day is about $40–$60. But 'no token cost' is not 'free': with Ornith you pay for the hardware and inference compute you run, and the 397B flagship is a data-centre-class footprint, whereas the 9B and 31B sizes fit a single accelerator. It is a metered bill versus a compute footprint you own.
Why is there no benchmark ranking or LLM Index Score on this page?
Because no evaluator measures both models the same way. Ornith 1.0 has no Artificial Analysis entry and is absent from the LiveBench snapshot our LLM Index Score is built from, so pairing it against Sonnet 5's scores would mean either ranking a scored model against an unscored one or cobbling DeepReinforce's self-run coding numbers against Anthropic's System Card — different harnesses, false precision. We show each vendor's own self-reported coding figures for reference, marked as ties, and rest the decision on what is genuinely comparable: license, deployment, price, context and reach.
What is self-scaffolding, and what does Ornith 1.0 give up for it?
Self-scaffolding is Ornith 1.0's defining training method: instead of running inside a fixed, human-written agent harness, it learns to generate its own task-specific scaffolds alongside its solutions, optimising both together — which DeepReinforce credits for strong agentic-coding results even at the smaller 9B and 31B sizes. What it gives up is breadth and reach. Ornith is a text-first coding specialist with a 256K context window (a quarter of Sonnet 5's 1M), no hosted option, and no independent benchmark; Sonnet 5 is a multimodal generalist with file input, four times the context and a verifiable third-party profile. Ornith trades generality and convenience for ownership and coding focus.
Spec, license and benchmark figures for Ornith 1.0 from DeepReinforce's primary sources — the Ornith-1.0-397B and -9B Hugging Face model cards, the Ornith-1 GitHub repository, and the official ornith.site — covering the MIT license, four sizes (9B/31B Dense, 35B/397B MoE), Gemma 4 + Qwen 3.5 base, 256K (262,144-token) context, self-hosting (no hosted API or per-token price) and self-reported coding benchmarks. Sonnet 5 spec, pricing and benchmark figures from the Claude Sonnet 5 System Card, Anthropic's pricing, and our 2026-07-21 snapshot (LLM Index Score 76.08; Artificial Analysis indices). No third-party evaluator scores both models under one methodology — Ornith 1.0 is not on Artificial Analysis or in our LiveBench snapshot — so no head-to-head ranking is published. Last verified 2026-07-25. Vendor-reported data, not our own hands-on runs.