# Claude Fable 5

Claude Fable 5 is an AI model released by Anthropic on Jun 9 2026. It has a 1M context window. At release it scored 92% on Next.js Evals, 63.3% on HealthBench Professional and 1932 on GDPval-AA.

## Facts

| Field | Value |
| --- | --- |
| Model | Claude Fable 5 |
| Developer | Anthropic |
| Release date | Tuesday, Jun 9 2026 |
| Licensing | Proprietary |
| Context window | 1M |

## API pricing

All rates in USD per 1,000,000 tokens, pay-as-you-go.

| Tier | Input | Cached input | 5 min cache write | 1 hr cache write | Output |
| --- | --- | --- | --- | --- | --- |
| All documented contexts | $10.00 | $1.00 | $12.50 | $20.00 | $50.00 |

Verified August 18, 2026 against the first-party source: https://platform.claude.com/docs/en/about-claude/pricing

## Benchmark scores published at release

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| BullshitBench v2 | 54% | [BullshitBench](https://github.com/petergpt/bullshit-benchmark) | Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better. |
| Gray Swan IPI (k = 1) | 0.4% | Lab | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. Gray Swan's indirect prompt injection benchmark measures how often such an attack succeeds when the attacker gets a single try. Lower is better. |
| Gray Swan IPI (k = 10) | 2.3% | Lab | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. This variant gives the attacker 10 tries and counts an attack as successful if any of them works. Lower is better. |
| Gray Swan IPI (k = 15) | 2.8% | Lab | Attackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. This variant gives the attacker 15 tries and counts an attack as successful if any of them works. Lower is better. |
| SWE-Bench Pro | 80.3% | Lab | Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better. |
| SWE-Bench Verified | 95.5% | Lab | Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better. |
| SWE-Bench Multilingual | 86.6% | Lab | Like SWE-Bench, but the coding problems span many programming languages, not just one. Tests how broadly the AI can code. Higher is better. |
| SWE-Bench Multimodal | 54.1% | Lab | Real bug reports that arrive with pictures attached — a screenshot, a mockup, a page rendering wrongly — so the AI has to read the image as well as the code to work out what to fix. Higher is better. |
| DeepSWE 1.1 | 70% | Lab | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. (Version 1.1 of the test.) Higher is better. |
| DeepSWE 1.0 | 66.1% | Lab | Artificial Analysis' independent test of deep, agentic software-engineering work — the AI has to plan and carry out substantial coding tasks end to end. Higher is better. |
| Next.js Evals | 92% | [Next.js Evals](https://nextjs.org/evals) | Vercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better. |
| Frontier-Bench v0.1 | 33.8% | Lab | A hard, ever-evolving set of real computer tasks — coding, system administration, data work, and more — that an AI agent has to complete on its own. Run by the Harbor / Laude Institute team as the successor to Terminal-Bench (v0.1 is the first release of the task set). The score is the share of tasks solved. Higher is better. |
| Terminal-Bench 4.0 | 44.55% | [Terminal-Bench](https://github.com/harbor-framework/terminal-bench), retrieved 2026-08-30 | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Version 4.0 recalibrated how much time, CPU and memory each task gets, removed eight tasks and fixed nineteen, so fewer runs fail for reasons that have nothing to do with the model. Scores are not comparable with earlier versions. Higher is better. |
| Terminal-Bench 2.1 | 88% | Lab | Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better. |
| Terminal-Bench-Science 0.1 | 24.7% | Lab | The same command-line setup as Terminal-Bench, pointed at scientific work: the AI has to drive research tooling and computational workflows through to a result, rather than administer a machine. Version 0.1 is the first release of the task set, and scores run lower than on the general board. Higher is better. |
| BrowseComp | 86.9% | [BenchLM](https://benchlm.ai), retrieved 2026-07-11 | Can the AI browse the web and track down hard-to-find answers? Higher is better. |
| Humanity's Last Exam (with tools) | 64.5% | [BenchLM](https://benchlm.ai), retrieved 2026-07-11 | Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better. |
| OSWorld-Verified | 85% | Lab | Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Higher is better. |
| Harvey's Legal Agent Benchmark | 11.25% | Lab | Harvey's test of whether an AI agent can complete real legal work — drafting and reviewing documents, working with spreadsheets and presentations, and navigating files the way a lawyer's assistant would. Higher is better. |
| TaxEval v2 | 76.94% | Lab | A set of real tax questions created by Vals AI — can the AI give accurate answers about tax rules and filings? Higher is better. |
| HealthBench Professional | 63.3% | Lab | Realistic health conversations graded against detailed rubrics written by physicians — can the AI respond the way a careful medical professional would? Higher is better. |
| MedScribe | 88.52% | Lab | Can the AI support doctors with their administrative work, like notes and paperwork? Created by Vals AI. Higher is better. |
| GDPval-AA | 1932 | Lab | Measures how well the AI does economically valuable knowledge work, judged against human experts. Shown as a rating (like a chess Elo) — higher is better. |
| GDPval-AA v2 | 1760 | Lab | economically valuable knowledge work (v2, re-based Elo) |
| threejseval | 1552 | [threejseval](https://threejseval.com/ranking) | Every model gets the same prompt — "the Eiffel Tower", "a glass fishbowl", "a robot arm picking toys into a box" — and builds a 3D scene in Three.js. Real people then see two scenes side by side, names hidden, and vote for the one they prefer. The votes become a chess-style Elo rating on threejseval.com, averaged across all the prompts. It measures whether the scene looks and moves right to a human eye, not whether the code passes a test. Higher is better. |

## About Claude Fable 5

Claude Fable 5, released June 9, 2026, is the first model of Anthropic's Claude 5 generation and the debut of the Mythos-class tier that sits above Opus in the lineup. It posted the largest single-release jump in the tracker's dataset: 95.5% on SWE-Bench Verified, 80.3% on SWE-Bench Pro, 88.0% on Terminal-Bench 2.1, and a WebDev Arena coding Elo of 1649 — over 80 points clear of the next model at launch.

Fable 5 ships with a 1M-token context window and leads agentic evaluations including OSWorld-Verified (85.0%) and Artificial Analysis GDPval (1932 Elo). The same underlying model without those additional dual-use safety measures shipped the same day as Claude Mythos 5, for vetted organisations only. Claude Sonnet 5 followed on June 30, 2026, and Claude Opus 5 brought the 5-series to the Opus tier on July 24, 2026 at half Fable's price. Fable 5 held the top of the lineup for twelve weeks, until Claude Fable 5.1 succeeded it on September 1, 2026 at the same price.

## Questions and answers

### When was Claude Fable 5 released?

Claude Fable 5 was released by Anthropic on Tuesday, Jun 9 2026.

### Who made Claude Fable 5?

Claude Fable 5 was built by Anthropic. AI safety company building the Claude family of models. Founded in 2021 by former OpenAI researchers.

### How much does Claude Fable 5 cost?

Claude Fable 5 costs $10.00 per million input tokens and $50.00 per million output tokens through the Anthropic API. Cached input is $1.00 per million tokens. Rates are pay-as-you-go API prices verified against Anthropic's published pricing on August 18, 2026.

### What benchmark scores did Claude Fable 5 get?

Claude Fable 5 reports 25 tracked benchmark scores — BullshitBench v2: 54%; Gray Swan IPI (k = 1): 0.4%; Gray Swan IPI (k = 10): 2.3%; Gray Swan IPI (k = 15): 2.8%; SWE-Bench Pro: 80.3%; SWE-Bench Verified: 95.5%; SWE-Bench Multilingual: 86.6%; SWE-Bench Multimodal: 54.1%; DeepSWE 1.1: 70%; DeepSWE 1.0: 66.1%; Next.js Evals: 92%; Frontier-Bench v0.1: 33.8%; Terminal-Bench 4.0: 44.55%; Terminal-Bench 2.1: 88%; Terminal-Bench-Science 0.1: 24.7%; BrowseComp: 86.9%; Humanity's Last Exam (with tools): 64.5%; OSWorld-Verified: 85%; Harvey's Legal Agent Benchmark: 11.25%; TaxEval v2: 76.94%; HealthBench Professional: 63.3%; MedScribe: 88.52%; GDPval-AA: 1932; GDPval-AA v2: 1760; threejseval: 1552. Scores are the figures published at release by Anthropic. It holds the best score among all models tracked here on Next.js Evals, HealthBench Professional and GDPval-AA.

### What is the context window of Claude Fable 5?

Claude Fable 5 has a context window of 1M. That is the maximum amount of input plus output the model can hold in a single request.

### Is Claude Fable 5 open source?

No. Claude Fable 5 is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after Claude Fable 5?

Anthropic's previous tracked release was Claude Opus 4.8 on May 28 2026, 12 days earlier. It was followed by Claude Mythos 5 on Jun 9 2026.


---

Canonical page: https://aireleasetracker.com/model/anthropic/claude-fable-5
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
