Claude Fable 5.1 is a Anthropic closed-API catalog model. CompareLLM Coding Reported rating is 1620. Active context window extends to 1M tokens. Commercial API token pricing is listed at $10/1M tok input and $50/1M tok output per 1M tokens. Anthropic frontier model for coding, agents, long-context reasoning, and multimodal document work. Numbers below are dated snapshots from empirical benchmark harnesses.
Independent evaluation answering: "Is Claude Fable 5.1 the right model for your workload & budget?"
Listed output price $50.00/1M tokens. Quality is the category rating above, not this price.
Dated snapshot metrics aggregated from official evaluators and API providers with visual relative score bars.
Our category lists (Coding, Reasoning, Chat, Agents). Named suites such as SWE-bench Pro are recipe inputs, not this rating.
LMArena, SWE-bench Pro, LiveBench and similar suites appear with a source and as-of date. They are not CompareLLM's rating.
How much of the recipe is present. Missing inputs stay missing; they are never filled with 0 or 50.
Compact streaming class. Exact tok/s stays in the fact sheet, not in compact cells.
Compact class for time-to-first-token. Exact milliseconds stay in the fact sheet.
Commercial API price per 1 Million prompt/output tokens (~750k words).
| Benchmark Metric & Meaning | Reported Score & Capability Fill | CompareLLM Review |
|---|---|---|
Context window (tokens) Context Window Capacity (tokens) | 1M tokens | Source not recorded · Sep 11, 2026 |
Output token price Cost to Generate (Output tokens) | $50/1M tok | Source not recorded · Sep 11, 2026 |
Cached input token price Cost to Reuse Cached Input | $0.25/1M tok | Source not recorded · Sep 11, 2026 |
Input token price Cost to Prompt (Input tokens) | $10/1M tok | Source not recorded · Sep 11, 2026 |
Maximum output (tokens) Maximum Response Length | 128k tokens | Source not recorded · Sep 11, 2026 |
Generation speed (tok/s) Writing Speed (tok/s) | <50 tok/s | Source not recorded · Sep 13, 2026 |
Time to first token Waiting Time (before it replies) | 2s+ | Source not recorded · Sep 13, 2026 |
Peers in the same closed API mid performance tier — not a jump to an unrelated frontier SKU.
Select any rival to launch a side-by-side empirical benchmark comparison with winner deltas.
Compare Claude Fable 5.1 against
Models may use different benchmarks and test settings. This is an indicative composite, not a controlled head-to-head comparison or community Elo. Admin-approved sentiment estimates fill categories without accepted benchmark results. Estimates are labelled and do not increase benchmark coverage. Coverage refers to the configured recipe, not confidence.
Release recent-models-2026-09-23-r1 · recipe reported-text-2026-09-11-r1 · method reported-with-estimates-v2 · research through 2026-09-23
Overall benchmark coverage includes missing applicable categories. Estimated categories contribute to the rating but add no benchmark coverage.
DeepSWE leaderboard (third-party report) · reported 2026-09-15
Config: DeepSWE v1.1 public leaderboard, reported as a 67.4% average over five trials. Recorded because the model had only one coding benchmark and so could not be rated on the category at all. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Repository-level work receives the largest share for coding relevance.
Anthropic (vendor-reported) · reported 2026-09-01
Config: Vendor launch disclosure; Terminal-Bench 4.0; reported model configuration. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Complement repository work with terminal-based tasks.
Direct vendor-reported launch result.
Missing benchmark families: frontiercode.
Editorial estimate, anchored to the Artificial Analysis Intelligence Index v4.3, on which Claude Fable 5.1 scores 53, and placed mid-range, because this category has too few measured results to define a band. AA's index is a composite of ten evaluations and is not this category's recipe, so this is an ordering anchor rather than a measurement. Any accepted result that clears the evidence thresholds replaces it.
As of 2026-09-23 · review on 2026-12-23 · Editorial estimate (not a benchmark result)
Sentiment source 1 →Anthropic (vendor-reported) · reported 2026-09-01
Config: Vendor launch disclosure; Humanity’s Last Exam with tools; reported model configuration. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Tool-assisted academic problems broaden the evidence, with a smaller weight because tools affect outcomes.
Tool-assisted result; not comparable to no-tools HLE.
Missing benchmark families: graduate-science, advanced-mathematics, interactive-abstraction.
Editorial estimate, anchored to the Artificial Analysis Intelligence Index v4.3, on which Claude Fable 5.1 scores 53, and placed mid-range, because this category has too few measured results to define a band. AA's index is a composite of ten evaluations and is not this category's recipe, so this is an ordering anchor rather than a measurement. Any accepted result that clears the evidence thresholds replaces it.
As of 2026-09-23 · review on 2026-12-23 · Editorial estimate (not a benchmark result)
Sentiment source 1 →Missing benchmark families: constrained-story-writing, instruction-compliance.
Anthropic (vendor-reported) · reported 2026-09-01
Config: Vendor launch disclosure; OSWorld 2.0 offline partial-credit score. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Desktop interaction is a complementary core agent workload; offline/partial-credit scope is disclosed.
Partial-credit score; strict score is a different measurement.
Anthropic (vendor-reported) · reported 2026-09-01
Config: Vendor launch disclosure; AutomationBench; reported model configuration. · Effort: reported_best
Admin-authored relevance policy, not empirical difficulty calibration. Include multistep workflow automation.
Direct vendor-reported launch result.
Missing benchmark families: professional-computer-tasks, web-research.
Editorial estimate, anchored to the Artificial Analysis Intelligence Index v4.3, on which Claude Fable 5.1 scores 53, and placed mid-range, because this category has too few measured results to define a band. AA's index is a composite of ten evaluations and is not this category's recipe, so this is an ordering anchor rather than a measurement. Any accepted result that clears the evidence thresholds replaces it.
As of 2026-09-23 · review on 2026-12-23 · Editorial estimate (not a benchmark result)
Sentiment source 1 →Missing benchmark families: conversation-quality, support-task-completion, instruction-compliance.
Editorial estimate, anchored to the Artificial Analysis Intelligence Index v4.3, on which Claude Fable 5.1 scores 53, and placed mid-range, because this category has too few measured results to define a band. AA's index is a composite of ten evaluations and is not this category's recipe, so this is an ordering anchor rather than a measurement. Any accepted result that clears the evidence thresholds replaces it.
As of 2026-09-23 · review on 2026-12-23 · Editorial estimate (not a benchmark result)
Sentiment source 1 →Missing benchmark families: visual-grounding, spatial-reconstruction, visual-symbol-recognition.
What this model is available as, what it takes to run, and where its identity comes from.
Available only through the provider's hosted API.
claude-fable-5-1anthropic/claude-fable-5.1Concise empirical overview formatted for citations and prompt context

Anthropic's Claude Opus 5.5 scores 58 on the Artificial Analysis Intelligence Index, five points clear of anything else in our catalog, and lists at $4 input and $20 output per million tokens. That is 20% below Opus 5 and 60% below Fable 5.1 and GPT-6 Astra.

Seventeen models from thirteen vendors reached the market in September, and the top of the Intelligence Index moved from 51 to 53. No vendor publishes a roadmap, so here is what the release record actually supports.
Plain-English methodology and leaderboard answers
Elo is a crowd vote on which hidden answer people liked more — not a school test. What is preference Elo? · Methodology
Follow this model in your watchlist, set it as your global comparison baseline, or assign it to your custom production stack.
Track updates & rank changes
Compare all models against this
Assign to custom architecture
Compare side-by-side vs all
Plotted against all active catalog models (50th percentile = catalog median).