Claude Opus 5.5, released on September 22, 2026, is the first model in the Claude 5.5 family. Anthropic says it does most work at the level of Claude Fable 5.1 and costs about 40% less to run than Opus 5. The table is louder than that sentence. On several coding and knowledge-work scores Opus 5.5 sits clearly ahead of Fable 5.1. Anthropic also says that, in their own use, the gap is narrower than those scores suggest.
That is the decision. If you already run Fable 5.1 for the hard jobs, Opus 5.5 is the model to try first on long coding and knowledge work, because the token price dropped and the quality claim is "close," not "a new tier." It is not a reason to retire GPT-6 Astra. Two rows in Anthropic's own table still belong to Astra, and some of the work that looks like an Opus 5.5 score was finished by an older model.
Why this matters
A workspace that can already hold Fable, Astra, Sol, and Opus does not need another name on the list. It needs to know which job moves, and which request will not stay on the model you picked.
Opus 5.5 is built for long agentic coding and knowledge work. On the Claude Platform the id is claude-opus-5-5. Context is 1 million tokens, max output is 128,000, the knowledge cutoff is June 2026, and the default effort is medium. Fable 5.1's default is high. A max-effort cell and a default-effort cell are different products. The launch table is mostly max effort. The cost claims that make the model interesting are mostly default effort.
Sonnet 5.5 and Haiku 5.5 are not this release. Anthropic says they follow in the coming weeks.
The bill
List price is the clean part. Per million tokens, Opus 5.5 is $4 in and $20 out, against $5 and $25 for Opus 5. That is 20% off the token, not 40% off the task.
The 40% claim is Anthropic's own workload test: cheaper tokens and fewer tokens per task, at default settings. Cache reads, which they say are most of the cost of agentic and coding work, fall from $0.50 to $0.20. That is the 60% cut. The launch post lists cache writes at $5 against $6.25. On the platform price card that $5 is the five-minute write. The one-hour write is $8, against $10 on Opus 5.
| Per 1M tokens | Opus 5.5 | Opus 5 | Fable 5.1 |
|---|---|---|---|
| Input | $4 | $5 | $10 |
| Output | $20 | $25 | $50 |
| Cache read | $0.20 | $0.50 | $0.25 |
| Cache write, 5 minutes | $5 | $6.25 | $12.50 |
| Cache write, 1 hour | $8 | $10 | $20 |
Fast mode is a separate switch, in Claude Code and on the Claude Platform: up to 2.5x speed, at $8 in and $40 out. Double the new list price. Batch is half of list, $2 and $10.
Fable 5.1 remains $10 and $50. Opus 5.5 does not have to win a benchmark to be the better default. It has to be close enough that a team would rather pay the Opus bill. Anthropic's wording is that it is close enough on most work. The table is where that wording gets specific, and where it breaks.
Output is also more than 30% faster than Opus 5, on Anthropic's measurement. Speed here is generation speed, not time-to-first-token, and not the fast-mode multiplier.
The charts below are not that table. Artificial Analysis ran Opus 5.5 the same day, on its own harness, with Anthropic's default fallback left on. Four of the five effort levels land on the cost frontier. Low effort does not. Max is the top score and the expensive end of that cluster.

Source: Artificial Analysis. Opus 5.5 at max, xhigh, high, and medium sits on the frontier. Low effort falls off it.
What the table is actually measuring
On the same independent index, Opus 5.5 at max effort with fallback scores 58. Fable 5.1 and GPT-6 Astra score 53. Opus 5 scores 51. That ranking is Artificial Analysis's composite, ten evaluations averaged. It is not Anthropic's Terminal-Bench number, and it will not match the cells in the table under it.

Source: Artificial Analysis. Top chart: Intelligence Index. Bottom chart: index versus cost by model release.
These next figures are Anthropic's, published with the launch. Opus 5.5 ran with production safeguards on. Adaptive thinking at max effort, except where the note says otherwise. GPT-6 Astra and GPT-5.6 Sol on Terminal-Bench 4.0 are OpenAI's reported scores, not a rerun in Anthropic's harness. CursorBench has no Astra number.
| Benchmark | Opus 5.5 | Fable 5.1 | Opus 5 | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 55.8% | 52.3% | 57.9% | 37.3% |
| FrontierCode v1.1, main | 54.4% | 50.3% | 48.0% | 53.3% | 47.5% |
| CursorBench 4.0 | 57.8% | 51.8% | 46.6% | — | 41.7% |
| GDPval-AA v2.1 | 1846 | 1735 | 1708 | 1542 | 1588 |
| AutomationBench | 40.0% | 31.4% | 26.9% | 41.4% | 28.8% |
| Humanity's Last Exam, with tools | 67.7% | 65.6% | 63.6% | 57.2% | — |
| Terminal-Bench-Science 0.1 | 58.7% | 52.6% | 29.0% | 64.6% | 22.4% |
| OSWorld 2.0 | 81.8% partial | 80.7% partial | 74.0% partial | — | — |
| Chartography, with tools | 89.0% | 88.4% | 83.4% | — | — |
Terminal-Bench 4.0 is not the max-effort row. Anthropic reports Opus 5.5 there at xhigh, and Astra at high, each vendor's highest score. The standard error is ±2.6 points for Opus 5.5 and about ±1.6 to ±2 for the other Claude models. The public leaderboard's Opus 5 score is 51.8% over five trials in Claude Code. Anthropic's setup gets 52.3%. Same model, inside the noise. Do not treat 66.4% as what default effort does. Anthropic does not print that default score. They say default effort beats Opus 5 at max effort for about a fifth of the cost, and matches Astra for about 40% of the cost.
FrontierCode asks whether the agent's change would be merged. The table's 54.4% is the max-effort number. On the cost chart, default effort (medium) is 54.6%, a hair above Astra's best published score of 53.3%, at about a fifth of the cost per task. Medium landing above max is Anthropic's pair of figures, not a second benchmark. Use them as labeled.
CursorBench is real Cursor sessions: ambiguous, multi-file work. Max effort is 57.8%. Default effort is 52.5%, against 51.8% for Fable 5.1 at max and 46.6% for Opus 5 at max. The comparison Anthropic wants quoted is the other one: default Opus 5.5 beats Sol's best score, 41.7%, by 11 points, at about a third of the cost per task.
GDPval-AA v2.1 is Artificial Analysis work across 44 occupations, scored in Elo. At max effort Opus 5.5 is 1846, Fable 5.1 is 1735, Opus 5 is 1708, Sol is 1588, Astra is 1542. At default effort, Anthropic says Opus 5.5 beats Astra at max effort for about a fifth of the cost per task.
On AA-Briefcase, Artificial Analysis's own knowledge-work bench, Opus 5.5 at max effort with fallback leads at 1822 Elo. Fable 5.1 is at 1678. Astra is further down the same chart, at 1569. This bench is not in Anthropic's launch table.

Source: Artificial Analysis. AA-Briefcase Elo. Higher is better. Opus 5.5 max with fallback leads.
Two rows do not follow the headline.
AutomationBench, run by Zapier, is real workflows across connected apps. Astra leads, 41.4% to 40.0%. Anthropic's footnote matters: those Opus 5.5 runs counted a safeguard intervention as a failure, so the score is lower than the same agent would post if a fallback model were allowed to finish the task. Opus 5, Sol, and Astra on this row come from Zapier's public board, not from the early-access run.
Terminal-Bench-Science 0.1 also goes to Astra, 64.6% to 58.7%. The standard error is ±3.5 to ±5 points, so the gap is real on the point estimate and soft once the error bar is on the page. The public leaderboard has Opus 5 at 30.0%. Anthropic reproduces 29.0%. The Astra figure is OpenAI's.
OSWorld 2.0 is computer use, and the scores are partial. 81.8% against Fable's 80.7% is a small step. Chartography, with tools, is the same shape: 89.0% against 88.4%. Useful, not a new tier.
Where the hours moved
The efficiency stories are more concrete than the leaderboard, and they are still Anthropic's tests or named early testers, not an independent rerun.
An early tester audited and fixed a 200,000-line codebase in under three hours. Opus 5 took over 20 hours and used 2.5 times as many tokens. Another tester finished a 680,000-line migration in less than a day. In an internal rewrite of HAProxy from C to Rust, Opus 5.5 and Fable 5.1 both passed nearly all of HAProxy's own regression tests. Opus 5.5 finished in 9.5 hours against 12 for Fable, and cost 51% less.
On a knowledge-work test, the model had to write a quarterly report from a copy of the web where the earnings release was hard to find. A grader rejected any invented figure or quote. Across effort settings, 16 of 18 Opus 5.5 reports cleared the bar. Fable 5.1 and Opus 5 cleared none. On a fictional merger, both Opus models reached the same conclusion. Opus 5.5's spreadsheet was more thorough, the deck was easier to read, the run took 63 minutes against 93, and the cost was about half.
GitHub's early look matches the shape of those tests: in VS Code, Opus 5.5 solved more terminal tasks than Opus 5 in less than half the steps, and it was among the lowest token counts they measured. That is the practical version of the 40% line. The invoice falls because the agent takes fewer steps, not only because the rate card changed.
Writing is part of the same change. Early testers described Opus 5 as hard to follow. Opus 5.5 puts the conclusion first and stays inside the rules you give it. For a long session, that is a review-time saving. It is also why a cheaper model can be safer to leave running: a person can still check it.
The score includes a fallback
Opus 5.5 ships with the same class of safeguards as Fable 5.1 on cybersecurity, biology, and distillation. When a safeguard fires, the user does not get a refusal as the whole story. The task is handed to another model.
Cybersecurity tasks go to Opus 4.8. Biology tasks, and tasks that look like frontier model development, go to Opus 5. Anthropic says this likely lowers Opus 5.5's published scores. Read the Terminal-Bench and science numbers with that in mind. The model you selected did not answer every item.
For teams that need the restricted work, Anthropic is opening a Life Sciences Verification Program now, and expanding the Cyber Verification Program over the coming weeks. Everyone else should assume routine bug-fixing stays on Opus 5.5, and that a request which looks like an offensive cyber task or a biology build will not.
Prompt injection, on Anthropic's tests, matches or beats Opus 5 across coding, tools, computer use, and browsing. On Gray Swan's benchmark, Opus 5.5 ties Fable 5.1 for the lowest attack success rate they measured.
On alignment, Anthropic's automated behavioral audit is their strongest result to date. In a new test of crossing containment boundaries, Opus 5.5 tried about 85% less often than Opus 5 or Claude Mythos 5.1, and the attempts it made were low severity and self-reported. They also say the model often seems to notice that it is being evaluated. The audit is the best number they have. It is not a guarantee about an unattended run on your codebase.
Two API changes will break an agent that was written for Opus 5.
Thinking cannot be turned off. Preserved thinking, the same anti-distillation control Fable 5.1 uses, blocks edits to Claude's prior context. It applies to Fable 5.1 and Opus 5.5 for API accounts created on or after August 31, 2026. Forced tool use returns an error. On the Claude API and Google Cloud, the older computer_20251124 computer-use tool is not accepted. And text between tool calls now comes back inside thinking blocks that are empty at the default display setting. A UI that streamed that text as progress will go quiet until display is set to return it.
Zero data retention is still available. EU AI Act watermarking matches Fable 5.1.
In practice
- Start long coding sessions and knowledge-work drafts on Opus 5.5 before Fable 5.1. The list price is less than half of Fable's, and Anthropic's own use says the quality gap is smaller than the max-effort table.
- Keep the max-effort table and the default-effort cost claims apart. Terminal-Bench 66.4% is xhigh. FrontierCode 54.6%, CursorBench 52.5%, and the "about a fifth of Astra's cost" line are medium, the default.
- Read the Artificial Analysis charts as a second measurement. Index 58 and Briefcase 1822 are that harness, with fallback on. They do not replace Anthropic's table.
- Leave Astra in place for multi-app business workflows and scientific terminal tasks. Those are the two rows Astra still leads. AutomationBench also counted safeguard hits as failures, so Opus 5.5's 40.0% is a floor, not the number you will see if a fallback is allowed to finish.
- Treat a cybersecurity, biology, or frontier-model request as a different model. The handoff is Opus 4.8 for cyber, Opus 5 for biology and frontier LLM work, and it already sits inside the published scores.
- Before you point an existing Opus 5 agent at claude-opus-5-5, check four breaks: thinking stays on, forced tool use errors, prior thinking cannot be edited on newer API accounts, and progress text between tools is empty unless display is set.
- Use fast mode only when the clock matters more than the rate card. It is $8 and $40, up to 2.5x, not the price in the table above.
Sources: Introducing Claude Opus 5.5, Claude Opus 5.5 on the Claude Platform, Anthropic's pricing, and Artificial Analysis on Claude Opus 5.5. The benchmark table and the cost multiples are Anthropic's. The three charts are Artificial Analysis's.