The cheaper model just complicated the buying decision. Claude Sonnet 5.5 beats Opus 5.5 on Anthropic’s Terminal-Bench evaluation, nearly matches it on professional knowledge work, and charges half as much for standard input and output tokens. Yet running Sonnet at maximum effort can cost more per benchmark task than running Opus.
That’s the central question behind this release: which model delivers acceptable work with the least wasted time and money?
For bounded tasks, Sonnet deserves a serious trial. For difficult reviews, long-context reconstruction and ambiguous projects, Opus still has advantages. Neither the cheapest token nor the highest isolated score settles the choice. The useful comparison includes effort, tools, retries and the quality of the finished result.
Table of Contents
1. Claude Sonnet 5.5 Benchmarks: The Complete Comparison
The table covers the capability inventory in Anthropic’s system card, Section 8, including results omitted from the launch chart. Figures are percentages unless labeled otherwise. Paired scores follow the order stated in the benchmark name.
These are headline configurations, generally Max effort, rather than every point on every effort curve. “n/r” means no comparison value reproduced here. Approximate values reflect rounded chart labels. External evaluations reported inside the card remain distinct from Anthropic’s own runs.
Claude Sonnet 5.5 Benchmarks vs Sonnet 5 and Opus 5.5
| Benchmark & Configuration | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| Coding & Terminal Tasks | |||
| SWE-Bench Pro | 81.3 | 63.2 | 89.9 |
| SWE-Bench Multilingual | 90.3 | 78.3 | 93.9 |
| SWE-Bench Multimodal | 54.3 | 28.1 | 61.4 |
| DeepSWE v1.1 | 71.0 | n/r | n/r |
| FrontierCode v1.1 Main, Max | 46.2 | 42.4 | 54.4 |
| FrontierCode Main / Extended, Xhigh | 52.1 / 64.4 | n/r | n/r |
| FrontierCode Extended, Max | 59.1 | n/r | n/r |
| Terminal-Bench 4.0 | 70.6 | 10.3 | 66.4, Xhigh |
| Terminal-Bench-Science 0.1 | 59.9 | n/r | 58.7 |
| FrontierSWE v2, Proximal harness | 61.9 | n/r | 62.3 |
| CursorBench 4.0, Cursor harness | 55.5 | 34.1 | 57.8 |
| ProgramBench, hidden-test pass rate | 79.7 | 77.3 | 91.2 |
| Multi-Agent ProgramBench | 1h async team beats 4h solo | n/r | n/r |
| Reasoning & Research | |||
| ArXivMath, Aug 2026, no tools / tools | 86.8 / 95.2 | n/r | 91.2 / 96.9 |
| Humanity’s Last Exam, no tools / tools | 56.9 / 64.5 | 43.1 / 54.9 | 64.4 / 67.7 |
| DRACO, modified evaluation | 87.0 | 80.8 | 87.4 |
| WANDR, modified evaluation, soft F1 | 70.0 | 48.3 | 72.3 |
| Vision & Computer Use | |||
| Chartography, no tools | 61.6 | 15.6 | 64.4 |
| Chartography, with tools | 90.2 | n/r | 89.0 |
| BenchCAD Vision2Code, no tools / tools, IoU | 0.747 / 0.963 | n/r | 0.730 / 0.962 |
| OSWorld 2.1, partial / strict pass | 80.1 / 43.5 | 57.0 / 25.6 | 81.8 / 48.7 |
| Professional Work & Automation | |||
| OfficeQA / OfficeQA Pro | 76.9 / 65.6 | 75.1 / 62.1 | 78.9 / 67.7 |
| Legal Agent Benchmark, all-pass / criterion-pass, High | 11.7 / 92.1 | n/r | n/r |
| Legal Agent Benchmark, all-pass / criterion-pass, Max | 10.0 / 93.1 | n/r | n/r |
| GDPval-AA v2.1, Elo | 1844 | 1449 | 1846 |
| AA-Briefcase v1.1, Elo | 1811 | 1359 | 1822 |
| Toolathlon-Verified, Pass@1 / Pass@3 / all-three-pass | 77.8 / 85.2 / 68.5 | 74.7 / 84.3 / 65.7 | 77.8 / 82.4 / 72.2 |
| AutomationBench, Zapier | 44.7 | 10.7 | 42.5* |
| Healthcare & Languages | |||
| HealthBench, raw / length-adjusted | 69.4 / 65.4 | n/r | 68.1 / approximately 61 |
| HealthBench Professional, raw / length-adjusted | 77.1 / 69.2 | n/r / 57.8 | 77.1 / 65.6 |
| PhysicianBench, Pass@1 | 63.2 | 37.4 | 68.4 |
| GMMLU, 42-language average | 92.1 | 89.2 | 94.3 |
| MILU, 11-language average | 91.6 | 89.3 | 93.1 |
| Life Sciences — Special Evaluation Conditions | |||
| BioMysteryBench, Human Solvable / Difficult | 89.2 / 44.7 | 87.5 / 32.9 | approximately 91 / 51 |
| LatchBio, SpatialBench / SingleCellBench | 72.5 / 59.1 | 68.9 / 56.5 | 72.0 / 61.2 |
| Morphology-To-Molecule Matching | 25.0 | 6.8 | approximately 34 |
| Medicinal Chemistry | 65.3 | approximately 41 | 63.5 |
| Protein Design, Sequence Generation / Library Ranking | 51.0 / 54.8 | 20.2 / 44.0 | approximately 60 / 56 |
| De Novo Protein Binder Design, 24h | 82.3 | approximately 73 | 82.6 |
| Biomedical Image Analysis, normalized score | 72.2 | approximately 42 | 71.4 |
| Protocols, Troubleshooting / Understanding V2 | 67.3 / 66.6 | 49.9 / 58.9 | approximately 74 / 69 |
Opus’s AutomationBench score includes rerunning refusals with fallbacks enabled, versus 40.0% previously. Life-science testing disabled Sonnet’s biology safeguards. Multi-agent results use special budgets and serving conditions. DRACO and WANDR use modified setups, so their scores should not be mixed with original leaderboards. BenchCAD uses a 1,000-file subset. Elo, partial credit and full-task success are different measures, not interchangeable percentages.
GDPval-AA and AA-Briefcase were independently run by Artificial Analysis on a prerelease deployment with a subsequently fixed structured-output bug. Anthropic expects any resulting score effect to be small. Treat these as dated measurements, not permanent rankings.
2. Claude Sonnet 5.5 Pricing: The Full Token Comparison
The published rates explain the appeal. They also expose a frequently missed detail: cache reads cost the same on both models.
Claude Sonnet 5.5 Pricing vs Opus 5.5
| Billing Category USD per 1M tokens | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Standard input | $2 | $4 |
| Standard output, including billed reasoning | $10 | $20 |
| Five-minute cache write | $2.50 | $5 |
| One-hour cache write | $4 | $8 |
| Cache read | $0.20 | $0.20 |
| Batch input | $1 | $2 |
| Batch output | $5 | $10 |
| Fast mode input | Not offered | $8 |
| Fast mode output | Not offered | $40 |
| US-only inference multiplier | 1.1× | 1.1× |
Anthropic applies standard token rates across the full one-million-token context. Server-side tools can add charges. Partner cloud regional pricing can differ, and Opus Fast mode is a separate first-party API option, incompatible with Batch. Check the official pricing documentation for the deployment you use. platform.claude.com
Consider an uncached request with 100,000 input tokens and 10,000 billable output tokens. Sonnet costs $0.30, versus $0.60 for Opus, excluding tools and other modifiers. If Sonnet instead generates 40,000 output tokens, its bill reaches $0.60. The cheaper rate has survived. The saving has disappeared.
3. What Changed From Sonnet 5?
Released September 28, 2026, Claude Sonnet 5.5 targets everyday coding and knowledge work. The Sonnet 5.5 vs Sonnet 5 comparison shows substantial gains without changing base token prices.
Its specifications include a one-million-token context, text and image input, text output, and a June 2026 knowledge cutoff. Standard maximum output is 128,000 tokens. The separate Batch API beta supports up to 300,000. Claude Platform Docs
Anthropic claims more than 30% faster output generation and up to 30% lower task costs in its testing. Those claims describe tested workloads, not a guaranteed discount on every prompt. A larger context allowance also says little about how reliably the model uses everything inside it. ProgramBench makes that distinction visible.
4. Does It Really Beat Opus at Coding?
Terminal-Bench tests work inside terminal environments, including scientific computing and engineering. Sonnet’s rise from 10.3% to 70.6% is published, not a typo. It does not mean every developer becomes seven times more productive.
The flagship comparison needs care. Sonnet’s reported standard error is ±2.5 percentage points, versus ±2.6 for Opus. Sonnet ran at Max and Opus at Xhigh, with safeguards and fallback behavior included. A 4.2-point lead is interesting, but insufficient to declare a universal winner.
Independent evidence strengthens the narrower finding. Artificial Analysis also places Sonnet ahead on Terminal-Bench, approximately 64% versus 60%, using its own evaluation setup. These numbers should remain separate from Anthropic’s results. Artificial Analysis
GPT-6 Sol provides another useful reference. Its published FrontierCode Main score is 49.3%, above Sonnet at Max but below Sonnet at Xhigh. That reversal demonstrates how selectively choosing an effort setting can change the headline. ArXivMath’s external GPT scores use different tools and grading conditions, so putting them beside Anthropic’s internal scores as a clean ranking would mislead readers.
For Sonnet 5.5 vs Opus 5.5 coding decisions, review quality matters too. CodeRabbit’s testing found Sonnet caught six of thirteen difficult issues, while Opus caught eight to ten. That small sample favors keeping Opus available for difficult reviews. www.coderabbit.ai
5. Beyond Coding: Strong Work, Uneven Reliability
GDPval-AA’s near-tie suggests Sonnet can produce competitive professional deliverables. It does not mean the models are equivalent across every occupation or task.
Look at completion criteria. Sonnet earns 80.1% partial credit on OSWorld but completes every checkpoint on 43.5% of attempts. Its legal evaluation reaches 92.1% average criterion satisfaction at High, yet only 11.7% of tasks pass everything. Missing one required filing detail can matter more than getting nine others right.
Tools also change the result. Chartography rises from 61.6% without tools to 90.2% with them. Research, visual reasoning and document production increasingly depend on the surrounding software.
Healthcare and life-science scores measure evaluation performance, not deployment readiness. Writing warmth and taste require separate human assessment. A spreadsheet benchmark cannot settle whether a conversation feels helpful.
For a writing workflow, compare the same brief across models and judge factual accuracy, structure, tone and editing time. For slide creation, inspect the exported deck rather than its preview alone. These checks capture the work left for the human, which a polished first impression can hide. They also make conflicting early user reports easier to interpret.
6. Why Cheaper Tokens Can Cost More

Artificial Analysis measured approximately 193,000 output tokens per Intelligence Index task at Max effort. This concerns its broader evaluation suite, not ordinary chat messages or specifically the Terminal-Bench win. Artificial Analysis
Its reported Max-effort cost is $7.60 per index task for Sonnet, compared with $5.98 for Opus. The metric combines weighted evaluation costs, including input, cache operations, reasoning and answer tokens. Multiplying 193,000 by the output rate alone will not reproduce that bill. Artificial Analysis
Anthropic’s efficiency claims can still hold for easier work or lower effort. Different workloads demand different amounts of reasoning.
For production, distinguish an attempt from a successful job. A cheap failed attempt may trigger another run, another review and another correction. Track total spending across those attempts, then divide by outputs that meet your acceptance criteria. That is the number your budget actually needs.
7. Choose Effort Before Choosing a Winner

Sonnet 5.5 effort levels range from Low through Medium, High, Xhigh and Max. Medium is the default in Claude apps and Claude Code. The API defaults to High.
For routine extraction or conversation, test Low or Medium. For bounded coding, start at Medium with explicit checks. Increase effort when a concrete failure suggests more reasoning would help.
AA’s configuration results make the tradeoff concrete. Sonnet High scores 47 at $1.08 per index task, while Opus Medium scores 51 at $1.34. Sonnet Xhigh scores 52 at $2.74, while Opus High scores 54 at $1.82. These are benchmark comparisons, not quotes for your workload. Artificial Analysis
Max can even hurt. On FrontierCode, extra review activity sometimes produced timeouts or out-of-scope edits, lowering Sonnet’s score below Xhigh. More checking is useful until it starts rewriting the assignment.
8. Faster Output Is Not Faster Completion
The advertised speed improvement compares Sonnet with Sonnet 5. It describes output generation, not a universal advantage over Opus.
A model that generates tokens quickly can still finish later if it reasons longer, calls more tools or repeats work. For an interactive application, measure time to the first useful answer. For coding agents, measure time to a verified patch.
CodeRabbit reported roughly half the review time of Sonnet 5 in its tests. That is encouraging workload-specific evidence, not proof that every project finishes twice as fast. Compare equivalent tasks, tools and acceptance standards before turning speed claims into staffing or delivery estimates. CodeRabbit
9. Using Claude Sonnet 5.5 in Claude Code and the API
The first-party API identifier is claude-sonnet-5-5. The model is also available through Amazon Bedrock, Google Cloud and Microsoft Foundry. Select the intended model in Claude Code and set effort explicitly when comparing runs.
Subscription allowances are a separate calculation. A Pro or Max subscription does not turn published API prices into a fixed message quota. Task complexity, conversation size and product limits matter, so “half the token price” does not promise twice the usable sessions.
The September 28 release note announces the model but no launch-specific usage reset. Check your account’s actual allowance rather than assuming an announcement refills it. API access and Claude subscriptions also remain distinct purchasing decisions. Claude Help Center
10. Safeguards, Refusals and Fallbacks
Some higher-risk cybersecurity requests can fall back to Sonnet 5. Certain frontier-model-development requests can also trigger fallback, while biology and reasoning-extraction safeguards can block requests without substituting another model.
The system card distinguishes first-party behavior from API configurations, where developers must opt into automatic fallback, and partner platforms may behave differently. The fallback is disclosed rather than silently changing the answer.
Anthropic says routine software development should be unaffected. Early complaints deserve investigation, but reports involving previous models do not establish Sonnet’s false-positive rate. Teams should test representative legitimate work and record whether interruptions come from model refusals, classifiers or application errors. Those are different problems with different remedies.
11. Should You Switch, and What Might Break?
Trial Sonnet on routine fixes, document drafts and clearly specified implementation work. Keep Opus in the comparison for architecture, difficult reviews and long-context tasks. Using Opus to plan and Sonnet to implement is reasonable, but coordination adds cost and needs measurement.
Before migrating, review the breaking API changes:
- Replace disabled thinking with between_tools where appropriate.
- Remove unsupported forced tool choices.
- Preserve thinking blocks according to their model, conversation and account restrictions.
- Update legacy computer-use tools and unsupported advisor selections.
- Check streaming behavior between tool calls. Claude Platform Docs
Then rerun representative tasks. Keep prompts, tools and success criteria consistent, and log retries alongside token totals. An upgrade that changes your application’s behavior deserves more than a model-ID swap.
Include an easy task, a typical task and a failure-prone task from your own work. Run each more than once. Set the acceptance test before reading the outputs, and include manual corrections in the comparison. Otherwise, it is surprisingly easy to reward the model that produces the most convincing explanation of an unfinished job.
12. Buy the Finished Result
Claude Sonnet 5.5 makes the middle of Anthropic’s range more capable. It also shows why model tiers no longer provide a reliable shortcut for cost or quality.
Start with the least expensive configuration that meets your standard. Escalate when evidence justifies it, and retain the stronger model wherever failures are expensive. A model earns its place by delivering usable work repeatedly.
Follow Binary Verse AI for benchmark analysis grounded in sources and practical tradeoffs. Share your own Sonnet-versus-Opus results with the task, effort setting, cost and outcome. Those details turn a launch-day opinion into evidence another builder can use.
1. Is Claude Sonnet 5.5 better than Opus 5.5?
Claude Sonnet 5.5 scores higher on Terminal-Bench 4.0 in Anthropic’s published evaluation, but that does not make it better across all coding or reasoning tasks. Opus 5.5 remains stronger on several other benchmarks, and Anthropic recommends it for complex, open-ended work requiring sustained judgment. Choose according to the workload and effort setting.
2. How much does Claude Sonnet 5.5 cost?
Standard API pricing is $2 per million input tokens and $10 per million output tokens. Cache reads cost $0.20 per million tokens; cache writes cost $2.50 for five minutes or $4 for one hour. Batch input and output rates are $1 and $5 respectively. These are API charges, separate from Claude subscription fees.
3. Why can Sonnet 5.5 cost more per task despite cheaper tokens?
A lower token price does not guarantee a lower bill. Longer reasoning, repeated inputs and tool calls can increase total consumption. At Max effort, Artificial Analysis reported approximately 193,000 output tokens per Intelligence Index task and a $7.60 cost-per-task metric. Those are evaluation results, not a fixed charge or typical token count for every request.
4. Which effort level should I use with Sonnet 5.5?
Medium is the default in Claude apps and Claude Code; the Claude API defaults to High. Start with the setting appropriate to your workflow, then compare quality, cost and completion time before increasing it. Max is not automatically best: Sonnet 5.5 scored lower at Max than Xhigh on FrontierCode.
5. Why does Sonnet 5.5 sometimes refuse a request or fall back to Sonnet 5?
Sonnet 5.5 includes safeguards that can route certain higher-risk cybersecurity requests to Sonnet 5. Other safeguard categories can block requests instead. Behavior depends on the product, provider and API fallback configuration. Anthropic says routine development should remain unaffected, but individual reports of legitimate requests being blocked require case-by-case investigation.
