Devlery
Blog/Anthropic

Claude Sonnet 5.5 Costs Half of Opus 5.5, and More per Task at max Effort

Anthropic released Claude Sonnet 5.5 on September 28. Pricing holds at Sonnet 5 levels, $2 per million input tokens, and Terminal-Bench 4.0 climbs from 10.3% to 70.6%, past Opus 5.5 at 66.4%.

Claude Sonnet 5.5 Costs Half of Opus 5.5, and More per Task at max Effort
AI 요약
  • Terminal-Bench 4.0 hits 70.6%, above Opus 5.5 at 66.4%.
  • Pricing is unchanged from Sonnet 5: $2 input and $10 output per million tokens.
  • Run it at max effort and the per-task bill comes out higher than Opus 5.5.

Anthropic released Claude Sonnet 5.5 (claude-sonnet-5-5) on September 28, 2026. The price list did not move a cent from Sonnet 5: $2 per million input tokens, $10 for output.

The same day, Claude Code v2.1.284 switched its default model to Sonnet 5.5. That came six days after v2.1.280 promoted the default to Opus 5.5 on September 22. Anyone who opened Claude Code this week watched the session default change twice.

From 10.3% to 70.6% on the test Sonnet 5 failed

Agentic coding moved the most. Terminal-Bench 4.0 measures whether a model drives a long terminal task to completion on its own. Sonnet 5.5 scored 70.6% there. Sonnet 5 scored 10.3% on the same test, and Opus 5.5, at twice the price, scored 66.4%. The model that costs half as much scored 4.2 points higher than the tier above it.

Every number in the announcement is measured at max effort. effort is the dial that sets how long the model deliberates before answering, with five steps: low, medium, high, xhigh, max.

Benchmark (max effort)Sonnet 5.5Sonnet 5Opus 5.5
Terminal-Bench 4.0 (agentic coding)70.6%10.3%66.4%
SWE-Bench Pro (bug fixing)81.3%63.2%89.9%
CursorBench 4.0 (long-running work in an editor)55.5%34.1%57.8%
OSWorld 2.1 (computer use)80.1%57.0%81.8%
AutomationBench (work automation)44.7%10.7%42.5%
GDPval-AA v2.1 (knowledge work, Elo)184414491846
Humanity's Last Exam (tool use)64.5%54.9%67.7%
HealthBench Professional69.2%57.8%65.6%

Sonnet 5.5 leads Opus 5.5 on three of the eight: Terminal-Bench 4.0, AutomationBench, and HealthBench Professional. Opus 5.5 still leads on the other five. The widest gap is SWE-Bench Pro, which patches bugs in real repositories, where 81.3% against 89.9% leaves 8.6 points on the table. GDPval-AA, the knowledge-work evaluation, is the closest at 1844 against 1846.

Speed went up too. Anthropic says output generation is more than 30% faster than Sonnet 5, and published customer measurements alongside it. Box reported 2.4x faster with 12% fewer tokens. Lovable, a coding tool, reported a third fewer tool calls and shell executions down by roughly half. Base44 measured average iterations across 118 builds falling from 7.7 on Opus 5 to 3.6.

Raise effort and the cost ordering flips

On list price, Sonnet 5.5 is exactly half of Opus 5.5. But the cost of finishing one task inverts at max effort.

The numbers come from benchmarking firm Artificial Analysis. At max effort, Sonnet 5.5 spent 193K output tokens per task and billed $7.60. Under the same conditions Opus 5.5 spent 119K and billed $5.98. The per-token rate is half, but it burns 1.6x the tokens, so the final invoice is larger.

Sonnet 5.5 (max effort)
$7.60
cost per task

193K output tokens · $10 per million tokens

Opus 5.5 (max effort)
$5.98
cost per task

119K output tokens · $20 per million tokens

The effort dial on Anthropic's Claude Sonnet 5.5 announcement page

Anthropic's claim of "up to 30% lower cost per task" is not wrong. It holds in the low and medium bands. That is likely why Claude Code and the Claude apps default effort to medium while Claude Platform (the API) defaults to high. If you run Claude Code on defaults, you are in the savings band. If you habitually crank effort to max, you now pay more for the same task.

Anthropic itself concedes there is no settled method for picking an effort level. Independent verification is thin: the 30% speed gain and the 30% cost reduction are both Anthropic's own measurements, and the per-task cost above is the only figure an outside party has reproduced.

Can you use it today

Yes, essentially anywhere. South Korea, Singapore, and the rest of APAC all sit on Anthropic's supported-countries list for both the commercial API and Claude.ai. There is no waitlist and no separate application.

ItemDetail
Who

Any developer with an API key, paid Claude plans, Claude Code, and GitHub Copilot Pro, Pro+, Max, Business, and Enterprise

Price

$2 per million input tokens / $10 output / $0.20 cache reads / $2.50 cache writes. Batch API takes 50% off. Identical to Sonnet 5

Regional availabilityBroadly available, APAC and Singapore included on both region lists
Requirements

None. Only GitHub Copilot rolls out in stages, and on Business and Enterprise an admin can block it in model policy (allowed by default)

Where it runs

Claude API, Amazon Bedrock (anthropic.claude-sonnet-5-5), Google Cloud, Microsoft Foundry

The context window is 1M tokens with no surcharge. Max output in a single response is 128K, or 300K in the Batch API beta. The training cutoff is June 2026. On Haiku 5.5, the small model in the same family, Anthropic said only "in the coming weeks."

What breaks before you migrate

Code that calls the API directly will error if you leave it alone. The migration guide lists four changes.

The one most callers will hit is the setting that turns off deliberation. Requests that used to send thinking as disabled no longer work. At high effort and below, switch to between_tools.

{
  "model": "claude-sonnet-5-5",
  "thinking": { "type": "between_tools" },
  "effort": "medium"
}

The other three are narrower. Requests that force a tool call by setting tool_choice to any or tool get a 400 error, the same behavior that disappeared from Fable 5.1 first and now applies to the Sonnet line. The Claude API and Google Cloud no longer accept the older computer-use tool computer_20251124. The Advisor tool rejects Opus 4.8, Opus 4.7, and Sonnet 5 when named as the advisor.

The safety layer changed as well. This is the first Sonnet model to carry Opus 5.5-grade cybersecurity protections: requests judged high-risk are handed off to Sonnet 5 instead. A classifier against distillation attacks was also added. Distillation attacks extract a model's reasoning traces to train another model on them.

If you use Claude Code, start by running /model today to see which model and effort level you are actually on. Then run one of your routine tasks at medium and compare the session cost against what Opus 5.5 charged last week. Cranking effort to max because a benchmark score looks thin is better left until after that comparison, because max is exactly where the per-task cost flips.