Claude Sonnet 5.5 Beats Opus 5.5 on Agentic Coding at Half the Tier Price

September 28, 2026 • news
AnthropicClaudeAI AgentsBenchmarks

Anthropic launched Claude Sonnet 5.5 on September 28, 2026, the second model in the Claude 5.5 family. It is positioned as a faster, leaner complement to Claude Opus 5.5 rather than a replacement for it. For developers running high-frequency agentic pipelines, the release has immediate operational significance: Anthropic reports a 30%-plus throughput gain and up to 30% lower per-task cost versus Sonnet 5, with no price increase at the API level. Claude Haiku 5.5, targeting high-volume and cost-sensitive deployments, is due in the coming weeks with no firm date announced.

Performance: Agentic Coding Leads the Gains

Anthropic's own benchmark data shows Sonnet 5.5's largest jump in agentic coding. On Terminal-Bench 4.0, Anthropic reports a score of 70.6%, versus 10.3% for Sonnet 5 and 66.4% for Opus 5.5 — meaning Sonnet 5.5 exceeds its senior sibling on this specific evaluation. On FrontierCode 1.1 (Main), Anthropic reports 46.2% at Max effort and 52.1% at Xhigh effort, compared to 42.4% for Sonnet 5 and 54.4% for Opus 5.5. CursorBench 4.0 shows 55.5% for Sonnet 5.5 versus 34.1% for Sonnet 5 and 57.8% for Opus 5.5.

Knowledge-work benchmarks show similarly compressed gaps with Opus 5.5. On GDPval-AA v2.1, Anthropic reports Sonnet 5.5 at 1844 versus Sonnet 5 at 1449 and Opus 5.5 at 1846. On AA-Briefcase v1.1, the scores are 1811, 1359, and 1822 respectively. Humanity's Last Exam (with tools) sits at 64.5% for Sonnet 5.5 versus 54.9% for Sonnet 5 and 67.7% for Opus 5.5. OSWorld 2.1 computer-use partial scores run 80.1%, 57.0%, and 81.8% for the same three models. As noted in our earlier analysis of how benchmark framing shapes product perception, these vendor-reported figures are best read alongside the effort-level cost curves Anthropic publishes rather than in isolation.

Efficiency gains in practice appear meaningful. Base44 reports, in Anthropic's early-tester disclosures, that Sonnet 5.5 completed 118 real app builds at 3.6 iterations per build on average, where Opus 5 took 7.7. Balyasny Asset Management reports Sonnet 5.5 used approximately 121,000 tokens per finance-task answer where Sonnet 5 used 497,000. Zendesk reports 20% faster ticket processing. Slack reports approximately 14% fewer output tokens on Slackbot evaluations without prompt changes. These are all vendor-cited reports surfaced through Anthropic's own announcement, not independently verified.

Pricing and Competitive Position

Anthropic prices Sonnet 5.5 identically to Sonnet 5: $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads. The cost story rests entirely on token efficiency — Anthropic reports up to 30% lower cost per task through reduced token consumption rather than any rate reduction.

Model Input ($/1M tokens) Output ($/1M tokens) Tier
Claude Sonnet 5.5 (Anthropic) $2.00 $10.00 Mid-tier
Claude Sonnet 5 (Anthropic) $2.00 $10.00 Mid-tier
Claude Opus 5 (Anthropic) $5.00 $25.00 Flagship
GPT-5.6 Sol (OpenAI) $4.00 $20.00 Mid-tier
GPT-5.6 Terra (OpenAI) $2.00 $12.00 Mid-tier
Gemini 3.1 Pro (Google) $2.00 $12.00 Mid-tier
Grok 4.5 (xAI) $2.00 $6.00 Mid-tier
Claude 4.5 Haiku (Anthropic) $1.00 $5.00 Budget

Competitor pricing from AI Mastery's pricing data as of 2026-09-14. Sonnet 5.5 pricing per Anthropic.

At list rate, Sonnet 5.5 matches GPT-5.6 Terra and Gemini 3.1 Pro on input but undercuts GPT-5.6 Sol by half on both input and output. If Anthropic's reported per-task token reductions hold in production, the effective cost advantage widens further against any competitor priced at or above $2/$10.

Safety: First Sonnet with Opus-Class Cyber Guardrails

Anthropic states that Sonnet 5.5's cybersecurity capabilities are comparable to Opus 5, making it the first Sonnet model subject to the same cyber safeguards and fallbacks previously reserved for Opus and Fable tier models. Biology safeguards remain at the same level as Sonnet 5. Anthropic explicitly scopes both restrictions to a narrow set of high-risk requests, stating that routine software development and most life sciences work are unaffected. On its automated behavioral audit, Anthropic reports Sonnet 5.5 improves on or matches Sonnet 5 on most alignment measures — a vendor self-assessment, not an independent audit.

AI Mastery Analysis

The architectural story here is efficiency engineering rather than capability discovery. Sonnet 5.5's Terminal-Bench 4.0 score exceeding Opus 5.5 (70.6% versus 66.4%, both Anthropic-reported) is a striking result that deserves scrutiny: Terminal-Bench tests multi-step CLI task completion, a workload that favors aggressive tool-call batching and tight loop control over sustained open-ended reasoning. Early-tester reports specifically call out Sonnet 5.5 batching tool calls more aggressively than Sonnet 5, reducing step count and cost — this is the mechanism driving many of the efficiency numbers. As we have argued in the context of constrained inference patterns, optimizing token paths through structured workflows frequently yields larger practical gains than raw benchmark headroom.

The more consequential decision for engineering teams is effort-level selection. Anthropic's own cost-versus-accuracy data shows Sonnet 5.5 at Low or Medium effort beating Sonnet 5's best score at roughly one-tenth the cost per task across multiple benchmarks. Teams defaulting to High effort across all workloads will pay a premium for marginal gains over what Medium effort already delivers. Teams evaluating the portability implications of migrating between model tiers should note that prompt tuning for effort levels is not transferable across providers, creating switching costs that list-price comparisons obscure.

With Haiku 5.5 still pending, developers building cost-sensitive high-volume pipelines are waiting on the model that typically anchors those architectures. The competitive pressure is substantial: AI Mastery's pricing data puts DeepSeek V4 Flash at $0.14/$0.28 and Gemini 3.1 Flash-Lite at $0.25/$1.50. Sonnet 5.5 is a meaningful efficiency upgrade at the mid tier, but the broader 5.5 family's value proposition for volume workloads remains incomplete until Haiku ships.

Sources

Frequently asked questions

How much does Claude Sonnet 5.5 cost per million tokens?

Anthropic prices Sonnet 5.5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads — identical to Sonnet 5. The cost reduction comes from token efficiency: Anthropic reports up to 30% lower cost per task through reduced token consumption.

How does Claude Sonnet 5.5 compare to Opus 5.5 on benchmarks?

On Terminal-Bench 4.0, Anthropic reports Sonnet 5.5 scores 70.6% versus Opus 5.5's 66.4%, making it the stronger performer on that agentic coding evaluation. Opus 5.5 leads on most other benchmarks, including FrontierCode 1.1 (54.4% vs 46.2% at Max effort) and GDPval-AA v2.1 (1846 vs 1844).

How many fewer tokens does Sonnet 5.5 use compared to Sonnet 5 in practice?

In Balyasny Asset Management's testing, Anthropic reports Sonnet 5.5 used approximately 121,000 tokens per finance-task answer versus 497,000 for Sonnet 5. Slack reports approximately 14% fewer output tokens on Slackbot evaluations without any prompt changes.

What cybersecurity safeguards does Claude Sonnet 5.5 have?

Anthropic states that Sonnet 5.5's cybersecurity capabilities are comparable to Opus 5, making it the first Sonnet model subject to the same cyber safeguards and fallbacks previously reserved for Opus and Fable tier models. Biology safeguards remain at the same level as Sonnet 5, and both target only a narrow set of high-risk requests.

When is Claude Haiku 5.5 being released?

Anthropic has announced Claude Haiku 5.5 will join the Claude 5.5 family in the coming weeks, targeting high-volume and cost-sensitive applications. No firm release date has been given.

Free interactive tools for the decisions this piece raises.

Related Reading