Anthropic Ships Haiku 5.5 With 1M Context, $0.10 Input Tokens

October 8, 2026 • news
AnthropicPricingAI Agents

Anthropic has released Claude Haiku 5.5, a text-and-image model with a 1-million-token context window and baseline pricing of $0.10 per million input tokens, according to MarkTechPost. The launch targets high-volume inference workloads: document compaction, classification, summarization, and subagent routing under larger models. Anthropic says the model is its fastest at standard speed, though Opus models remain faster in Fast Mode.

What the model changes

Claude Haiku 5.5 has a June 2026 knowledge cutoff. MarkTechPost reports standard API requests support up to 128,000 output tokens, while batch jobs in beta expand that to 300,000. It is Anthropic's first Haiku-class model with an adjustable effort setting; adaptive thinking is on by default, and the effort parameter defaults to medium.

The API enforces deterministic generation. Passing non-default temperature, top_p, or top_k values returns a 400 error, MarkTechPost notes. The model also uses a new tokenizer aligned with Sonnet 5.5 and Opus 5.5, which Anthropic says counts the same text payload as roughly 30% more tokens than Claude Haiku 4.5.

Anthropic states the model's biology safeguards match those on Sonnet 5, Sonnet 5.5, and Opus 5. Cybersecurity safeguards are stricter than Haiku 4.5 but somewhat less restrictive than those applied to other recent models. Updated Python and TypeScript SDKs add beta support for computer use and browser use.

Two-tier pricing and competitors

Pricing splits at 100,000 prompt tokens. Up to that threshold, inputs cost $0.10 per million tokens and outputs $0.50; cache reads are $0.01 and five-minute cache writes $0.125. Above 100,000 tokens, rates rise fivefold to $0.50 input and $2.50 output per million tokens. Cache reads and writes in that tier become $0.05 and $0.625. Batch processing takes another 50% off.

Haiku 4.5 listed $1.00 input and $5.00 output per million tokens. Anthropic says about 90% of Haiku 4.5 requests fell under the 100,000-token line. Accounting for the tokenizer change, it estimates Haiku 5.5 runs about 75% cheaper on average, even though the base-tier list price is 90% lower than Haiku 4.5.

Model Short-Prompt Pricing (per 1M In/Out) Long-Prompt Pricing (per 1M In/Out) Max Output Tokens
Claude Haiku 5.5 $0.10 / $0.50 $0.50 / $2.50 above 100K 128,000
GPT-6 Luna $0.10 / $0.50 $0.20 / $0.75 above 272K 128,000
Gemini 3.5 Flash-Lite $0.30 / $2.50 Flat rate 65,536

As MarkTechPost points out, GPT-6 Luna lists the same short-context price but does not move to its higher tier until 272,000 input tokens. That makes a 150,000-token prompt cheaper on Luna on list price.

Vendor-reported benchmarks

Anthropic reports Claude Haiku 5.5 reaches 72.4% on the OSWorld 2.1 offline subset, up from 15.7% for Haiku 4.5 and ahead of GPT-6 Luna at 48.9%. On Terminal-Bench 4.0, Anthropic reports 39.2% for Haiku 5.5, versus 16.4% for Luna and 0.0% for Haiku 4.5, while Sonnet 5.5 leads at 70.6%.

On FrontierCode 1.1 Main, Anthropic reports Haiku 5.5 scored 46.4%, compared with 42.4% for Luna and 52.1% for Sonnet 5.5. On Chartography without tools, it scored 46.4%, ahead of Luna's 29.1%. On Artificial Analysis's GDPval-AA v2.1 benchmark, Anthropic reports Haiku 5.5 scored 1620, ahead of Luna's 1437 and Haiku 4.5's 735. For Humanity's Last Exam, the model scored 45.9% without tools and 57.4% with tools.

Anthropic still directs complex agentic coding work to Sonnet 5.5 and Opus 5.5, framing Haiku 5.5 as a narrowly scoped subagent: fetching a 10-K revenue line, classifying a support ticket, or filling a web form while a larger model assembles the deliverable.

AI Mastery analysis

The hard-coded sampling constraints are the most important deployment signal. A 400 error on non-default temperature, top_p, or top_k means Haiku 5.5 is optimized for deterministic extraction and tool calls, not creative prose. Combined with pricing, the result is a model engineered as cheap connective tissue in hierarchical agent systems, not as a general-purpose chat model.

That fits the broader token efficiency repricing dynamic. In the same launch, Anthropic halved Sonnet 5.5 cache reads to $0.10 per million tokens, which the maker estimates cuts most agentic task costs by about 20%. Cheap cache reads on a flagship model plus a hyper-cheap routing model make multi-tier architectures financially practical, even though the new tokenizer reduces the advertised 90% base cost reduction by inflating token counts.

Sources

Frequently asked questions

How much does Claude Haiku 5.5 cost per million tokens?

For prompts up to 100,000 tokens, Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. Cache reads in that tier cost $0.01 per million tokens, and five-minute cache writes cost $0.125. Above 100,000 prompt tokens, input and output prices rise fivefold to $0.50 and $2.50 per million tokens.

Does Claude Haiku 5.5 have adjustable reasoning effort?

Yes. It is Anthropic's first Haiku-class model with an adjustable effort setting, and adaptive thinking is on by default with the effort parameter set to medium. Non-default temperature, top_p, or top_k values return a 400 error, so sampling control, not reasoning effort, is restricted.

What changed between Claude Haiku 4.5 and Claude Haiku 5.5?

Anthropic reports that Haiku 5.5 improves OSWorld 2.1 offline subset performance from 15.7% to 72.4%, and Terminal-Bench 4.0 from 0.0% to 39.2%. It is priced 90% lower than Haiku 4.5 for prompts up to 100,000 tokens. The new tokenizer counts the same text as roughly 30% more tokens, so Anthropic estimates the average cost saving at about 75% after adjustment.

How much context can Claude Haiku 5.5 process?

MarkTechPost reports Claude Haiku 5.5 keeps a 1-million-token context window. Standard API requests can generate up to 128,000 output tokens, while beta batch jobs support up to 300,000 output tokens. Prompts above 100,000 tokens move to the higher pricing tier.

Where is GPT-6 Luna cheaper than Claude Haiku 5.5?

For prompts above 100,000 tokens, Haiku 5.5 moves to $0.50 per million input and $2.50 per million output tokens. GPT-6 Luna keeps its $0.10/$0.50 list price until 272,000 input tokens, then moves to $0.20/$0.75. MarkTechPost notes that a 150,000-token prompt is therefore cheaper on Luna on list price.

Free interactive tools for the decisions this piece raises.

Related Reading