📊 See Claude Haiku 5.5 on the usage ranking

📱 Get the AI Rank app

Claude Haiku 5.5 API pricing starts at $0.10 per million input tokens and $0.50 per million output tokens on the Claude Platform, for prompts up to 100,000 tokens. Above that prompt length, the rates become $0.50 input and $2.50 output. These are USD token rates checked on October 9, 2026, rather than a flat price per message. Anthropic's Haiku product page.

For a short extraction or classification request, estimate input and output separately. For a long agent conversation, first check which prompt-length tier applies. A low headline rate alone cannot tell you the cost of a successful task.

The two standard price tiers

Prompt length per request Input, USD per million tokens Output, USD per million tokens
Up to 100,000 tokens $0.10 $0.50
Over 100,000 tokens $0.50 $2.50

Use the official model identifier claude-haiku-5-5 when checking the Claude Platform configuration. These rates describe that platform; check the provider's own price sheet if your invoice comes from another service. Official availability and pricing.

What the 100,000-token threshold means

Anthropic prices each request independently. Prompt length includes all input tokens, including cache reads and writes. A cache hit does not exempt a long request from the higher tier. Earlier requests keep their original prices. The higher tier applies to the request, rather than only the tokens exceeding the threshold. Long-context pricing rules.

For a plain request without caching or extra charges, the planning formula is:

cost = input_tokens / 1,000,000 × input_rate + output_tokens / 1,000,000 × output_rate

The following are calculated illustrations, not bills or measured usage:

Assumed request Calculation Estimated token cost
10,000 input + 2,000 output 0.01 × $0.10 + 0.002 × $0.50 $0.002
100,000 input + 2,000 output 0.10 × $0.10 + 0.002 × $0.50 $0.011
100,001 input + 2,000 output 0.100001 × $0.50 + 0.002 × $2.50 $0.0550005

The last two rows show why a growing conversation needs monitoring. They assume exact token counts, no cache, no tools and no other modifiers; they do not predict a real application's monthly bill.

Caching and batch requests need separate accounting

Cache reads, cache creation and uncached input should be tracked separately. Anthropic's price table lists cache-read rates of $0.01 per million tokens in the shorter tier and $0.05 in the longer tier; creation has different rates. Eligible Batch API input and output receive a 50% discount. Server-side tools can add charges. Verify the applicable feature rules before combining discounts. Official pricing documentation.

Do not multiply every input token by the cache-read rate. Only tokens actually reported in that category qualify. Likewise, a scheduled workload is not automatically a Batch API request.

A useful budget worksheet

Before changing your model routing, record these fields for representative requests:

  1. Actual model ID and platform that bills you.
  2. Total prompt length and applicable tier.
  3. Uncached input, cache writes, cache reads and output counts.
  4. Additional tool charges and applicable discounts.
  5. Whether the task passed, retry count and human correction time.

Then sum the request costs for the entire task. Divide spend by successful tasks to assess the workflow. For example, 1,000 requests matching the first illustration would cost $2 in token charges under its assumptions; repeated attempts and other charges must be added separately. This is arithmetic for planning, not a claim of tested savings.

What AI Rank can and cannot tell you

In AI Rank's Arena overall snapshot dated October 8, 2026, claude-haiku-5-5 is absent. The older claude-haiku-4-5-20251001 is present. This does not give Haiku 5.5 a zero score, and the older model's rating cannot be transferred to it. Source: AI Rank overall API, using Arena's text_style_control policy; original Arena data, CC BY 4.0.

Use the Haiku 4.5 profile and top-rated model board to find evaluation candidates. Those preference signals do not determine API prices or cost per successful task. For publisher-reported performance and evaluation conditions, see our separate Haiku 5.5 benchmark guide.

→ Track it on AI Rank: Claude Haiku 4.5 — Arena rating, rank and 90-day history

📱 Follow it in the AI Rank app

Start with a small representative workload, retain its token accounting and success criteria, and compare the full task cost before expanding traffic.

Compare before you choose: pricing changes and rankings move daily — check the current model rankings on AI Rank before you commit.

Or 📱 get the AI Rank app to follow them on your phone.

Prepared with AI assistance and checked against cited official sources and AI Rank snapshots. No paid API experiment was conducted for this article.