Official API pricing and token cost benchmarks for Claude Haiku 4 5 by Anthropic. Currently billed at $1.00 / 1M in and $5.00 / 1M out. Claude Haiku 4 5 is a flagship frontier intelligence model engineered for complex reasoning, multi-step problem solving, architecture design, and high-precision agentic tasks.
≈ $0.0010 per 1K tokens
5.0x higher than input
⚡ Saves 90% on cache hits
~267 standard pages
Standard per-token rate cards can feel abstract. Below is what calling Claude Haiku 4 5 actually costs across standardized enterprise workloads and engineering production loops.
How context length and prefix caching impact your bill at scale, and engineering strategies to optimize token throughput.
$0.200
1 full load of 200,000 tokens at standard input rates.
$0.020
Saves 90% on subsequent queries hitting pre-warmed cache.
5.0x
Output tokens cost 5.0x more than standard input tokens.
Because generation costs 5.0x more than input processing, using strict JSON schemas (Structured Outputs) to suppress conversational pleasantries directly lowers per-query cost by 40%–70%.
Prefix caching yields up to 90% savings. Always place static instructions, schemas, and reference code at the top of your prompt to maximize cache hit rates across agentic loops.
Current pricing specifications for Claude Haiku 4 5 via Anthropic.
| Specification / Metric | Published Rate / Limit |
|---|---|
| Standard Input TokensPrompt, system instructions, and tool schemas | $1.00 / 1M tokens |
| Standard Output TokensGenerated text, reasoning steps, and tool calls | $5.00 / 1M tokens |
| Prompt Cache ReadCache hit on pre-loaded static prompt prefixes | $0.100 / 1M tokens (90% discount) |
| Prompt Cache Write / CreationInitial write to ephemeral prompt cache memory | $1.25 / 1M tokens |
| Context Window SizeMaximum combined input tokens per single request | 200,000 tokens (~267 pages) |
| Max Generation LengthMaximum output tokens generated per completion | 64,000 tokens |
| Batch API PricingAsynchronous workloads completed within 24 hours | 50% off standard token rates |
Compare Claude Haiku 4 5 against its closest industry rivals across standard chat, RAG, and classification workloads:
Comparable frontier-tier models with their current per-token rates and cost differential.
| Model | Provider | Context | Input / 1M | Output / 1M | Input Cost Delta |
|---|---|---|---|---|---|
| Claude Sonnet 5 | Anthropic | 1.0M | $2.00 | $10.00 | +100% higher |
| Gemini 3.1 Pro | 1.0M | $2.00 | $12.00 | +100% higher | |
| Gpt 5.4 | OpenAI | 1.1M | $2.50 | $15.00 | +150% higher |
| Gpt 5.6 | OpenAI | 922K | $4.00 | $20.00 | +300% higher |
| Claude Opus 5 | Anthropic | 1.0M | $5.00 | $25.00 | +400% higher |
Explore higher-capacity or lower-latency alternatives across the Anthropic ecosystem.
Tier Positioning: Claude Haiku 4 5 is a flagship frontier intelligence model engineered for complex reasoning, multi-step problem solving, architecture design, and high-precision agentic tasks. Compared to other frontier-tier models, Claude Haiku 4 5 offers a competitive rate card at $1.00 / 1M input and $5.00 / 1M output tokens.
Technical pricing details, caching rules, and budgeting considerations.
Claude Haiku 4 5 costs $1.00 per 1 million input tokens and $5.00 per 1 million output tokens. If prompt caching is active, cache reads cost $0.100 per 1M tokens.
Yes. Most major providers (including OpenAI, Anthropic, and Google) offer a 50% discount on standard token pricing for asynchronous Batch API jobs completed within a 24-hour turnaround window.
Prompt caching allows the provider to reuse pre-computed attention states from recurring prompt prefixes. Repeated prompt segments cost $0.100/1M tokens instead of $1.00/1M, saving ~90%.
Use our interactive AI Cost Calculator to input your exact daily request volume, average prompt token count, completion length, and prompt cache hit ratio. It generates precise cost forecasts across all models simultaneously.
Simulate real-world token consumption, model mixtures, and prompt cache savings with the DevFlow AI Cost Calculator.