Model and compare LLM API spend across providers for any workload.
Free LLM API cost calculator. Enter your workload — input/output/cache tokens, calls per day, and time horizon — and compare monthly spend across GPT-5, Claude, Gemini, Llama, DeepSeek and more. Accounts for Anthropic prompt caching and Gemini context cache tiers. Export CSV/JSON or share a scenario URL. 100% client-side; nothing leaves your browser.
Keywords: llm pricing calculator, api cost calculator, llm cost comparison, token cost calculator
Tags: ai, llm, cost, pricing, api, budget, calculator
AI Cost Calculator is also known as: LLM Pricing Calculator, API Cost Calculator, Token Cost Comparison, LLM Cost Estimator, AI API Budget Tool, Model Cost Comparison, Prompt Cache Calculator.
Select the models you want to compare in the model picker — search the catalog, filter by provider or tier (frontier, fast, standard), and add up to 12 models, or use the one-click Compare Frontier and Compare Fast/Cheap presets.
Describe your workload: set the average input and output tokens per request using the sliders or number fields.
If you use prompt caching, set cache read and cache write tokens per call — these inputs unlock automatically when a selected model supports cache tiers like Anthropic prompt caching or Gemini context caching.
Set your request volume with calls per day and the time horizon in days to project monthly spend, and switch between Standard and Batch pricing to see how provider batch discounts change the ranking.
Review the results table, sorted by monthly cost with the cheapest model highlighted, and switch between Per-call detail, Monthly spend, and Per-1M-token price views.
Copy the results as CSV, download them as JSON, or copy a shareable scenario URL so teammates can open your exact comparison.
Compare 121+ models in one view: per-token pricing for GPT-5, Claude, Gemini, Llama, DeepSeek, Qwen, Grok, Mistral, Cohere, GLM, and MiniMax — a full LLM cost comparison without opening a dozen provider pricing pages.
True monthly AI cost projection: model your exact workload with input tokens, output tokens, calls per day, and a configurable time horizon instead of guessing from list prices.
Prompt cache cost modeling: price cache read and cache write tiers for Anthropic prompt caching and Gemini context caching, and see your net savings versus running the same workload uncached.
Batch API mode: toggle standard vs batch pricing to apply published 50% batch discounts (OpenAI, Anthropic, Google) — providers without a published batch rate are clearly flagged as estimates.
Cheapest model for your workload: results are ranked by projected monthly spend with the winner badged, so the token cost calculator accounts for your actual input/output mix rather than headline rates.
Live pricing with a visible as-of date: rates are refreshed from a provider-reflective dataset (LiteLLM) with a static fallback catalog, and every row links to the official provider pricing page to verify before you budget.
Flexible model discovery: search, filter by provider and tier, and sort by input price, output price, context window, or newest additions to shortlist candidates fast.
Export and share: copy a CSV for spreadsheets, download JSON for your own tooling, or share a URL that encodes your entire scenario — no account required.
100% client-side and private: all math runs in your browser and your scenario never leaves it, so confidential workload estimates stay internal.