LLM API Cost Calculator — Compare 121+ Models (GPT-5, Claude, Gemini, Llama, DeepSeek, Qwen, xAI, Mistral, Cohere)

AI Cost Calculator

Model and compare LLM API spend across providers for any workload.

Free LLM API cost calculator. Enter your workload — input/output/cache tokens, calls per day, and time horizon — and compare monthly spend across GPT-5, Claude, Gemini, Llama, DeepSeek and more. Accounts for Anthropic prompt caching and Gemini context cache tiers. Export CSV/JSON or share a scenario URL. 100% client-side; nothing leaves your browser.

Keywords: llm pricing calculator, api cost calculator, llm cost comparison, token cost calculator

Tags: ai, llm, cost, pricing, api, budget, calculator

Browse all 35 Developer Tools tools →

AI Cost Calculator is also known as: LLM Pricing Calculator, API Cost Calculator, Token Cost Comparison, LLM Cost Estimator, AI API Budget Tool, Model Cost Comparison, Prompt Cache Calculator.

How to AI Cost Calculator Online

  1. Select the models you want to compare in the model picker — search the catalog, filter by provider or tier (frontier, fast, standard), and add up to 12 models, or use the one-click Compare Frontier and Compare Fast/Cheap presets.

  2. Describe your workload: set the average input and output tokens per request using the sliders or number fields.

  3. If you use prompt caching, set cache read and cache write tokens per call — these inputs unlock automatically when a selected model supports cache tiers like Anthropic prompt caching or Gemini context caching.

  4. Set your request volume with calls per day and the time horizon in days to project monthly spend, and switch between Standard and Batch pricing to see how provider batch discounts change the ranking.

  5. Review the results table, sorted by monthly cost with the cheapest model highlighted, and switch between Per-call detail, Monthly spend, and Per-1M-token price views.

  6. Copy the results as CSV, download them as JSON, or copy a shareable scenario URL so teammates can open your exact comparison.

AI Cost Calculator Features

  • Compare 121+ models in one view: per-token pricing for GPT-5, Claude, Gemini, Llama, DeepSeek, Qwen, Grok, Mistral, Cohere, GLM, and MiniMax — a full LLM cost comparison without opening a dozen provider pricing pages.

  • True monthly AI cost projection: model your exact workload with input tokens, output tokens, calls per day, and a configurable time horizon instead of guessing from list prices.

  • Prompt cache cost modeling: price cache read and cache write tiers for Anthropic prompt caching and Gemini context caching, and see your net savings versus running the same workload uncached.

  • Batch API mode: toggle standard vs batch pricing to apply published 50% batch discounts (OpenAI, Anthropic, Google) — providers without a published batch rate are clearly flagged as estimates.

  • Cheapest model for your workload: results are ranked by projected monthly spend with the winner badged, so the token cost calculator accounts for your actual input/output mix rather than headline rates.

  • Live pricing with a visible as-of date: rates are refreshed from a provider-reflective dataset (LiteLLM) with a static fallback catalog, and every row links to the official provider pricing page to verify before you budget.

  • Flexible model discovery: search, filter by provider and tier, and sort by input price, output price, context window, or newest additions to shortlist candidates fast.

  • Export and share: copy a CSV for spreadsheets, download JSON for your own tooling, or share a URL that encodes your entire scenario — no account required.

  • 100% client-side and private: all math runs in your browser and your scenario never leaves it, so confidential workload estimates stay internal.

Frequently Asked Questions

How does this LLM API cost calculator work?
Enter your workload — input tokens, output tokens, optional cache read/write tokens per call, calls per day, and a time horizon. The calculator multiplies each token amount by the model's published per-million-token rate, sums the tiers into a per-call cost, and projects monthly spend. Because every selected model is scored against the same workload, the results are directly comparable.
How do I estimate my monthly LLM API spend?
Take a representative request and note its average input and output token counts, then estimate how many calls your app makes per day. Enter those numbers, set the horizon to 30 days, and the Monthly column shows projected spend per model. If you use prompt caching, fill in the cache token fields for a more accurate estimate.
Why are output tokens more expensive than input tokens?
Generating text requires the model to run a full forward pass for every token produced, while processing input tokens can be parallelized. Providers therefore charge a higher per-million-token rate for output — often 3–10x the input rate — which is why your output volume usually dominates the bill. The Per-call detail view breaks out both costs so you can see the split.
How does prompt caching affect LLM API costs?
Cached tokens are billed at separate tiers: cache reads are much cheaper than regular input, while some providers (notably Anthropic prompt caching) also charge a premium on cache writes. The calculator prices both tiers per model and shows a Savings column comparing your cached workload against the uncached input rate. Cache inputs are enabled only when a selected model publishes cache tiers.
Which LLM is the cheapest for my workload?
It depends on your token mix — a model with cheap input can lose on output-heavy workloads. The results table is sorted by projected monthly cost for your exact scenario and badges the cheapest option, so you can compare frontier, fast, and standard tiers side by side. Always weigh quality and latency before routing production traffic to the lowest bidder.
Is the pricing data current?
The calculator loads live pricing from a provider-reflective dataset sourced via LiteLLM and displays the date the rates were last updated, falling back to a bundled static catalog if the live feed is unavailable. Every row also links to the official provider pricing page. LLM prices change frequently, so verify against the provider before committing a budget.
Are Batch API discounts included?
Yes. Switch pricing mode from Standard to Batch to apply published batch discounts — typically 50% off list rates for OpenAI Batch, Anthropic Batch, and Google Gemini Batch. For providers without a published batch rate, the calculator applies an estimated 50% and flags those rows as estimates. Negotiated enterprise rates, fine-tuning, and taxes are excluded, so treat the output as a planning number rather than an invoice.
Is the tool free, and is my data private?
Yes — the AI cost calculator is completely free with no account required. All computation runs 100% client-side in your browser and your scenario is never sent to a server. The Share scenario button simply encodes your inputs in the URL, and you can also display results in USD, EUR, GBP, or INR.