Official API pricing and token cost benchmarks for Glm 5 by Zhipu AI. Currently billed at $0.950 / 1M in and $3.15 / 1M out. Glm 5 is a balanced, cost-efficient standard model designed for general enterprise automation, structured data extraction, summarization, and daily developer workflows.
≈ $0.000950 per 1K tokens
3.3x higher than input
No prompt caching published
~0 standard pages
Standard per-token rate cards can feel abstract. Below is what calling Glm 5 actually costs across standardized enterprise workloads and engineering production loops.
How context length and prefix caching impact your bill at scale, and engineering strategies to optimize token throughput.
$0.000000
1 full load of 0 tokens at standard input rates.
$0.000000
Prompt caching rate not published for this model.
3.3x
Output tokens cost 3.3x more than standard input tokens.
Because generation costs 3.3x more than input processing, using strict JSON schemas (Structured Outputs) to suppress conversational pleasantries directly lowers per-query cost by 40%–70%.
Keep prompt templates modular so system instructions can be re-used efficiently across parallel calls.
Current pricing specifications for Glm 5 via Baseten.
| Specification / Metric | Published Rate / Limit |
|---|---|
| Standard Input TokensPrompt, system instructions, and tool schemas | $0.950 / 1M tokens |
| Standard Output TokensGenerated text, reasoning steps, and tool calls | $3.15 / 1M tokens |
| Prompt Cache ReadCache hit on pre-loaded static prompt prefixes | Standard input rate |
| Prompt Cache Write / CreationInitial write to ephemeral prompt cache memory | Standard input rate |
| Context Window SizeMaximum combined input tokens per single request | 0 tokens (~0 pages) |
| Max Generation LengthMaximum output tokens generated per completion | 0 tokens |
| Batch API PricingAsynchronous workloads completed within 24 hours | 50% off standard token rates |
Comparable standard-tier models with their current per-token rates and cost differential.
| Model | Provider | Context | Input / 1M | Output / 1M | Input Cost Delta |
|---|---|---|---|---|---|
| Llama 4 Scout | Meta | 131K | $0.100 | $0.300 | -89% cheaper |
| Deepseek V3.2 | DeepSeek | 164K | $0.269 | $0.400 | -72% cheaper |
| Grok 4.3 | xAI | 1.0M | $1.25 | $2.50 | +32% higher |
| Qwen3.8 Max | Qwen | 992K | $2.00 | $6.00 | +111% higher |
Tier Positioning: Glm 5 is a balanced, cost-efficient standard model designed for general enterprise automation, structured data extraction, summarization, and daily developer workflows. Compared to other standard-tier models, Glm 5 offers a competitive rate card at $0.950 / 1M input and $3.15 / 1M output tokens.
Calculate accurate costs for Glm 5 with your custom tokens and requests in the DevFlow AI Cost Calculator
Technical pricing details, caching rules, and budgeting considerations.
Glm 5 costs $0.950 per 1 million input tokens and $3.15 per 1 million output tokens. If prompt caching is active, cache reads cost standard rates per 1M tokens.
Yes. Most major providers (including OpenAI, Anthropic, and Google) offer a 50% discount on standard token pricing for asynchronous Batch API jobs completed within a 24-hour turnaround window.
Prompt caching rates are not published for this specific tier or are included within base token pricing. Structure your system prompts to remain concise to minimize raw input cost.
Use our interactive AI Cost Calculator to input your exact daily request volume, average prompt token count, completion length, and prompt cache hit ratio. It generates precise cost forecasts across all models simultaneously.
Simulate real-world token consumption, model mixtures, and prompt cache savings with the DevFlow AI Cost Calculator.