Prompt Cache Savings Calculator
Compare uncached and cached input token costs and estimate monthly prompt caching savings.
Prompt Cache Savings Calculator
Compare normal and cached input pricing for repeated prompt content.
Frequently asked questions
Practical answers for applying this calculator to a production API billing or usage plan.
How does prompt caching reduce API cost?
Eligible repeated input tokens are billed at a lower cached-input rate. Savings depend on the provider's cache rules, the share of reusable prompt content, and the achieved cache hit rate.
Does a high cache hit rate guarantee savings?
Only when cached reads are cheaper than normal input and cache write or storage charges do not outweigh the discount. Model those extra charges separately when they apply.
Which tokens should count as cacheable?
Use stable system prompts, tool definitions, long reference documents, and repeated conversation prefixes that meet the provider's caching requirements.
Related API billing guides
Read the concepts behind the calculator and adapt the examples to your own API pricing model.
How to Price an API: Cost, Margin, Usage and Billing Models
A practical API pricing guide for turning delivery cost, margin, usage volume, free tiers, and billing models into a paid API price.
Read guideAPI Charges Explained: How API Costs and Pricing Work
Understand API charges, billable units, cost per 1,000 requests, token pricing, credits, free tiers, and invoice examples.
Read guideAPI Billing Models: Usage-Based, Subscription, Credits and BYOK
Compare API billing models for developer products, including usage-based billing, subscriptions, credits, prepaid wallets, and bring-your-own-key.
Read guide