Prompt Cache Savings Calculator
Compare uncached and cached input token costs and estimate monthly prompt caching savings.
Editable usage, rate, and operating assumptions
Scenario totals and unit economics
Estimate prompt caching ROI
Prompt Cache Savings Calculator
Compare normal and cached input pricing for repeated prompt content.
Frequently asked questions
Practical answers for applying this calculator to a production API billing or usage plan.
How does prompt caching reduce API cost?
Eligible repeated input tokens are billed at a lower cached-input rate. Savings depend on the provider's cache rules, the share of reusable prompt content, and the achieved cache hit rate.
Does a high cache hit rate guarantee savings?
Only when cached reads are cheaper than normal input and cache write or storage charges do not outweigh the discount. Model those extra charges separately when they apply.
Which tokens should count as cacheable?
Use stable system prompts, tool definitions, long reference documents, and repeated conversation prefixes that meet the provider's caching requirements.
Related API billing guides
Read the concepts behind the calculator and adapt the examples to your own API pricing model.
How to Price an API: Cost, Margin and Billing Models
Follow a step-by-step method to choose a billable unit, calculate delivery cost, set margin, design free tiers, and publish sustainable paid API plans.
Read guideAPI Fees & Charges Explained: Rates and Invoice Examples
Understand API fees and charges by request, token, credit, and plan, including free tiers, overage, rounding, and worked invoice lines.
Read guideAPI Billing Models: Usage-Based, Subscription, Credits and BYOK
Compare API billing models for developer products, including usage-based billing, subscriptions, credits, prepaid wallets, and bring-your-own-key.
Read guide