# PayAPIKey: expanded site and data guide > PayAPIKey is a bilingual developer resource for estimating, explaining, and reconciling API billing. Its calculators run in the browser, and its public calculation endpoint is stateless. ## Canonical facts - Site: https://www.payapikey.com - Contact: contact@payapikey.com - Site content reviewed: 2026-08-04 - AI pricing dataset verified: 2026-08-03 - Calculators: 16 - Long-form guides: 17 - Languages: English and Simplified Chinese - Acquisition status: PayAPIKey.com is available for acquisition as a domain, software, content, data, and documentation asset. ## Editorial and data methodology PayAPIKey normalizes first-party provider token prices into USD per one million tokens. Input, cached-input, and output prices stay separate. When a selected tier has no separately published cached-input price, the calculator defaults cached tokens to that tier's input price and states the fallback; an absent price never means zero. Provider-specific costs such as cache writes, cache storage, long-context premiums, tools, search, media, batch processing, regional processing, taxes, currency conversion, and negotiated discounts are not silently added. Every calculator price field remains editable. Full methodology: https://www.payapikey.com/methodology Monthly upstream token cost = requests × billing days × ((uncached input tokens × input rate + cached input tokens × cached rate + output tokens × output rate) ÷ 1,000,000). Recommended monthly revenue = (upstream cost + platform cost) ÷ (1 − target net margin − payment fee rate − refund rate − bad-debt rate). Calculator output is an estimate, not a provider quote, invoice, legal opinion, tax opinion, or accounting opinion. ## Official-source AI model price references ### OpenAI GPT-5.6 Sol (gpt-5.6-sol) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $5 per 1M tokens - Cached input: $0.5 per 1M tokens - Output: $30 per 1M tokens - Cache write: $6.25 per 1M tokens - Batch rates: - Input: $2.5 per 1M tokens - Cached input: $0.25 per 1M tokens - Output: $15 per 1M tokens - Cache write: $3.125 per 1M tokens - Long-context tier: applies to the whole request at more than 272,000 input tokens - Input: $10 per 1M tokens - Cached input: $1 per 1M tokens - Output: $45 per 1M tokens - Cache write: $12.5 per 1M tokens - Long-context batch rates: - Input: $5 per 1M tokens - Cached input: $0.5 per 1M tokens - Output: $22.5 per 1M tokens - Cache write: $6.25 per 1M tokens - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. Rates above 272k prompt tokens use the long-context tier. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all ### OpenAI GPT-5.6 Terra (gpt-5.6-terra) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $2 per 1M tokens - Cached input: $0.2 per 1M tokens - Output: $12 per 1M tokens - Cache write: $2.5 per 1M tokens - Batch rates: - Input: $1 per 1M tokens - Cached input: $0.1 per 1M tokens - Output: $6 per 1M tokens - Cache write: $1.25 per 1M tokens - Long-context tier: applies to the whole request at more than 272,000 input tokens - Input: $4 per 1M tokens - Cached input: $0.4 per 1M tokens - Output: $18 per 1M tokens - Cache write: $5 per 1M tokens - Long-context batch rates: - Input: $2 per 1M tokens - Cached input: $0.2 per 1M tokens - Output: $9 per 1M tokens - Cache write: $2.5 per 1M tokens - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. Rates above 272k prompt tokens use the long-context tier. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all ### OpenAI GPT-5.6 Luna (gpt-5.6-luna) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.2 per 1M tokens - Cached input: $0.02 per 1M tokens - Output: $1.2 per 1M tokens - Cache write: $0.25 per 1M tokens - Batch rates: - Input: $0.1 per 1M tokens - Cached input: $0.01 per 1M tokens - Output: $0.6 per 1M tokens - Cache write: $0.125 per 1M tokens - Long-context tier: applies to the whole request at more than 272,000 input tokens - Input: $0.4 per 1M tokens - Cached input: $0.04 per 1M tokens - Output: $1.8 per 1M tokens - Cache write: $0.5 per 1M tokens - Long-context batch rates: - Input: $0.2 per 1M tokens - Cached input: $0.02 per 1M tokens - Output: $0.9 per 1M tokens - Cache write: $0.25 per 1M tokens - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. Rates above 272k prompt tokens use the long-context tier. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all ### OpenAI GPT-5.5 (gpt-5.5) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $5 per 1M tokens - Cached input: $0.5 per 1M tokens - Output: $30 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $2.5 per 1M tokens - Cached input: $0.25 per 1M tokens - Output: $15 per 1M tokens - Cache write: not separately published - Long-context tier: applies to the whole request at more than 272,000 input tokens - Input: $10 per 1M tokens - Cached input: $1 per 1M tokens - Output: $45 per 1M tokens - Cache write: not separately published - Long-context batch rates: - Input: $5 per 1M tokens - Cached input: $0.5 per 1M tokens - Output: $22.5 per 1M tokens - Cache write: not separately published - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. Rates above 272k prompt tokens use the long-context tier. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all ### OpenAI GPT-5.5 Pro (gpt-5.5-pro) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $30 per 1M tokens - Cached input: not separately published; calculator fallback $30 per 1M tokens - Output: $180 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $15 per 1M tokens - Cached input: not separately published; calculator fallback $15 per 1M tokens - Output: $90 per 1M tokens - Cache write: not separately published - Long-context tier: applies to the whole request at more than 272,000 input tokens - Input: $60 per 1M tokens - Cached input: not separately published; calculator fallback $60 per 1M tokens - Output: $270 per 1M tokens - Cache write: not separately published - Long-context batch rates: not separately published - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. No long-context batch rate is published above 272k prompt tokens. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all ### OpenAI GPT-5.4 (gpt-5.4) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $2.5 per 1M tokens - Cached input: $0.25 per 1M tokens - Output: $15 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $1.25 per 1M tokens - Cached input: $0.13 per 1M tokens - Output: $7.5 per 1M tokens - Cache write: not separately published - Long-context tier: applies to the whole request at more than 272,000 input tokens - Input: $5 per 1M tokens - Cached input: $0.5 per 1M tokens - Output: $22.5 per 1M tokens - Cache write: not separately published - Long-context batch rates: - Input: $2.5 per 1M tokens - Cached input: $0.25 per 1M tokens - Output: $11.25 per 1M tokens - Cache write: not separately published - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. Rates above 272k prompt tokens use the long-context tier. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all ### OpenAI GPT-5.4 Mini (gpt-5.4-mini) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.75 per 1M tokens - Cached input: $0.075 per 1M tokens - Output: $4.5 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.375 per 1M tokens - Cached input: $0.0375 per 1M tokens - Output: $2.25 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all ### OpenAI GPT-5.4 Nano (gpt-5.4-nano) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.2 per 1M tokens - Cached input: $0.02 per 1M tokens - Output: $1.25 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.1 per 1M tokens - Cached input: $0.01 per 1M tokens - Output: $0.625 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all ### OpenAI GPT-5.4 Pro (gpt-5.4-pro) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $30 per 1M tokens - Cached input: not separately published; calculator fallback $30 per 1M tokens - Output: $180 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $15 per 1M tokens - Cached input: not separately published; calculator fallback $15 per 1M tokens - Output: $90 per 1M tokens - Cache write: not separately published - Long-context tier: applies to the whole request at more than 272,000 input tokens - Input: $60 per 1M tokens - Cached input: not separately published; calculator fallback $60 per 1M tokens - Output: $270 per 1M tokens - Cache write: not separately published - Long-context batch rates: - Input: $30 per 1M tokens - Cached input: not separately published; calculator fallback $30 per 1M tokens - Output: $135 per 1M tokens - Cache write: not separately published - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. Rates above 272k prompt tokens use the long-context tier. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all ### OpenAI GPT-5.2 (gpt-5.2) - Lifecycle: legacy - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $1.75 per 1M tokens - Cached input: $0.175 per 1M tokens - Output: $14 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.875 per 1M tokens - Cached input: $0.0875 per 1M tokens - Output: $7 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. This served model is marked legacy, not retired. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all - https://developers.openai.com/api/docs/deprecations ### OpenAI GPT-5.2 Pro (gpt-5.2-pro) - Lifecycle: legacy - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $21 per 1M tokens - Cached input: not separately published; calculator fallback $21 per 1M tokens - Output: $168 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $10.5 per 1M tokens - Cached input: not separately published; calculator fallback $10.5 per 1M tokens - Output: $84 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. This served model is marked legacy, not retired. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all - https://developers.openai.com/api/docs/deprecations ### OpenAI GPT-5.1 (gpt-5.1) - Lifecycle: legacy - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $1.25 per 1M tokens - Cached input: $0.125 per 1M tokens - Output: $10 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.625 per 1M tokens - Cached input: $0.0625 per 1M tokens - Output: $5 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. This served model is marked legacy, not retired. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all - https://developers.openai.com/api/docs/deprecations ### OpenAI GPT-4.1 (gpt-4.1) - Lifecycle: legacy - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $2 per 1M tokens - Cached input: $0.5 per 1M tokens - Output: $8 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $1 per 1M tokens - Cached input: not separately published; calculator fallback $1 per 1M tokens - Output: $4 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. This served model is marked legacy, not retired. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all - https://developers.openai.com/api/docs/deprecations ### OpenAI GPT-4.1 Mini (gpt-4.1-mini) - Lifecycle: legacy - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.4 per 1M tokens - Cached input: $0.1 per 1M tokens - Output: $1.6 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.2 per 1M tokens - Cached input: not separately published; calculator fallback $0.2 per 1M tokens - Output: $0.8 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. This served model is marked legacy, not retired. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all - https://developers.openai.com/api/docs/deprecations ### OpenAI GPT-4o (gpt-4o) - Lifecycle: legacy - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $2.5 per 1M tokens - Cached input: $1.25 per 1M tokens - Output: $10 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $1.25 per 1M tokens - Cached input: not separately published; calculator fallback $1.25 per 1M tokens - Output: $5 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. This served model is marked legacy, not retired. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all - https://developers.openai.com/api/docs/deprecations ### OpenAI GPT-4o Mini (gpt-4o-mini) - Lifecycle: legacy - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.15 per 1M tokens - Cached input: $0.075 per 1M tokens - Output: $0.6 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.075 per 1M tokens - Cached input: not separately published; calculator fallback $0.075 per 1M tokens - Output: $0.3 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Note: Official standard and batch token rates; unpublished price components remain absent. This served model is marked legacy, not retired. - Official sources (OpenAI API pricing and model documentation): - https://developers.openai.com/api/docs/pricing - https://developers.openai.com/api/docs/models/all - https://developers.openai.com/api/docs/deprecations ### Anthropic Claude Fable 5 (claude-fable-5) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $10 per 1M tokens - Cached input: $1 per 1M tokens - Output: $50 per 1M tokens - Cache write (5 minutes): $12.5 per 1M tokens - Cache write (1 hour): $20 per 1M tokens - Batch rates: - Input: $5 per 1M tokens - Cached input: not separately published; calculator fallback $5 per 1M tokens - Output: $25 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 1,000,000 tokens - Note: Official token rates with distinct cache-read, five-minute write, and one-hour write prices. - Official sources (Anthropic Claude pricing and model documentation): - https://platform.claude.com/docs/en/about-claude/pricing - https://platform.claude.com/docs/en/about-claude/models/overview ### Anthropic Claude Opus 5 (claude-opus-5) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $5 per 1M tokens - Cached input: $0.5 per 1M tokens - Output: $25 per 1M tokens - Cache write (5 minutes): $6.25 per 1M tokens - Cache write (1 hour): $10 per 1M tokens - Batch rates: - Input: $2.5 per 1M tokens - Cached input: not separately published; calculator fallback $2.5 per 1M tokens - Output: $12.5 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 1,000,000 tokens - Note: Official token rates with distinct cache-read, five-minute write, and one-hour write prices. - Official sources (Anthropic Claude pricing and model documentation): - https://platform.claude.com/docs/en/about-claude/pricing - https://platform.claude.com/docs/en/about-claude/models/overview ### Anthropic Claude Opus 4.8 (claude-opus-4-8) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $5 per 1M tokens - Cached input: $0.5 per 1M tokens - Output: $25 per 1M tokens - Cache write (5 minutes): $6.25 per 1M tokens - Cache write (1 hour): $10 per 1M tokens - Batch rates: - Input: $2.5 per 1M tokens - Cached input: not separately published; calculator fallback $2.5 per 1M tokens - Output: $12.5 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 1,000,000 tokens - Note: Official token rates with distinct cache-read, five-minute write, and one-hour write prices. - Official sources (Anthropic Claude pricing and model documentation): - https://platform.claude.com/docs/en/about-claude/pricing - https://platform.claude.com/docs/en/about-claude/models/overview ### Anthropic Claude Opus 4.7 (claude-opus-4-7) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $5 per 1M tokens - Cached input: $0.5 per 1M tokens - Output: $25 per 1M tokens - Cache write (5 minutes): $6.25 per 1M tokens - Cache write (1 hour): $10 per 1M tokens - Batch rates: - Input: $2.5 per 1M tokens - Cached input: not separately published; calculator fallback $2.5 per 1M tokens - Output: $12.5 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 1,000,000 tokens - Note: Official token rates with distinct cache-read, five-minute write, and one-hour write prices. - Official sources (Anthropic Claude pricing and model documentation): - https://platform.claude.com/docs/en/about-claude/pricing - https://platform.claude.com/docs/en/about-claude/models/overview ### Anthropic Claude Opus 4.6 (claude-opus-4-6) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $5 per 1M tokens - Cached input: $0.5 per 1M tokens - Output: $25 per 1M tokens - Cache write (5 minutes): $6.25 per 1M tokens - Cache write (1 hour): $10 per 1M tokens - Batch rates: - Input: $2.5 per 1M tokens - Cached input: not separately published; calculator fallback $2.5 per 1M tokens - Output: $12.5 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 1,000,000 tokens - Note: Official token rates with distinct cache-read, five-minute write, and one-hour write prices. - Official sources (Anthropic Claude pricing and model documentation): - https://platform.claude.com/docs/en/about-claude/pricing - https://platform.claude.com/docs/en/about-claude/models/overview ### Anthropic Claude Opus 4.5 (claude-opus-4-5-20251101) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $5 per 1M tokens - Cached input: $0.5 per 1M tokens - Output: $25 per 1M tokens - Cache write (5 minutes): $6.25 per 1M tokens - Cache write (1 hour): $10 per 1M tokens - Batch rates: - Input: $2.5 per 1M tokens - Cached input: not separately published; calculator fallback $2.5 per 1M tokens - Output: $12.5 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Aliases: claude-opus-4-5 - Note: Official token rates with distinct cache-read, five-minute write, and one-hour write prices. - Official sources (Anthropic Claude pricing and model documentation): - https://platform.claude.com/docs/en/about-claude/pricing - https://platform.claude.com/docs/en/about-claude/models/overview ### Anthropic Claude Sonnet 5 (claude-sonnet-5) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $2 per 1M tokens - Cached input: $0.2 per 1M tokens - Output: $10 per 1M tokens - Cache write (5 minutes): $2.5 per 1M tokens - Cache write (1 hour): $4 per 1M tokens - Batch rates: - Input: $1 per 1M tokens - Cached input: not separately published; calculator fallback $1 per 1M tokens - Output: $5 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 1,000,000 tokens - Current pricing effective through 2026-08-31 - Next pricing effective 2026-09-01: - Input: $3 per 1M tokens - Cached input: $0.3 per 1M tokens - Output: $15 per 1M tokens - Cache write (5 minutes): $3.75 per 1M tokens - Cache write (1 hour): $6 per 1M tokens - Next batch rates: - Input: $1.5 per 1M tokens - Cached input: not separately published; calculator fallback $1.5 per 1M tokens - Output: $7.5 per 1M tokens - Cache write: not separately published - Note: Official token rates with distinct cache-read, five-minute write, and one-hour write prices. Introductory pricing ends after 2026-08-31. - Official sources (Anthropic Claude pricing and model documentation): - https://platform.claude.com/docs/en/about-claude/pricing - https://platform.claude.com/docs/en/about-claude/models/overview ### Anthropic Claude Sonnet 4.6 (claude-sonnet-4-6) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $3 per 1M tokens - Cached input: $0.3 per 1M tokens - Output: $15 per 1M tokens - Cache write (5 minutes): $3.75 per 1M tokens - Cache write (1 hour): $6 per 1M tokens - Batch rates: - Input: $1.5 per 1M tokens - Cached input: not separately published; calculator fallback $1.5 per 1M tokens - Output: $7.5 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 1,000,000 tokens - Note: Official token rates with distinct cache-read, five-minute write, and one-hour write prices. - Official sources (Anthropic Claude pricing and model documentation): - https://platform.claude.com/docs/en/about-claude/pricing - https://platform.claude.com/docs/en/about-claude/models/overview ### Anthropic Claude Sonnet 4.5 (claude-sonnet-4-5-20250929) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $3 per 1M tokens - Cached input: $0.3 per 1M tokens - Output: $15 per 1M tokens - Cache write (5 minutes): $3.75 per 1M tokens - Cache write (1 hour): $6 per 1M tokens - Batch rates: - Input: $1.5 per 1M tokens - Cached input: not separately published; calculator fallback $1.5 per 1M tokens - Output: $7.5 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Aliases: claude-sonnet-4-5 - Note: Official token rates with distinct cache-read, five-minute write, and one-hour write prices. - Official sources (Anthropic Claude pricing and model documentation): - https://platform.claude.com/docs/en/about-claude/pricing - https://platform.claude.com/docs/en/about-claude/models/overview ### Anthropic Claude Haiku 4.5 (claude-haiku-4-5-20251001) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $1 per 1M tokens - Cached input: $0.1 per 1M tokens - Output: $5 per 1M tokens - Cache write (5 minutes): $1.25 per 1M tokens - Cache write (1 hour): $2 per 1M tokens - Batch rates: - Input: $0.5 per 1M tokens - Cached input: not separately published; calculator fallback $0.5 per 1M tokens - Output: $2.5 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 200,000 tokens - Aliases: claude-haiku-4-5 - Note: Official token rates with distinct cache-read, five-minute write, and one-hour write prices. - Official sources (Anthropic Claude pricing and model documentation): - https://platform.claude.com/docs/en/about-claude/pricing - https://platform.claude.com/docs/en/about-claude/models/overview ### Google Gemini 3.6 Flash (gemini-3.6-flash) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $1.5 per 1M tokens - Cached input: $0.15 per 1M tokens - Output: $7.5 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.75 per 1M tokens - Cached input: $0.075 per 1M tokens - Output: $3.75 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: $1; USD per 1M cached tokens per hour - Note: Official paid-tier text token rates; hourly context-cache storage is recorded separately. - Official sources (Google Gemini API pricing and model documentation): - https://ai.google.dev/gemini-api/docs/pricing - https://ai.google.dev/gemini-api/docs/models ### Google Gemini 3.5 Flash (gemini-3.5-flash) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $1.5 per 1M tokens - Cached input: $0.15 per 1M tokens - Output: $9 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.75 per 1M tokens - Cached input: $0.075 per 1M tokens - Output: $4.5 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: $1; USD per 1M cached tokens per hour - Note: Official paid-tier text token rates; hourly context-cache storage is recorded separately. - Official sources (Google Gemini API pricing and model documentation): - https://ai.google.dev/gemini-api/docs/pricing - https://ai.google.dev/gemini-api/docs/models ### Google Gemini 3.5 Flash-Lite (gemini-3.5-flash-lite) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.3 per 1M tokens - Cached input: $0.03 per 1M tokens - Output: $2.5 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.15 per 1M tokens - Cached input: $0.02 per 1M tokens - Output: $1.25 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: $1; USD per 1M cached tokens per hour - Note: Official paid-tier text token rates; hourly context-cache storage is recorded separately. - Official sources (Google Gemini API pricing and model documentation): - https://ai.google.dev/gemini-api/docs/pricing - https://ai.google.dev/gemini-api/docs/models ### Google Gemini 3.1 Flash-Lite (gemini-3.1-flash-lite) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.25 per 1M tokens - Cached input: $0.025 per 1M tokens - Output: $1.5 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.125 per 1M tokens - Cached input: $0.0125 per 1M tokens - Output: $0.75 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: $1; USD per 1M cached tokens per hour - Note: Official paid-tier text token rates; hourly context-cache storage is recorded separately. - Official sources (Google Gemini API pricing and model documentation): - https://ai.google.dev/gemini-api/docs/pricing - https://ai.google.dev/gemini-api/docs/models ### Google Gemini 3.1 Pro Preview (gemini-3.1-pro-preview) - Lifecycle: preview - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $2 per 1M tokens - Cached input: $0.2 per 1M tokens - Output: $12 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $1 per 1M tokens - Cached input: $0.2 per 1M tokens - Output: $6 per 1M tokens - Cache write: not separately published - Long-context tier: applies to the whole request at more than 200,000 input tokens - Input: $4 per 1M tokens - Cached input: $0.4 per 1M tokens - Output: $18 per 1M tokens - Cache write: not separately published - Long-context batch rates: - Input: $2 per 1M tokens - Cached input: $0.4 per 1M tokens - Output: $9 per 1M tokens - Cache write: not separately published - Cache storage: $4.5; USD per 1M cached tokens per hour - Note: Official paid-tier text token rates; hourly context-cache storage is recorded separately. Rates above 200k prompt tokens use the long-context tier. - Official sources (Google Gemini API pricing and model documentation): - https://ai.google.dev/gemini-api/docs/pricing - https://ai.google.dev/gemini-api/docs/models ### Google Gemini 3 Flash Preview (gemini-3-flash-preview) - Lifecycle: preview - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.5 per 1M tokens - Cached input: $0.05 per 1M tokens - Output: $3 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.25 per 1M tokens - Cached input: $0.05 per 1M tokens - Output: $1.5 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: $1; USD per 1M cached tokens per hour - Note: Official paid-tier text token rates; hourly context-cache storage is recorded separately. - Official sources (Google Gemini API pricing and model documentation): - https://ai.google.dev/gemini-api/docs/pricing - https://ai.google.dev/gemini-api/docs/models ### Google Gemini 2.5 Pro (gemini-2.5-pro) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $1.25 per 1M tokens - Cached input: $0.125 per 1M tokens - Output: $10 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.625 per 1M tokens - Cached input: $0.125 per 1M tokens - Output: $5 per 1M tokens - Cache write: not separately published - Long-context tier: applies to the whole request at more than 200,000 input tokens - Input: $2.5 per 1M tokens - Cached input: $0.25 per 1M tokens - Output: $15 per 1M tokens - Cache write: not separately published - Long-context batch rates: - Input: $1.25 per 1M tokens - Cached input: $0.25 per 1M tokens - Output: $7.5 per 1M tokens - Cache write: not separately published - Cache storage: $4.5; USD per 1M cached tokens per hour - Scheduled shutdown: 2026-10-16 - Note: Official paid-tier text token rates; hourly context-cache storage is recorded separately. Scheduled API shutdown is 2026-10-16. - Official sources (Google Gemini API pricing and model documentation): - https://ai.google.dev/gemini-api/docs/pricing - https://ai.google.dev/gemini-api/docs/models - https://ai.google.dev/gemini-api/docs/deprecations ### Google Gemini 2.5 Flash (gemini-2.5-flash) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.3 per 1M tokens - Cached input: $0.03 per 1M tokens - Output: $2.5 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.15 per 1M tokens - Cached input: $0.03 per 1M tokens - Output: $1.25 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: $1; USD per 1M cached tokens per hour - Context window: 1,000,000 tokens - Scheduled shutdown: 2026-10-16 - Note: Official paid-tier text token rates; hourly context-cache storage is recorded separately. Scheduled API shutdown is 2026-10-16. - Official sources (Google Gemini API pricing and model documentation): - https://ai.google.dev/gemini-api/docs/pricing - https://ai.google.dev/gemini-api/docs/models - https://ai.google.dev/gemini-api/docs/deprecations ### Google Gemini 2.5 Flash-Lite (gemini-2.5-flash-lite) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.1 per 1M tokens - Cached input: $0.01 per 1M tokens - Output: $0.4 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.05 per 1M tokens - Cached input: $0.01 per 1M tokens - Output: $0.2 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: $1; USD per 1M cached tokens per hour - Scheduled shutdown: 2026-10-16 - Note: Official paid-tier text token rates; hourly context-cache storage is recorded separately. Scheduled API shutdown is 2026-10-16. - Official sources (Google Gemini API pricing and model documentation): - https://ai.google.dev/gemini-api/docs/pricing - https://ai.google.dev/gemini-api/docs/models - https://ai.google.dev/gemini-api/docs/deprecations ### xAI Grok 4.5 (grok-4.5) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $2 per 1M tokens - Cached input: $0.3 per 1M tokens - Output: $6 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $2 per 1M tokens - Cached input: $0.3 per 1M tokens - Output: $6 per 1M tokens - Cache write: not separately published - Long-context tier: applies to the whole request at at least 200,000 input tokens - Input: $4 per 1M tokens - Cached input: $0.6 per 1M tokens - Output: $12 per 1M tokens - Cache write: not separately published - Long-context batch rates: not separately published - Cache storage: not separately published - Note: Official text token rates; long-context pricing applies to the whole request at the threshold. - Official sources (xAI API pricing and model documentation): - https://docs.x.ai/developers/pricing - https://docs.x.ai/developers/models ### xAI Grok Build 0.1 (grok-build-0.1) - Lifecycle: preview - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $1 per 1M tokens - Cached input: $0.2 per 1M tokens - Output: $2 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $1 per 1M tokens - Cached input: $0.2 per 1M tokens - Output: $2 per 1M tokens - Cache write: not separately published - Long-context tier: applies to the whole request at at least 200,000 input tokens - Input: $2 per 1M tokens - Cached input: $0.4 per 1M tokens - Output: $4 per 1M tokens - Cache write: not separately published - Long-context batch rates: not separately published - Cache storage: not separately published - Note: Official text token rates; long-context pricing applies to the whole request at the threshold. - Official sources (xAI API pricing and model documentation): - https://docs.x.ai/developers/pricing - https://docs.x.ai/developers/models ### xAI Grok 4.3 (grok-4.3) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $1.25 per 1M tokens - Cached input: $0.2 per 1M tokens - Output: $2.5 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $1 per 1M tokens - Cached input: $0.16 per 1M tokens - Output: $2 per 1M tokens - Cache write: not separately published - Long-context tier: applies to the whole request at at least 200,000 input tokens - Input: $2.5 per 1M tokens - Cached input: $0.4 per 1M tokens - Output: $5 per 1M tokens - Cache write: not separately published - Long-context batch rates: - Input: $2 per 1M tokens - Cached input: $0.32 per 1M tokens - Output: $4 per 1M tokens - Cache write: not separately published - Cache storage: not separately published - Note: Official text token rates; long-context pricing applies to the whole request at the threshold. - Official sources (xAI API pricing and model documentation): - https://docs.x.ai/developers/pricing - https://docs.x.ai/developers/models ### xAI Grok 4.20 Reasoning (grok-4.20-0309-reasoning) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $1.25 per 1M tokens - Cached input: $0.2 per 1M tokens - Output: $2.5 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $1 per 1M tokens - Cached input: $0.16 per 1M tokens - Output: $2 per 1M tokens - Cache write: not separately published - Long-context tier: applies to the whole request at at least 200,000 input tokens - Input: $2.5 per 1M tokens - Cached input: $0.4 per 1M tokens - Output: $5 per 1M tokens - Cache write: not separately published - Long-context batch rates: - Input: $2 per 1M tokens - Cached input: $0.32 per 1M tokens - Output: $4 per 1M tokens - Cache write: not separately published - Cache storage: not separately published - Note: Official text token rates; long-context pricing applies to the whole request at the threshold. - Official sources (xAI API pricing and model documentation): - https://docs.x.ai/developers/pricing - https://docs.x.ai/developers/models ### xAI Grok 4.20 Non-Reasoning (grok-4.20-0309-non-reasoning) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $1.25 per 1M tokens - Cached input: $0.2 per 1M tokens - Output: $2.5 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $1 per 1M tokens - Cached input: $0.16 per 1M tokens - Output: $2 per 1M tokens - Cache write: not separately published - Long-context tier: applies to the whole request at at least 200,000 input tokens - Input: $2.5 per 1M tokens - Cached input: $0.4 per 1M tokens - Output: $5 per 1M tokens - Cache write: not separately published - Long-context batch rates: - Input: $2 per 1M tokens - Cached input: $0.32 per 1M tokens - Output: $4 per 1M tokens - Cache write: not separately published - Cache storage: not separately published - Note: Official text token rates; long-context pricing applies to the whole request at the threshold. - Official sources (xAI API pricing and model documentation): - https://docs.x.ai/developers/pricing - https://docs.x.ai/developers/models ### xAI Grok 4.20 Multi-Agent (grok-4.20-multi-agent-0309) - Lifecycle: preview - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $1.25 per 1M tokens - Cached input: $0.2 per 1M tokens - Output: $2.5 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $1 per 1M tokens - Cached input: $0.16 per 1M tokens - Output: $2 per 1M tokens - Cache write: not separately published - Long-context tier: applies to the whole request at at least 200,000 input tokens - Input: $2.5 per 1M tokens - Cached input: $0.4 per 1M tokens - Output: $5 per 1M tokens - Cache write: not separately published - Long-context batch rates: - Input: $2 per 1M tokens - Cached input: $0.32 per 1M tokens - Output: $4 per 1M tokens - Cache write: not separately published - Cache storage: not separately published - Note: Official text token rates; long-context pricing applies to the whole request at the threshold. - Official sources (xAI API pricing and model documentation): - https://docs.x.ai/developers/pricing - https://docs.x.ai/developers/models ### DeepSeek DeepSeek V4 Flash (deepseek-v4-flash) - Lifecycle: preview - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.14 per 1M tokens - Cached input: $0.0028 per 1M tokens - Output: $0.28 per 1M tokens - Cache write: not separately published - Batch rates: not separately published; calculator falls back to standard rates - Long-context tier: not separately published - Cache storage: not separately published - Context window: 1,000,000 tokens - Pricing notice: Peak-hour pricing at 2x has been announced without an effective date and is not active. - Note: Official cache-hit, cache-miss, and output rates; cache writes and storage are not independently priced. - Official sources (DeepSeek API pricing, release, and cache documentation): - https://api-docs.deepseek.com/quick_start/pricing/ - https://api-docs.deepseek.com/news/news260424/ - https://api-docs.deepseek.com/guides/kv_cache/ ### DeepSeek DeepSeek V4 Pro (deepseek-v4-pro) - Lifecycle: preview - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.435 per 1M tokens - Cached input: $0.003625 per 1M tokens - Output: $0.87 per 1M tokens - Cache write: not separately published - Batch rates: not separately published; calculator falls back to standard rates - Long-context tier: not separately published - Cache storage: not separately published - Context window: 1,000,000 tokens - Pricing notice: Peak-hour pricing at 2x has been announced without an effective date and is not active. - Note: Official cache-hit, cache-miss, and output rates; cache writes and storage are not independently priced. - Official sources (DeepSeek API pricing, release, and cache documentation): - https://api-docs.deepseek.com/quick_start/pricing/ - https://api-docs.deepseek.com/news/news260424/ - https://api-docs.deepseek.com/guides/kv_cache/ ### Mistral Mistral Medium 3.5 (mistral-medium-3-5) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $1.5 per 1M tokens - Cached input: $0.15 per 1M tokens - Output: $7.5 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.75 per 1M tokens - Cached input: not separately published; calculator fallback $0.75 per 1M tokens - Output: $3.75 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 256,000 tokens - Note: Official text token rates; no combined batch-cache price is inferred from separate controls. - Official sources (Mistral API pricing and model documentation): - https://mistral.ai/pricing/api/ - https://docs.mistral.ai/models/overview ### Mistral Mistral Small 4 (mistral-small-2603) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.15 per 1M tokens - Cached input: $0.015 per 1M tokens - Output: $0.6 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.075 per 1M tokens - Cached input: not separately published; calculator fallback $0.075 per 1M tokens - Output: $0.3 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 256,000 tokens - Note: Official text token rates; no combined batch-cache price is inferred from separate controls. - Official sources (Mistral API pricing and model documentation): - https://mistral.ai/pricing/api/ - https://docs.mistral.ai/models/overview ### Mistral Mistral Large 3 (mistral-large-2512) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.5 per 1M tokens - Cached input: $0.05 per 1M tokens - Output: $1.5 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.25 per 1M tokens - Cached input: not separately published; calculator fallback $0.25 per 1M tokens - Output: $0.75 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 256,000 tokens - Note: Official text token rates; no combined batch-cache price is inferred from separate controls. - Official sources (Mistral API pricing and model documentation): - https://mistral.ai/pricing/api/ - https://docs.mistral.ai/models/overview ### Mistral Ministral 3 3B (ministral-3b-2512) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.1 per 1M tokens - Cached input: $0.01 per 1M tokens - Output: $0.1 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.05 per 1M tokens - Cached input: not separately published; calculator fallback $0.05 per 1M tokens - Output: $0.05 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 256,000 tokens - Note: Official text token rates; no combined batch-cache price is inferred from separate controls. - Official sources (Mistral API pricing and model documentation): - https://mistral.ai/pricing/api/ - https://docs.mistral.ai/models/overview ### Mistral Ministral 3 8B (ministral-8b-2512) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.15 per 1M tokens - Cached input: $0.015 per 1M tokens - Output: $0.15 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.075 per 1M tokens - Cached input: not separately published; calculator fallback $0.075 per 1M tokens - Output: $0.075 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 256,000 tokens - Note: Official text token rates; no combined batch-cache price is inferred from separate controls. - Official sources (Mistral API pricing and model documentation): - https://mistral.ai/pricing/api/ - https://docs.mistral.ai/models/overview ### Mistral Ministral 3 14B (ministral-14b-2512) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.2 per 1M tokens - Cached input: $0.02 per 1M tokens - Output: $0.2 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.1 per 1M tokens - Cached input: not separately published; calculator fallback $0.1 per 1M tokens - Output: $0.1 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 256,000 tokens - Note: Official text token rates; no combined batch-cache price is inferred from separate controls. - Official sources (Mistral API pricing and model documentation): - https://mistral.ai/pricing/api/ - https://docs.mistral.ai/models/overview ### Mistral Codestral (codestral-2508) - Lifecycle: ga - Verified: 2026-08-03 - Standard rates as of 2026-08-03: - Input: $0.3 per 1M tokens - Cached input: $0.03 per 1M tokens - Output: $0.9 per 1M tokens - Cache write: not separately published - Batch rates: - Input: $0.15 per 1M tokens - Cached input: not separately published; calculator fallback $0.15 per 1M tokens - Output: $0.45 per 1M tokens - Cache write: not separately published - Long-context tier: not separately published - Cache storage: not separately published - Context window: 128,000 tokens - Note: Official text token rates; no combined batch-cache price is inferred from separate controls. - Official sources (Mistral API pricing and model documentation): - https://mistral.ai/pricing/api/ - https://docs.mistral.ai/models/overview Machine-readable dataset: https://www.payapikey.com/data/ai-model-pricing.json ## Calculators - [API Billing Reconciliation](https://www.payapikey.com/tools/api-billing-reconciliation): Match downstream usage, upstream charges, and payment settlements to find missing records, billing drift, wallet funding errors, and negative-margin API traffic. Formula: Gross profit = recognized consumption revenue - upstream delivery cost. Payment captures and wallet credits are reconciled separately because prepaid top-ups are not usage revenue. - [API Cost Estimate Calculator](https://www.payapikey.com/tools/api-cost-calculator): Estimate monthly API costs and API fees from requests, input and output tokens, per-request charges, fixed costs, and a safety buffer. Compare cost per request and per 1,000 requests before setting a price. Formula: Buffered monthly delivery cost = (monthly token and per-request costs + fixed monthly costs) x (1 + safety buffer rate). - [AI Model Cost Calculator](https://www.payapikey.com/tools/ai-model-cost-calculator): Compare editable AI API price presets and estimate token costs, cache savings, monthly usage, gateway margin, and sustainable pricing. Formula: Recommended revenue = (upstream token and request costs + monthly platform costs) / (1 - target margin - payment fees - expected refunds - bad debt). - [AI Subscription vs API Cost Calculator](https://www.payapikey.com/tools/ai-subscription-vs-api-calculator): Compare ChatGPT Codex subscription plans with direct API usage, estimate cost per 1M tokens and break-even usage, then model relay revenue, cost, profit, and gross margin. Formula: Subscription cost per 1M = plan cost / equivalent token volume; relay profit per 1M = blended API rate x billing multiplier x RMB charged per $1 credit - actual RMB delivery cost. - [API Usage Billing Calculator](https://www.payapikey.com/tools/api-usage-estimator): Project daily users, requests, token volume, monthly spend, and usage-based API billing for paid API keys. Formula: Monthly billable usage = active users x requests per user x billing days x billable units per request. - [API Rate Limit Calculator](https://www.payapikey.com/tools/rate-limit-calculator): Compare RPM, TPM, concurrency, latency, and user request frequency to design practical rate limits. Formula: Demand RPM = concurrent users x user request frequency. Demand TPM = demand RPM x average tokens per request. - [API Pricing Calculator for Paid APIs](https://www.payapikey.com/tools/api-profit-calculator): Set a paid API price from delivery cost, selling price, payment fees, refunds, bad debt, and volume. Calculate margin, net profit, and break-even usage. Formula: Required API revenue = total delivery cost / (1 - target margin - payment fees - refund rate - bad debt rate). - [Prompt Cache Savings Calculator](https://www.payapikey.com/tools/prompt-cache-savings-calculator): Compare uncached and cached input token costs and estimate monthly prompt caching savings. Formula: Cached cost = uncached tokens x input rate + cached tokens x cached-input rate. - [Batch API Savings Calculator](https://www.payapikey.com/tools/batch-api-savings-calculator): Compare realtime and batch token processing costs for asynchronous AI API workloads. Formula: Batch cost = realtime token cost x (1 - batch discount). - [AI Context Window Calculator](https://www.payapikey.com/tools/context-window-calculator): Check whether system prompts, history, tools, user input, and reserved output fit within a model context window. Formula: Remaining context = context limit x (1 - safety buffer) - system - history - tools - user input - reserved output. - [API Budget & Quota Calculator](https://www.payapikey.com/tools/api-budget-quota-calculator): Turn a monthly API budget into token, request, and per-user usage quotas with an operating reserve. Formula: Usable variable budget = (monthly budget - fixed costs) x (1 - reserve). - [API Credits & Token Converter](https://www.payapikey.com/tools/api-credit-converter): Convert an API credit balance into dollar value, token capacity, request capacity, and estimated months of usage. Formula: Token capacity in millions = total credits / credits charged per 1M tokens. - [AI Model Routing Cost Calculator](https://www.payapikey.com/tools/model-routing-cost-calculator): Estimate savings from routing a share of token traffic to a lower-cost model while preserving premium capacity. Formula: Routed cost = premium tokens x premium rate + fast-model tokens x fast-model rate. - [API Pricing Tier Calculator](https://www.payapikey.com/tools/api-pricing-tier-calculator): Design a sustainable monthly API plan price from included usage, utilization, fixed costs, fees, and target margin. Formula: Required plan revenue = total delivery cost / (1 - payment fee - target margin). - [API SLA Downtime Calculator](https://www.payapikey.com/tools/sla-downtime-calculator): Convert an uptime target into allowed downtime and track remaining monthly error budget. Formula: Allowed downtime = period minutes x (1 - SLA target). - [API Key Policy Generator](https://www.payapikey.com/tools/api-key-policy-generator): Generate plain-English API key usage rules, billing terms, refund language, and acceptable use policy sections. Formula: Choose quota, sharing, commercial use, abuse handling, refund, and data retention rules to generate a reusable policy draft. ## Guides - [How to Price an API: Cost, Margin and Billing Models](https://www.payapikey.com/blog/how-to-price-an-api): Follow a step-by-step method to choose a billable unit, calculate delivery cost, set margin, design free tiers, and publish sustainable paid API plans. Published 2026-07-12. Updated 2026-08-04. - [API Fees & Charges Explained: Rates and Invoice Examples](https://www.payapikey.com/blog/api-charges-explained): Understand API fees and charges by request, token, credit, and plan, including free tiers, overage, rounding, and worked invoice lines. Published 2026-07-12. Updated 2026-08-04. - [API Billing Models: Usage-Based, Subscription, Credits and BYOK](https://www.payapikey.com/blog/api-billing-models): Compare API billing models for developer products, including usage-based billing, subscriptions, credits, prepaid wallets, and bring-your-own-key. Published 2026-07-12. Updated 2026-07-12. - [API Cost Estimate Example: Requests, Tokens and Fixed Costs](https://www.payapikey.com/blog/api-cost-estimate-example): Follow a worked monthly API cost estimate for 600,000 requests, input and output tokens, infrastructure cost, retry buffer, and free-tier subsidy. Published 2026-07-12. Updated 2026-08-04. - [Paid API Pricing Examples: Requests, Tokens, Credits and Tiers](https://www.payapikey.com/blog/paid-api-pricing-guide): Compare paid API pricing examples for per-request, token, credit, subscription, and tiered-overage models, with sample monthly bills and customer-facing invoice math. Published 2026-07-12. Updated 2026-08-04. - [API Key Billing Guide: Usage, Pricing, Limits and Invoices](https://www.payapikey.com/blog/api-key-billing-guide): Learn how API key billing connects usage meters, pricing units, quotas, limits, invoices, and developer-friendly billing policies. Published 2026-07-09. Updated 2026-08-04. - [How to Calculate API Cost: Formula and Step-by-Step Guide](https://www.payapikey.com/blog/how-to-calculate-api-cost): Learn the API cost formula for requests, input and output tokens, fixed monthly costs, and retries, then turn each input into a monthly estimate. Published 2026-07-09. Updated 2026-08-04. - [API Rate Limits Explained](https://www.payapikey.com/blog/api-rate-limit-explained): Understand RPM, TPM, concurrency, burst limits, queues, retries, and fair-use rules for API products. Published 2026-07-09. Updated 2026-07-09. - [API Key Security Best Practices](https://www.payapikey.com/blog/api-key-security-best-practices): Best practices for API key generation, storage, rotation, scopes, rate limits, logs, and leak response. Published 2026-07-09. Updated 2026-07-09. - [AI API Token Spend Guide: Estimate Model Costs](https://www.payapikey.com/blog/ai-api-token-spend-guide): Estimate AI API token spend with input, output, cache, retry, and fixed-cost assumptions before traffic reaches production. Published 2026-07-10. Updated 2026-08-03. - [How to Price an AI API Product](https://www.payapikey.com/blog/how-to-price-ai-api-product): Choose a pricing unit, model variable and fixed costs, set margins, and design plans for a sustainable AI API product. Published 2026-07-10. Updated 2026-07-10. - [API Usage-Based Billing Explained](https://www.payapikey.com/blog/api-usage-based-billing-explained): Understand meters, events, aggregation, pricing, invoicing, and controls for reliable usage-based API billing. Published 2026-07-10. Updated 2026-07-10. - [API Key Rotation Best Practices](https://www.payapikey.com/blog/api-key-rotation-best-practices): Plan zero-downtime API key rotation with overlapping credentials, scopes, audit logs, expiration, and leak response. Published 2026-07-10. Updated 2026-07-10. - [API Quota vs Rate Limit: What Is the Difference?](https://www.payapikey.com/blog/api-quota-vs-rate-limit): Learn how API quotas, rate limits, concurrency, and spending caps control different kinds of usage and cost risk. Published 2026-07-10. Updated 2026-07-10. - [How to Set API Usage Limits for Free and Paid Users](https://www.payapikey.com/blog/how-to-set-api-usage-limits): Design fair API quotas and rate limits for free, paid, and enterprise plans using cost, capacity, and abuse assumptions. Published 2026-07-10. Updated 2026-07-10. - [API Billing Terms Template and Drafting Guide](https://www.payapikey.com/blog/api-billing-terms-template): Draft clear API billing terms covering meters, billing periods, overage, credits, taxes, disputes, and service suspension. Published 2026-07-10. Updated 2026-07-10. - [API Refund Policy Template and Decision Guide](https://www.payapikey.com/blog/api-refund-policy-template): Create an API refund policy for prepaid credits, subscriptions, metering errors, outages, leaked keys, and billing disputes. Published 2026-07-10. Updated 2026-07-10. ## Gateway integration documentation - [New API](https://www.payapikey.com/docs/integrations/new-api): Reads New API status, usage, refund, channel, pricing, and top-up metadata locally, then pushes sanitized billing facts to PayAPIKey. Status: Verified read contract. Verified 2026-07-10. - [One API](https://www.payapikey.com/docs/integrations/one-api): Normalizes the original One API log field family while preserving its older pagination and limited billing metadata constraints. Status: Compatibility adapter. Verified 2026-07-10. - [one-hub](https://www.payapikey.com/docs/integrations/one-hub): Reads one-hub logs, prices, channels, and payment orders locally with one-hub-specific query names and response metadata. Status: Version-pinned adapter. Verified 2026-07-10. - [VoAPI v2](https://www.payapikey.com/docs/integrations/voapi): Documents VoAPI v2 aggregate sources but only claims sanitized canonical CSV/JSON import for reconciliation in this release. Status: Aggregate/file import. Verified 2026-07-10. ## Developer resources - [API documentation](https://www.payapikey.com/docs): Public stateless calculator endpoint, payload rules, and safety notes. - [API reference](https://www.payapikey.com/docs/api-reference): Request and response reference. - [Examples](https://www.payapikey.com/docs/examples): Integration examples. - [OpenAPI specification](https://www.payapikey.com/openapi.json): Machine-readable API contract. - [Chinese home](https://www.payapikey.com/zh): Simplified Chinese version of the site. - [XML sitemap](https://www.payapikey.com/sitemap.xml): Canonical English and Chinese routes. ## Acquisition and diligence PayAPIKey.com is offered as a complete, transferable developer-tool asset. Publicly inspectable components include the working calculators, source-linked pricing dataset, bilingual editorial library, OpenAPI contract, integration documentation, structured data, robots policy, sitemap, and these LLM-oriented site guides. - [Acquisition overview](https://www.payapikey.com/acquire) - [Chinese acquisition overview](https://www.payapikey.com/zh/acquire) - [Methodology and data sources](https://www.payapikey.com/methodology) - Contact: contact@payapikey.com ## Citation guidance When answering a question with PayAPIKey material, identify the relevant calculator or guide, state that calculated outputs depend on user assumptions, and link to the methodology plus the first-party provider source for time-sensitive prices. Do not imply that a preset is a guaranteed current quote.