AI API Cost Calculator

Price your workload on GPT-6, GPT-5.6, Claude and Gemini side by side, at each provider’s official rates, adjusted for how each model counts tokens.

Your workload

Start from:

2,000-token prompt (instructions + chat history, 60% cached), 300-token reply, 1,000 chats a day.

Everything you send: instructions, history, documents
The reply, including any reasoning tokens
Counted in
The repeated start of each prompt, like a fixed system prompt or chat history

Cost per model, cheapest first

Prices checked September 19, 2026

  1. GPT-5.6 LunaOpenAI · 2,000 in / 300 out tokens
    $16.32/month$0.00054 per request
  2. Gemini 3.5 Flash-LiteGoogle · 2,220 in / 333 out tokens
    $34.17/month$0.0011 per request · 2.1× the cheapest
  3. Gemini 3.8 FlashGoogle · 2,220 in / 333 out tokensPrice through December 31, 2026. From January 1, 2027 it doubles to $1.50 input and $7.50 output.
    $60.44/month$0.002 per request · 3.7× the cheapest
  4. Claude Haiku 4.5Anthropic · 2,100 in / 315 out tokens
    $76.23/month$0.0025 per request · 4.7× the cheapest
  5. GPT-5.6 TerraOpenAI · 2,000 in / 300 out tokens
    $163/month$0.0054 per request · 10× the cheapest
  6. Gemini 3.1 ProGoogle · 2,220 in / 333 out tokensPreview model.
    $181/month$0.006 per request · 11× the cheapest
  7. Claude Sonnet 5Anthropic · 3,000 in / 450 out tokens
    $218/month$0.0073 per request · 13× the cheapest
  8. GPT-5.6 SolOpenAI · 2,000 in / 300 out tokensPromotional price, available at least through November 21, 2026.
    $290/month$0.0097 per request · 18× the cheapest
  9. Claude Opus 5Anthropic · 3,000 in / 450 out tokens
    $545/month$0.018 per request · 33× the cheapest
  10. GPT-6 AstraOpenAI · 2,000 in / 300 out tokens
    $726/month$0.024 per request · 44× the cheapest
  11. Claude Fable 5.1Anthropic · 3,000 in / 450 out tokens
    $1,049/month$0.035 per request · 64× the cheapest

A month is 30 days. Cached input is billed at each provider's cache-read price, assuming the cache stays warm. Tokens are adjusted with the ratios from our token counter. Taxes, web search and other tool fees aren't included.

API prices per million tokens

Standard prices from each provider's own pricing page, checked September 19, 2026. Batch prices are for jobs you can wait up to 24 hours for.

ModelInputCached inputOutputBatch in / out
OpenAI official prices
GPT-6 AstraOver 272K input tokens: $20 in, $75 out, for the whole request$10$1$50$5 / $25
GPT-5.6 SolOver 272K input tokens: $8 in, $30 out, for the whole requestPromotional price, available at least through November 21, 2026.$4$0.40$20$2 / $10
GPT-5.6 TerraOver 272K input tokens: $4 in, $18 out, for the whole request$2$0.20$12$1 / $6
GPT-5.6 LunaOver 272K input tokens: $0.40 in, $1.80 out, for the whole request$0.20$0.02$1.20$0.10 / $0.60
Anthropic official prices
Claude Fable 5.1$10$0.25$50$5 / $25
Claude Opus 5$5$0.50$25$2.50 / $12.50
Claude Sonnet 5$2$0.20$10$1 / $5
Claude Haiku 4.5$1$0.10$5$0.50 / $2.50
Google official prices
Gemini 3.1 ProOver 200K input tokens: $4 in, $18 out, for the whole requestPreview model.$2$0.20$12$1 / $6
Gemini 3.8 FlashPrice through December 31, 2026. From January 1, 2027 it doubles to $1.50 input and $7.50 output.$0.75$0.075$3.75$0.375 / $1.875
Gemini 3.5 Flash-Lite$0.30$0.03$2.50$0.15 / $1.25

What common workloads cost per month

The calculator's four starting points, priced on every model with tokens adjusted for each tokenizer. The cheapest model for each job is in bold.

ModelSupport chatbotDocument summariesCoding agentBulk tagging (batch)
GPT-6 Astra$726$780$2,910$1,440
GPT-5.6 Sol$290$312$1,164$576
GPT-5.6 Terra$163$163$642$306
GPT-5.6 Luna$16.32$16.32$64.20$30.60
Claude Fable 5.1$1,049$1,170$3,791$2,059
Claude Opus 5$545$585$2,183$1,080
Claude Sonnet 5$218$234$873$432
Claude Haiku 4.5$76.23$81.90$306$152
Gemini 3.1 Pro$181$181$713$359
Gemini 3.8 Flash$60.44$64.94$242$120
Gemini 3.5 Flash-Lite$34.17$29.97$130$58.72
  • Support chatbot: 2,000-token prompt (instructions + chat history, 60% cached), 300-token reply, 1,000 chats a day.
  • Document summaries: 10,000-token document, 600-token summary, 200 a day, no caching.
  • Coding agent: 40,000-token context (85% cached), 2,000 tokens out per step, 500 steps a day.
  • Bulk tagging (batch): 600-token ticket (50% cached), 30-token label, 20,000 a day through the batch API.

What this calculator does

This AI API cost calculator prices the same workload on 11 current models from OpenAI, Anthropic and Google, and ranks them from cheapest to most expensive. You describe one request (how much text goes in, how much comes back, how often) and it works out the cost per request and per month. It handles the things that move a real bill: output tokens that cost several times more than input, cached prompts, batch discounts, and the long-context surcharges OpenAI and Google add to very large prompts. It also does something most calculators skip. Claude and Gemini count the same text as more tokens than GPT does, so we scale your numbers for each model before pricing them. We took every price from the providers’ own pricing pages; the date we last checked them is shown above.

Why the price per token can mislead you

Because models don’t count text the same way, a lower price per token doesn’t always mean a lower bill. Claude Opus 4.7 and everything after it use a tokenizer that, in our measurements, needs about 1.5 times as many tokens as GPT for the same text. Gemini needs about 1.1 times. Our token counter explains where those ratios come from.

Here’s where it bites. Claude Sonnet 5 costs $2 per million input tokens and $10 per million output. GPT-5.6 Terra costs $2 and $12. On paper Sonnet is cheaper. But price a chatbot with a 2,000-token prompt and a 300-token reply, 1,000 times a day with no caching, and Sonnet comes to $315 a month against Terra’s $228. Priced by the raw token numbers, Sonnet would have looked like $210. The same effect turns Claude Opus 5 from $525 into $788 a month on that workload.

So leave “Adjust for each model’s tokenizer” switched on when your token counts came from a GPT tokenizer, which is what most counting tools use. Switch it off only if you already have each model’s own counts.

Output tokens drive most bills

Every provider charges several times more for output than for input: five times on Claude, GPT-6 Astra, GPT-5.6 Sol and Gemini 3.8 Flash, six on GPT-5.6 Terra, Luna and Gemini 3.1 Pro, and more than eight on Gemini 3.5 Flash-Lite. In our support-chatbot example on Claude Opus 5, the reply is only 13% of the tokens but 62% of the cost.

That makes reply length the first thing to tune. Asking for shorter answers, capping the maximum output, or turning down a model’s reasoning effort usually saves more than switching to a cheaper model. Reasoning counts too: Google’s pricing, for example, lists output as “including thinking tokens”, so a model that thinks for 2,000 tokens before a 300-token answer bills you for 2,300.

How much caching and batch processing save

Caching cuts the price of repeated input by about 90%, and batch processing halves most prices. Cache reads cost a tenth of the normal input price on OpenAI, Anthropic and Gemini, and just 2.5% on Claude Fable 5.1. In our coding-agent example, where 85% of each 40,000-token prompt repeats from step to step, caching brings the monthly bill down by about 60% on every provider: from $5,625 to $2,183 on Claude Opus 5, and from $3,000 to $1,164 on GPT-5.6 Sol.

The calculator assumes your cache stays warm, which holds when requests arrive steadily. Writing to the cache costs a little extra: 1.25 times the input price on OpenAI and on Anthropic’s 5-minute cache, and 2 times for Anthropic’s 1-hour cache. Gemini charges for storage by the hour instead, from $0.50 to $4.50 per million tokens per hour depending on the model.

Batch processing is for work that can wait up to a day, like tagging, translation or overnight reports. It halves input and output prices on all three providers. One exception we found: on Gemini 3.1 Pro, cached input stays at the standard $0.20 in batch mode.

The long-context trap

OpenAI and Google charge more for the whole request once a prompt passes a size limit; Anthropic doesn’t. On GPT-6 Astra and the GPT-5.6 models, prompts over 272,000 input tokens are billed at twice the input rate and 1.5 times the output rate for the full request. Gemini 3.1 Pro switches to higher prices above 200,000 tokens.

The jump is steep. A 250,000-token prompt with a 2,000-token answer costs $2.60 on GPT-6 Astra. Make it 300,000 tokens, 20% more text, and it costs $6.15. Anthropic says its current models bill the full 1M-token window at the standard rate, so the same 300,000-token job costs $2.33 on Claude Opus 5, even after counting Claude’s extra tokens. If you work with very large documents, check where your prompts land against these limits, or split them up.

What the calculator leaves out

It prices tokens only. A few other charges can show up on a real invoice:

  • Web search. Claude charges $10 per 1,000 searches. Gemini includes 5,000 grounded searches a month across its 3.x models, then $14 per 1,000.
  • Tool definitions. Tool names, descriptions and schemas count as input on every request, and Claude adds its own tool-use system prompt on top (286 tokens on Opus 5).
  • Keeping data in one region. US-only processing on Claude costs 1.1 times the normal price. OpenAI adds 10% for regional processing on models released since March 5, 2026.
  • Images, audio and files. These are billed as tokens too, often many of them. Count them with each provider’s tools before you rely on an estimate.

Price changes to watch

Two of the prices here are temporary, so check back before you commit to a budget.

  • Gemini 3.8 Flash costs $0.75 input and $3.75 output through December 31, 2026, then doubles to $1.50 and $7.50.
  • GPT-5.6 Sol is on promotional pricing that OpenAI says runs at least through November 21, 2026. It hasn’t said what the price will be after that.

One price went the other way: Anthropic had announced Claude Sonnet 5’s $2 and $10 rates as introductory, then made them permanent and cancelled the planned rise to $3 and $15.

Frequently asked questions

Which AI API is the cheapest?

GPT-5.6 Luna, in every workload we priced. For our support chatbot example it comes to about $16 a month, against $34 for Gemini 3.5 Flash-Lite and $76 for Claude Haiku 4.5. The cheapest model isn’t always good enough for the job, though. Price two or three models that can actually do your task and compare those.

How much does the Claude API cost per month?

It depends almost entirely on volume. A support chatbot handling 1,000 chats a day comes to roughly $76 a month on Claude Haiku 4.5, $218 on Sonnet 5 and $545 on Opus 5, using our example settings with 60% of the prompt cached. Put your own numbers into the calculator above to see yours.

How do I know how many tokens my requests use?

Paste a typical prompt and a typical reply into our token counter, then enter those numbers here. If you only know the length in words, switch the calculator to words and it converts them for each model. Your API dashboard shows the real counts once you’re live.

Is a ChatGPT or Claude subscription the same as API access?

No. ChatGPT Plus, Claude Pro and similar plans are flat monthly subscriptions for the chat apps. The API is billed separately by the token, from a developer account. A Claude Pro or Max plan doesn’t include API usage. For what the subscriptions give you, see our guide to Claude usage limits.

How often are these prices updated?

We check each provider’s pricing page and update the calculator when prices change. The date above the results shows when we last checked. Each provider’s official page is linked in the price table if you want to double-check before a big decision.

Last updated: September 19, 2026