Reference · 62 models

LLM API pricing

What developers pay per token for every current model from the five largest API vendors. Prices come from each vendor's own pricing page, not from resellers. Paying for a chat app instead? See chatbot plans.

By vendor

Open a vendor for cache, batch, long-context and peak-hour rates, plus the billing rules behind them.

VendorModelsInput per 1MCheapest requestChecked
OpenAI29$0.05 – $150.00GPT-5 nano $0.00025 for 1K in + 500 outSep 16, 2026
Anthropic13$1.00 – $10.00Claude Haiku 4.5 $0.0035 for 1K in + 500 outSep 16, 2026
Google Gemini11$0.10 – $2.00Gemini 2.5 Flash-Lite $0.0003 for 1K in + 500 outSep 16, 2026
xAI Grok7$1.00 – $2.00Grok Build 0.1 $0.002 for 1K in + 500 outSep 16, 2026
DeepSeek2$0.15 – $0.66DeepSeek V4.1 Flash $0.00045 for 1K in + 500 outSep 16, 2026

Every model, cheapest first

US dollars per 1M tokens, standard processing. Sorted by the cost of one request with 1,000 input and 500 output tokens. DeepSeek rows are off-peak prices; rows marked "promo price" are temporary: GPT-5.6 Sol's price runs at least through November 21, 2026, and the Gemini Flash launch prices through December 31, 2026.

ModelVendorInputCached inputOutput1K in + 500 out
GPT-5 nanoOpenAI$0.05$0.005$0.40$0.00025
Gemini 2.5 Flash-LiteGoogle Gemini$0.10$0.01$0.40$0.0003
GPT-4.1 nanoOpenAI$0.10$0.025$0.40$0.0003
DeepSeek V4.1 Flashoff-peakDeepSeek$0.15$0.003$0.60$0.00045
GPT-4o miniOpenAI$0.15$0.075$0.60$0.00045
GPT-5.6 LunaOpenAI$0.20$0.02$1.20$0.0008
GPT-5.4 nanoOpenAI$0.20$0.02$1.25$0.00083
Gemini 3.1 Flash-LiteGoogle Gemini$0.25$0.025$1.50$0.001
GPT-4.1 miniOpenAI$0.40$0.10$1.60$0.0012
GPT-5 miniOpenAI$0.25$0.025$2.00$0.0013
Gemini 2.5 FlashGoogle Gemini$0.30$0.03$2.50$0.0016
Gemini 3.5 Flash-LiteGoogle Gemini$0.30$2.50$0.0016
DeepSeek V4 Prooff-peakDeepSeek$0.66$0.022$1.98$0.0017
Gemini 3 Flash PreviewGoogle Gemini$0.50$0.05$3.00$0.002
Grok Build 0.1xAI Grok$1.00$0.20$2.00$0.002
Grok 4.20 multi-agentxAI Grok$1.25$0.20$2.50$0.0025
Grok 4.20 non-reasoningxAI Grok$1.25$0.20$2.50$0.0025
Grok 4.20 reasoningxAI Grok$1.25$0.20$2.50$0.0025
Grok 4.3xAI Grok$1.25$0.20$2.50$0.0025
Gemini 3.6 Flashpromo priceGoogle Gemini$0.75$0.075$3.75$0.0026
Gemini 3.7 Flashpromo priceGoogle Gemini$0.75$3.75$0.0026
Gemini 3.8 Flashpromo priceGoogle Gemini$0.75$0.075$3.75$0.0026
GPT-5.4 miniOpenAI$0.75$0.075$4.50$0.003
o3-miniOpenAI$1.10$0.55$4.40$0.0033
o4-miniOpenAI$1.10$0.275$4.40$0.0033
Claude Haiku 4.5Anthropic$1.00$0.10$5.00$0.0035
Grok 4.5xAI Grok$2.00$0.30$6.00$0.005
Grok 4.6xAI Grok$2.00$0.50$6.00$0.005
Gemini 3.5 FlashGoogle Gemini$1.50$0.15$9.00$0.006
GPT-4.1OpenAI$2.00$0.50$8.00$0.006
o3OpenAI$2.00$0.50$8.00$0.006
Gemini 2.5 ProGoogle Gemini$1.25$0.125$10.00$0.0063
GPT-5OpenAI$1.25$0.125$10.00$0.0063
GPT-5.1OpenAI$1.25$0.125$10.00$0.0063
Claude Sonnet 5Anthropic$2.00$0.20$10.00$0.007
GPT-4oOpenAI$2.50$1.25$10.00$0.0075
Gemini 3.1 Pro PreviewGoogle Gemini$2.00$0.20$12.00$0.008
GPT-5.6 TerraOpenAI$2.00$0.20$12.00$0.008
GPT-5.2OpenAI$1.75$0.175$14.00$0.0088
GPT-5.3 CodexOpenAI$1.75$0.175$14.00$0.0088
GPT-5.4OpenAI$2.50$0.25$15.00$0.01
Claude Sonnet 4.5Anthropic$3.00$0.30$15.00$0.011
Claude Sonnet 4.6Anthropic$3.00$0.30$15.00$0.011
GPT-5.6 Solpromo priceOpenAI$4.00$0.40$20.00$0.014
Claude Opus 4.5Anthropic$5.00$0.50$25.00$0.018
Claude Opus 4.6Anthropic$5.00$0.50$25.00$0.018
Claude Opus 4.7Anthropic$5.00$0.50$25.00$0.018
Claude Opus 4.8Anthropic$5.00$0.50$25.00$0.018
Claude Opus 5Anthropic$5.00$0.50$25.00$0.018
GPT-5.5OpenAI$5.00$0.50$30.00$0.02
Claude Fable 5Anthropic$10.00$1.00$50.00$0.035
Claude Fable 5.1Anthropic$10.00$0.25$50.00$0.035
Claude Mythos 5Anthropic$10.00$1.00$50.00$0.035
Claude Mythos 5.1Anthropic$10.00$0.25$50.00$0.035
GPT-6 AstraOpenAI$10.00$1.00$50.00$0.035
o1OpenAI$15.00$7.50$60.00$0.045
o3-proOpenAI$20.00$80.00$0.06
GPT-5 ProOpenAI$15.00$120.00$0.075
GPT-5.2 ProOpenAI$21.00$168.00$0.11
GPT-5.4 ProOpenAI$30.00$180.00$0.12
GPT-5.5 ProOpenAI$30.00$180.00$0.12
o1-proOpenAI$150.00$600.00$0.45

How to read this table

Input is what you pay for the tokens you send: the system prompt, the conversation so far, retrieved documents and the user's message. Output is what you pay for the tokens the model writes back. Cached input is the discounted rate for prompt text the vendor has already processed and stored, such as a long system prompt reused across requests. All three are list prices in US dollars per 1 million tokens.

Per-token prices are hard to compare in your head, so the last column prices one ordinary request: 1,000 input tokens and 500 output tokens, roughly a short chat turn with some context. The table is sorted by that figure. It ranges from $0.00025 for GPT-5 nano to $0.45 for o1-pro. If your workload is different, for example long documents in and short answers out, the ranking can change, so look at the input and output columns separately.

The price you actually pay can be lower or higher than the list price. Batch processing, caching, peak hours, long prompts and launch promotions all change it, and each vendor page explains its own rules.

Why LLM API prices differ so much

The 62 models above are all priced per token, yet a request can cost 1,800 times more on one model than on another. These are the main reasons.

Model size and tier

Every vendor sells a ladder of models, from small and fast to large and slow. Bigger models need more GPU memory and compute for every token, and the price follows. Inside a single OpenAI generation the gap is wide: GPT-5.4 nano costs $0.20 per 1M input tokens and GPT-5.4 Pro costs $30.00, 150x more. Names like nano, mini, Flash-Lite, Flash, Haiku, Sonnet and Opus mark these tiers.

Output costs more than input

A model reads your whole prompt in one parallel pass, but writes its answer one token at a time. That makes output tokens more expensive to serve. Across this table, output is priced at 2x to 8.3x the input rate, with a median of 5x. Workloads that generate long answers, code or reports are therefore dominated by the output price.

Reasoning and "Pro" models

Reasoning models think before they answer, and those thinking tokens are billed as output even when you never see them. Pro variants spend even more compute per request. That is why the most expensive rows here are reasoning and Pro models, and why a cheap-looking reasoning model can cost more per task than its per-token price suggests.

Caching discounts

Reusing the same prompt prefix lets the vendor skip work, and most pass part of the saving on. Cached input costs between 2% and 50% of the normal input price across the models that offer it. 8 models, mostly Pro variants, list no cached rate at all. Anthropic also charges extra to write to the cache, so caching only pays off when the same prefix is reused often.

Long prompts cost more on some models

15 of the 62 models switch to higher rates once a prompt passes a threshold, 200K or 272K tokens. Above it, the input price doubles and the output price rises 1.5x to 2x. Anthropic's Claude 4.6 and later models bill their full 1M-token context at standard rates instead.

Tokenizers are not the same

A token is not a fixed unit. Each vendor splits text with its own tokenizer, so the same paragraph becomes a different number of tokens on different models. Anthropic says its tokenizer for Claude 4.7 and later produces about 30% more tokens for the same text. Two models with the same per-token price can still produce different bills for identical prompts.

Time of day and promotional pricing

DeepSeek charges double during peak hours: DeepSeek V4 Pro is $0.66 input off-peak and $1.32 at peak. OpenAI's GPT-5.6 Sol costs $4.00 / $20.00 on a promotion that runs at least through November 21, 2026, and Google sells 3 Gemini Flash models at launch prices through December 31, 2026; Gemini 3.8 Flash goes from $0.75 to $1.50 per 1M input tokens in 2027. Prices in this table are what the vendors charge today, not what they will charge next year.

Newer is not always pricier

Vendors often launch new models at lower prices than the ones they replace. OpenAI's o1 still lists $15.00 input and $60.00 output, while the newer GPT-5.4 costs $2.50 and $15.00. Staying on an old model out of habit can be the expensive choice.

Resellers and open-weight models

Open-weight models can be hosted by anyone, so the same model shows up at different prices on different platforms. For DeepSeek V4 Pro, DeepSeek charges $0.66 input and $1.98 output off-peak, while OpenRouter lists $0.58 and $1.74. Cloud platforms and routers also add their own markups, discounts and regional pricing. This page uses the vendor's own price, because that is the reference the others are measured against.

What a month of traffic costs

100,000 requests per month, each with 1,000 input and 500 output tokens. The last column assumes 800 of the 1,000 input tokens hit the cache, and ignores cache-write charges.

ModelVendorPer monthWith 80% cached input
GPT-5.6 SolOpenAI$1,400$1,112
Claude Sonnet 5Anthropic$700$556
Gemini 3.1 Pro PreviewGoogle Gemini$800$656
Grok 4.6xAI Grok$500$380
DeepSeek V4 ProDeepSeek$165$114

For your own numbers, use the LLM API cost calculator. To estimate by hand, multiply your average input tokens by the input price and your average output tokens by the output price, divide each by 1,000,000, add them, and multiply by your monthly request count. Then check the vendor page for batch, caching and long-context rules that apply to your traffic.

Questions

What is the cheapest LLM API?

On list prices, GPT-5 nano from OpenAI is the cheapest model we track: $0.05 per 1M input tokens and $0.40 per 1M output tokens, or about $0.00025 for a request with 1,000 input and 500 output tokens. The cheapest model is rarely the right one for hard tasks, so compare a few tiers on your own prompts before committing.

What is the most expensive LLM API?

o1-pro from OpenAI costs $150.00 input and $600.00 output per 1M tokens, about $0.45 per 1,000-in/500-out request. That is roughly 1,800 times the cheapest model on this page.

Why are output tokens more expensive than input tokens?

Input tokens are processed in one parallel pass, while output tokens are generated one at a time, and each new token needs another pass through the model. Across the 62 models here, output costs between 2x and 8.3x the input price, with a median of 5x.

Are reasoning tokens billed?

Yes. OpenAI, Anthropic and Google bill the tokens a model spends thinking before it answers as output tokens, even when the thinking itself is not returned. A reasoning model with a modest output price can cost more per task than a pricier model that answers directly.

How many words are in 1 million tokens?

It depends on the tokenizer and the language. OpenAI's rule of thumb for English is about four characters, or three quarters of a word, per token, so 1M tokens is roughly 750,000 words. Anthropic says its tokenizer for Claude 4.7 and later models produces about 30% more tokens for the same text, and non-English text usually needs more tokens than English.

How we collect these prices

  • Every price is copied by hand from the vendor's own pricing page, which is what you are actually billed. Each vendor page links to its source and shows the date we checked it.
  • Resellers such as OpenRouter often list different rates. We use them only to notice that a price may have changed, then re-check the vendor. Every confirmed change goes into the LLM changelog.
  • The per-request figure is our arithmetic on list prices. Token counts for the same text differ between vendors' tokenizers, so the same prompt does not produce the same number of tokens everywhere.