Reference · 62 models
LLM API pricing
What developers pay per token for every current model from the five largest API vendors. Prices come from each vendor's own pricing page, not from resellers. Paying for a chat app instead? See chatbot plans.
By vendor
Open a vendor for cache, batch, long-context and peak-hour rates, plus the billing rules behind them.
| Vendor | Models | Input per 1M | Cheapest request | Checked |
|---|---|---|---|---|
| OpenAI | 29 | $0.05 – $150.00 | GPT-5 nano $0.00025 for 1K in + 500 out | Sep 16, 2026 |
| Anthropic | 13 | $1.00 – $10.00 | Claude Haiku 4.5 $0.0035 for 1K in + 500 out | Sep 16, 2026 |
| Google Gemini | 11 | $0.10 – $2.00 | Gemini 2.5 Flash-Lite $0.0003 for 1K in + 500 out | Sep 16, 2026 |
| xAI Grok | 7 | $1.00 – $2.00 | Grok Build 0.1 $0.002 for 1K in + 500 out | Sep 16, 2026 |
| DeepSeek | 2 | $0.15 – $0.66 | DeepSeek V4.1 Flash $0.00045 for 1K in + 500 out | Sep 16, 2026 |
Every model, cheapest first
US dollars per 1M tokens, standard processing. Sorted by the cost of one request with 1,000 input and 500 output tokens. DeepSeek rows are off-peak prices; rows marked "promo price" are temporary: GPT-5.6 Sol's price runs at least through November 21, 2026, and the Gemini Flash launch prices through December 31, 2026.
| Model | Vendor | Input | Cached input | Output | 1K in + 500 out |
|---|---|---|---|---|---|
| GPT-5 nano | OpenAI | $0.05 | $0.005 | $0.40 | $0.00025 |
| Gemini 2.5 Flash-Lite | Google Gemini | $0.10 | $0.01 | $0.40 | $0.0003 |
| GPT-4.1 nano | OpenAI | $0.10 | $0.025 | $0.40 | $0.0003 |
| DeepSeek V4.1 Flashoff-peak | DeepSeek | $0.15 | $0.003 | $0.60 | $0.00045 |
| GPT-4o mini | OpenAI | $0.15 | $0.075 | $0.60 | $0.00045 |
| GPT-5.6 Luna | OpenAI | $0.20 | $0.02 | $1.20 | $0.0008 |
| GPT-5.4 nano | OpenAI | $0.20 | $0.02 | $1.25 | $0.00083 |
| Gemini 3.1 Flash-Lite | Google Gemini | $0.25 | $0.025 | $1.50 | $0.001 |
| GPT-4.1 mini | OpenAI | $0.40 | $0.10 | $1.60 | $0.0012 |
| GPT-5 mini | OpenAI | $0.25 | $0.025 | $2.00 | $0.0013 |
| Gemini 2.5 Flash | Google Gemini | $0.30 | $0.03 | $2.50 | $0.0016 |
| Gemini 3.5 Flash-Lite | Google Gemini | $0.30 | — | $2.50 | $0.0016 |
| DeepSeek V4 Prooff-peak | DeepSeek | $0.66 | $0.022 | $1.98 | $0.0017 |
| Gemini 3 Flash Preview | Google Gemini | $0.50 | $0.05 | $3.00 | $0.002 |
| Grok Build 0.1 | xAI Grok | $1.00 | $0.20 | $2.00 | $0.002 |
| Grok 4.20 multi-agent | xAI Grok | $1.25 | $0.20 | $2.50 | $0.0025 |
| Grok 4.20 non-reasoning | xAI Grok | $1.25 | $0.20 | $2.50 | $0.0025 |
| Grok 4.20 reasoning | xAI Grok | $1.25 | $0.20 | $2.50 | $0.0025 |
| Grok 4.3 | xAI Grok | $1.25 | $0.20 | $2.50 | $0.0025 |
| Gemini 3.6 Flashpromo price | Google Gemini | $0.75 | $0.075 | $3.75 | $0.0026 |
| Gemini 3.7 Flashpromo price | Google Gemini | $0.75 | — | $3.75 | $0.0026 |
| Gemini 3.8 Flashpromo price | Google Gemini | $0.75 | $0.075 | $3.75 | $0.0026 |
| GPT-5.4 mini | OpenAI | $0.75 | $0.075 | $4.50 | $0.003 |
| o3-mini | OpenAI | $1.10 | $0.55 | $4.40 | $0.0033 |
| o4-mini | OpenAI | $1.10 | $0.275 | $4.40 | $0.0033 |
| Claude Haiku 4.5 | Anthropic | $1.00 | $0.10 | $5.00 | $0.0035 |
| Grok 4.5 | xAI Grok | $2.00 | $0.30 | $6.00 | $0.005 |
| Grok 4.6 | xAI Grok | $2.00 | $0.50 | $6.00 | $0.005 |
| Gemini 3.5 Flash | Google Gemini | $1.50 | $0.15 | $9.00 | $0.006 |
| GPT-4.1 | OpenAI | $2.00 | $0.50 | $8.00 | $0.006 |
| o3 | OpenAI | $2.00 | $0.50 | $8.00 | $0.006 |
| Gemini 2.5 Pro | Google Gemini | $1.25 | $0.125 | $10.00 | $0.0063 |
| GPT-5 | OpenAI | $1.25 | $0.125 | $10.00 | $0.0063 |
| GPT-5.1 | OpenAI | $1.25 | $0.125 | $10.00 | $0.0063 |
| Claude Sonnet 5 | Anthropic | $2.00 | $0.20 | $10.00 | $0.007 |
| GPT-4o | OpenAI | $2.50 | $1.25 | $10.00 | $0.0075 |
| Gemini 3.1 Pro Preview | Google Gemini | $2.00 | $0.20 | $12.00 | $0.008 |
| GPT-5.6 Terra | OpenAI | $2.00 | $0.20 | $12.00 | $0.008 |
| GPT-5.2 | OpenAI | $1.75 | $0.175 | $14.00 | $0.0088 |
| GPT-5.3 Codex | OpenAI | $1.75 | $0.175 | $14.00 | $0.0088 |
| GPT-5.4 | OpenAI | $2.50 | $0.25 | $15.00 | $0.01 |
| Claude Sonnet 4.5 | Anthropic | $3.00 | $0.30 | $15.00 | $0.011 |
| Claude Sonnet 4.6 | Anthropic | $3.00 | $0.30 | $15.00 | $0.011 |
| GPT-5.6 Solpromo price | OpenAI | $4.00 | $0.40 | $20.00 | $0.014 |
| Claude Opus 4.5 | Anthropic | $5.00 | $0.50 | $25.00 | $0.018 |
| Claude Opus 4.6 | Anthropic | $5.00 | $0.50 | $25.00 | $0.018 |
| Claude Opus 4.7 | Anthropic | $5.00 | $0.50 | $25.00 | $0.018 |
| Claude Opus 4.8 | Anthropic | $5.00 | $0.50 | $25.00 | $0.018 |
| Claude Opus 5 | Anthropic | $5.00 | $0.50 | $25.00 | $0.018 |
| GPT-5.5 | OpenAI | $5.00 | $0.50 | $30.00 | $0.02 |
| Claude Fable 5 | Anthropic | $10.00 | $1.00 | $50.00 | $0.035 |
| Claude Fable 5.1 | Anthropic | $10.00 | $0.25 | $50.00 | $0.035 |
| Claude Mythos 5 | Anthropic | $10.00 | $1.00 | $50.00 | $0.035 |
| Claude Mythos 5.1 | Anthropic | $10.00 | $0.25 | $50.00 | $0.035 |
| GPT-6 Astra | OpenAI | $10.00 | $1.00 | $50.00 | $0.035 |
| o1 | OpenAI | $15.00 | $7.50 | $60.00 | $0.045 |
| o3-pro | OpenAI | $20.00 | — | $80.00 | $0.06 |
| GPT-5 Pro | OpenAI | $15.00 | — | $120.00 | $0.075 |
| GPT-5.2 Pro | OpenAI | $21.00 | — | $168.00 | $0.11 |
| GPT-5.4 Pro | OpenAI | $30.00 | — | $180.00 | $0.12 |
| GPT-5.5 Pro | OpenAI | $30.00 | — | $180.00 | $0.12 |
| o1-pro | OpenAI | $150.00 | — | $600.00 | $0.45 |
How to read this table
Input is what you pay for the tokens you send: the system prompt, the conversation so far, retrieved documents and the user's message. Output is what you pay for the tokens the model writes back. Cached input is the discounted rate for prompt text the vendor has already processed and stored, such as a long system prompt reused across requests. All three are list prices in US dollars per 1 million tokens.
Per-token prices are hard to compare in your head, so the last column prices one ordinary request: 1,000 input tokens and 500 output tokens, roughly a short chat turn with some context. The table is sorted by that figure. It ranges from $0.00025 for GPT-5 nano to $0.45 for o1-pro. If your workload is different, for example long documents in and short answers out, the ranking can change, so look at the input and output columns separately.
The price you actually pay can be lower or higher than the list price. Batch processing, caching, peak hours, long prompts and launch promotions all change it, and each vendor page explains its own rules.
Why LLM API prices differ so much
The 62 models above are all priced per token, yet a request can cost 1,800 times more on one model than on another. These are the main reasons.
Model size and tier
Every vendor sells a ladder of models, from small and fast to large and slow. Bigger models need more GPU memory and compute for every token, and the price follows. Inside a single OpenAI generation the gap is wide: GPT-5.4 nano costs $0.20 per 1M input tokens and GPT-5.4 Pro costs $30.00, 150x more. Names like nano, mini, Flash-Lite, Flash, Haiku, Sonnet and Opus mark these tiers.
Output costs more than input
A model reads your whole prompt in one parallel pass, but writes its answer one token at a time. That makes output tokens more expensive to serve. Across this table, output is priced at 2x to 8.3x the input rate, with a median of 5x. Workloads that generate long answers, code or reports are therefore dominated by the output price.
Reasoning and "Pro" models
Reasoning models think before they answer, and those thinking tokens are billed as output even when you never see them. Pro variants spend even more compute per request. That is why the most expensive rows here are reasoning and Pro models, and why a cheap-looking reasoning model can cost more per task than its per-token price suggests.
Caching discounts
Reusing the same prompt prefix lets the vendor skip work, and most pass part of the saving on. Cached input costs between 2% and 50% of the normal input price across the models that offer it. 8 models, mostly Pro variants, list no cached rate at all. Anthropic also charges extra to write to the cache, so caching only pays off when the same prefix is reused often.
Long prompts cost more on some models
15 of the 62 models switch to higher rates once a prompt passes a threshold, 200K or 272K tokens. Above it, the input price doubles and the output price rises 1.5x to 2x. Anthropic's Claude 4.6 and later models bill their full 1M-token context at standard rates instead.
Tokenizers are not the same
A token is not a fixed unit. Each vendor splits text with its own tokenizer, so the same paragraph becomes a different number of tokens on different models. Anthropic says its tokenizer for Claude 4.7 and later produces about 30% more tokens for the same text. Two models with the same per-token price can still produce different bills for identical prompts.
Time of day and promotional pricing
DeepSeek charges double during peak hours: DeepSeek V4 Pro is $0.66 input off-peak and $1.32 at peak. OpenAI's GPT-5.6 Sol costs $4.00 / $20.00 on a promotion that runs at least through November 21, 2026, and Google sells 3 Gemini Flash models at launch prices through December 31, 2026; Gemini 3.8 Flash goes from $0.75 to $1.50 per 1M input tokens in 2027. Prices in this table are what the vendors charge today, not what they will charge next year.
Newer is not always pricier
Vendors often launch new models at lower prices than the ones they replace. OpenAI's o1 still lists $15.00 input and $60.00 output, while the newer GPT-5.4 costs $2.50 and $15.00. Staying on an old model out of habit can be the expensive choice.
Resellers and open-weight models
Open-weight models can be hosted by anyone, so the same model shows up at different prices on different platforms. For DeepSeek V4 Pro, DeepSeek charges $0.66 input and $1.98 output off-peak, while OpenRouter lists $0.58 and $1.74. Cloud platforms and routers also add their own markups, discounts and regional pricing. This page uses the vendor's own price, because that is the reference the others are measured against.
What a month of traffic costs
100,000 requests per month, each with 1,000 input and 500 output tokens. The last column assumes 800 of the 1,000 input tokens hit the cache, and ignores cache-write charges.
| Model | Vendor | Per month | With 80% cached input |
|---|---|---|---|
| GPT-5.6 Sol | OpenAI | $1,400 | $1,112 |
| Claude Sonnet 5 | Anthropic | $700 | $556 |
| Gemini 3.1 Pro Preview | Google Gemini | $800 | $656 |
| Grok 4.6 | xAI Grok | $500 | $380 |
| DeepSeek V4 Pro | DeepSeek | $165 | $114 |
For your own numbers, use the LLM API cost calculator. To estimate by hand, multiply your average input tokens by the input price and your average output tokens by the output price, divide each by 1,000,000, add them, and multiply by your monthly request count. Then check the vendor page for batch, caching and long-context rules that apply to your traffic.
Questions
What is the cheapest LLM API?
On list prices, GPT-5 nano from OpenAI is the cheapest model we track: $0.05 per 1M input tokens and $0.40 per 1M output tokens, or about $0.00025 for a request with 1,000 input and 500 output tokens. The cheapest model is rarely the right one for hard tasks, so compare a few tiers on your own prompts before committing.
What is the most expensive LLM API?
o1-pro from OpenAI costs $150.00 input and $600.00 output per 1M tokens, about $0.45 per 1,000-in/500-out request. That is roughly 1,800 times the cheapest model on this page.
Why are output tokens more expensive than input tokens?
Input tokens are processed in one parallel pass, while output tokens are generated one at a time, and each new token needs another pass through the model. Across the 62 models here, output costs between 2x and 8.3x the input price, with a median of 5x.
Are reasoning tokens billed?
Yes. OpenAI, Anthropic and Google bill the tokens a model spends thinking before it answers as output tokens, even when the thinking itself is not returned. A reasoning model with a modest output price can cost more per task than a pricier model that answers directly.
How many words are in 1 million tokens?
It depends on the tokenizer and the language. OpenAI's rule of thumb for English is about four characters, or three quarters of a word, per token, so 1M tokens is roughly 750,000 words. Anthropic says its tokenizer for Claude 4.7 and later models produces about 30% more tokens for the same text, and non-English text usually needs more tokens than English.
How we collect these prices
- Every price is copied by hand from the vendor's own pricing page, which is what you are actually billed. Each vendor page links to its source and shows the date we checked it.
- Resellers such as OpenRouter often list different rates. We use them only to notice that a price may have changed, then re-check the vendor. Every confirmed change goes into the LLM changelog.
- The per-request figure is our arithmetic on list prices. Token counts for the same text differ between vendors' tokenizers, so the same prompt does not produce the same number of tokens everywhere.