Calculate Your LLM API Cost
Use this page to compare current LLM API prices across providers. When you want to turn those rates into a real workload estimate, open our LLM Cost Calculator and enter your prompt tokens, output tokens, cache usage, and request volume.
- Estimate cost per request, per 1,000 requests, per day, and per month
- Compare GPT, Claude, Gemini, Grok, Kimi, DeepSeek, and Qwen models side by side
- Model cache usage, context-based pricing tiers, and supported pricing modes
Last pricing review: 21 August 2026.
Introduction
LLM API pricing has become harder to compare than a single “price per token” suggests. Providers now separate input and output rates, discount cached prompts, introduce long-context pricing thresholds, and offer Batch, Flex, Priority, or other processing tiers.
This LLM pricing comparison keeps the first question simple: what does each model currently charge for API tokens?
The table below compares current standard text-token rates across leading GPT, Claude, Gemini, Grok, Kimi, DeepSeek, and Qwen models. Prices are shown per 1 million tokens so you can compare vendors on the same unit.
For actual application budgeting, token price is only the starting point. The amount of context you send, output length, caching, retries, tool calls, and the number of requests required to complete a task can all change the real cost.
Table of Contents
1. LLM API Pricing Comparison at a Glance
The table below compares standard paid API pricing per 1M text tokens. “Cached” refers to the published cache-read or cache-hit rate where available.
For models with context-dependent pricing, the relevant prompt tier is shown in the model name. Temporary promotions, enterprise contracts, media generation, search tools, and other usage-based fees are excluded from this compact comparison unless specifically noted.
These headline rates are useful for comparing vendors, but they do not tell you what a complete application request will cost. A model with a lower token rate can still cost more if it generates longer outputs, requires more retries, or uses more calls to complete the same task.
| Provider | Model | Pricing (per 1M tokens) |
|---|---|---|
| OpenAI | GPT-5.6 Sol (≤272K prompt) | In: $5.00 · Cached: $0.50 · Out: $30.00 |
| OpenAI | GPT-5.6 Terra (≤272K prompt) | In: $2.00 · Cached: $0.20 · Out: $12.00 |
| OpenAI | GPT-5.6 Luna (≤272K prompt) | In: $0.20 · Cached: $0.02 · Out: $1.20 |
| OpenAI | GPT-5.1 | In: $1.25 · Cached: $0.125 · Out: $10.00 |
| OpenAI | GPT-4.1 | In: $2.00 · Cached: $0.50 · Out: $8.00 |
| Anthropic | Claude Fable 5 | In: $10.00 · Cached: $1.00 · Out: $50.00 |
| Anthropic | Claude Opus 5 | In: $5.00 · Cached: $0.50 · Out: $25.00 |
| Anthropic | Claude Sonnet 5 | In: $2.00 · Cached: $0.20 · Out: $10.00 |
| Anthropic | Claude Haiku 4.5 | In: $1.00 · Cached: $0.10 · Out: $5.00 |
| Gemini 3.7 Flash (promo through Dec 31, 2026) | In: $0.75 · Cached: $0.075 · Out: $3.75 | |
| Gemini 3.6 Flash (promo through Dec 31, 2026) | In: $0.75 · Cached: $0.075 · Out: $3.75 | |
| Gemini 3.5 Flash-Lite | In: $0.30 · Cached: $0.03 · Out: $2.50 | |
| Gemini 3.1 Pro Preview (≤200K prompt) | In: $2.00 · Cached: $0.20 · Out: $12.00 | |
| xAI | Grok 4.6 (<200K prompt) | In: $2.00 · Cached: $0.50 · Out: $6.00 |
| xAI | Grok 4.6 (≥200K prompt) | In: $4.00 · Cached: $1.00 · Out: $12.00 |
| xAI | Grok 4.5 (<200K prompt) | In: $2.00 · Cached: $0.30 · Out: $6.00 |
| xAI | Grok 4.5 (≥200K prompt) | In: $4.00 · Cached: $0.60 · Out: $12.00 |
| xAI | Grok 4.3 (<200K prompt) | In: $1.25 · Cached: $0.20 · Out: $2.50 |
| Kimi | Kimi K3 | In: $3.00 · Cached: $0.30 · Out: $15.00 |
| DeepSeek | DeepSeek V4 Pro | In: $0.435 · Cached: $0.003625 · Out: $0.87 |
| DeepSeek | DeepSeek V4 Flash | In: $0.14 · Cached: $0.0028 · Out: $0.28 |
| Alibaba / Qwen | Qwen3.8-Max (Global, ≤1M) | In: $1.65 · Cached*: $0.137 · Out: $4.951 |
| Alibaba / Qwen | Qwen3.7-Plus (Global, non-thinking, ≤256K) | In: $0.276 · Cached*: $0.028 · Out: $1.101 |
Standard paid text-token rates are shown unless noted. “Cached” means cache-read/cache-hit pricing, not cache creation or storage. GPT-5.6 rows show pricing for prompts up to 272K tokens; requests above 272K are billed at 2× the standard input rate and 1.5× the standard output rate for the full request. Claude Sonnet 5’s $2 input / $10 output rate is now its standard price; the previously planned September 1 increase will not occur. Gemini 3.7 Flash and Gemini 3.6 Flash use introductory Standard pricing through December 31, 2026. Grok 4.6 and Grok 4.5 long-context pricing is shown separately. Qwen rows use Global list prices and explicit-cache read pricing; temporary promotional discounts are excluded. Batch, Flex, Priority/Fast, regional, search/tool, media, cache-creation, and cache-storage fees are not included in this compact table.
2. Cheapest LLM API Pricing in 2026
If you are comparing models purely by published token price, the lowest-cost options in the current table are concentrated among DeepSeek, Qwen, and Google’s efficiency-focused models.
DeepSeek V4 Flash currently has one of the lowest standard direct API rates in this comparison at $0.14 per 1M cache-miss input tokens and $0.28 per 1M output tokens. DeepSeek V4 Pro remains inexpensive at $0.435 input and $0.87 output per 1M tokens.
For Alibaba’s Global deployment scope, Qwen3.7-Plus starts at $0.276 per 1M input tokens for prompts up to 256K in non-thinking mode, while Google prices Gemini 3.5 Flash-Lite at $0.30 input and $2.50 output per 1M tokens under Standard processing.
However, the cheapest LLM API is not automatically the cheapest model for your application. Quality, output length, latency, retries, tool calls, and the number of model requests required to finish a task can outweigh a small difference in token price.
Use the published rates to shortlist models, then compare the candidates using the same workload and quality criteria.
3. LLM Pricing Updates: What Changed in 2026
This page is updated as model lineups and API prices change. The latest review was completed on 21 August 2026.
Current additions and major changes include:
- OpenAI GPT-5.6: Current Standard API rates below the long-context threshold are $5/$30 per 1M input/output tokens for Sol, $2/$12 for Terra, and $0.20/$1.20 for Luna. Requests above 272K input tokens use higher long-context rates.
- Anthropic: Claude Sonnet 5 remains $2 per 1M input tokens and $10 per 1M output tokens. Anthropic has made this the standard price, and the previously planned September 1 increase to $3/$15 will not occur.
- Google: Gemini 3.7 Flash is now generally available with introductory Standard pricing of $0.75 input, $0.075 cached input, and $3.75 output per 1M tokens through December 31, 2026.
- xAI: Grok 4.6 is now the current flagship API model, with a 500K context window. Pricing is $2/$0.50/$6 per 1M input/cached/output tokens below 200K prompt tokens and $4/$1/$12 at 200K or above.
- Kimi: Kimi K3 brings a 1M-token context window with separate cache-hit, cache-miss, and output pricing.
- DeepSeek: DeepSeek V4 Pro and V4 Flash replaced the older
deepseek-chatanddeepseek-reasonerAPI naming scheme. - Qwen: Qwen3.7 Max and Qwen3.7 Plus are part of Alibaba Model Studio’s current generation, with deployment-scope and context-dependent pricing.
This section should be updated whenever you refresh the pricing table. That gives visitors a quick explanation of what actually changed instead of silently replacing numbers.
4. How to Compare LLM API Pricing Fairly
A useful API price comparison starts by putting every provider on the same unit.
Most providers quote text-token pricing per 1 million tokens, with separate rates for input and output. Cached input may receive another rate.
For a simple request:
Input cost = input tokens ÷ 1,000,000 × input price
Output cost = output tokens ÷ 1,000,000 × output price
Total request cost = input cost + output cost
If cached input is involved, calculate those tokens separately using the provider’s cache-hit rate.
The important part is consistency. When comparing two models, use the same expected prompt size, output size, cache assumptions, and request volume. Otherwise you are comparing different workloads rather than different prices.
5. Input and output token pricing
LLM APIs usually charge input and output tokens at different rates.
Input tokens include the prompt, system instructions, conversation history, retrieved context, and other text sent to the model.
Output tokens are the tokens generated by the model, including billable reasoning tokens where the provider includes them in output usage.
Output pricing is often higher than input pricing, so applications that generate long responses can have a very different cost profile from applications that process large prompts and return short answers.
This is why comparing only the advertised input price can give a misleading picture of total API cost.
6. From Token Price to Real LLM API Cost
Token pricing answers one question:
How much does the provider charge for tokens?
Your application budget answers another:
How much will my workload actually cost?
The difference matters because two models with similar API prices can produce very different bills once real usage is included.
Your total LLM API cost can depend on:
- input tokens per request;
- output tokens per request;
- cache-hit and cache-write behavior;
- requests per user or task;
- API calls per day;
- retries and failed requests;
- agent loops and tool calls;
- long-context pricing thresholds;
- Standard, Batch, Flex, or Priority processing.
Use the pricing comparison on this page to identify models worth evaluating. Then use the LLM Cost Calculator to apply the same workload to each candidate and compare cost per request, per 1,000 requests, per day, and per month.
7. Compare LLM Price, Quality, and Latency Before Switching Vendors
A lower API price alone is rarely enough to justify switching models or vendors.
For a meaningful side-by-side comparison, test the same representative workloads against each candidate and track both cost and performance.
| Metric | What it tells you |
|---|---|
| Quality or task success rate | Whether the model actually solves the workload |
| Cost per request | Direct API cost under the same workload |
| Cost per successful task | Whether lower token pricing survives retries and failures |
| Median and p95 latency | Typical and worst-case response time |
| Tokens per completed task | How efficiently the model uses input and output tokens |
| Retries and tool calls | Hidden sources of additional API spend |
| Cache-hit rate | Whether repeated context produces meaningful savings |
| Context requirements | Whether requests cross higher-priced context tiers |
For vendor decisions, cost per successful task can be more useful than cost per token.
A model may charge less per million tokens but still cost more in production if it requires additional retries, longer outputs, or more tool calls to complete the same job.
This framework is particularly useful when engineering, product, procurement, or finance teams need evidence to justify switching from one LLM provider to another.
8. How to Reduce LLM API Costs
- Right-size the model. Use lower-cost models for classification, extraction, routing, summarization, and other tasks that do not require frontier-level reasoning.
- Trim unnecessary context. Repeated instructions, old conversation turns, and oversized retrieval results all add billable input tokens.
- Control output length. Output tokens are often more expensive than input tokens, so unnecessary verbosity can materially increase cost.
- Use prompt caching deliberately. Stable system prompts, instructions, and repeated context can become much cheaper when cache-hit pricing is supported.
- Consider Batch or Flex processing. Workloads that do not require immediate responses may qualify for cheaper processing tiers.
- Watch long-context thresholds. Crossing a provider’s context threshold can increase the token price for a request.
- Limit agent loops and retries. A single user request can result in many paid API calls when an agent repeatedly reasons, retries, or invokes tools.
- Measure production usage. Track token consumption, latency, retries, cache hits, and cost by feature instead of relying only on pre-launch estimates.
9. LLM Pricing Comparison: 2025 vs 2026
This page was originally published in 2025 and is continuously updated as providers release new models and change their API pricing.
If you arrived here through a search for LLM API pricing comparison 2025, cheapest LLM API pricing 2025, or an older provider-specific pricing query, the main table now shows the current 2026 rates, not a frozen 2025 snapshot.
We retain selected older models, such as GPT-5.1 and GPT-4.1, when their pricing remains useful for existing production deployments, migration planning, or historical comparison.
For current budgeting, always use the latest pricing-review date shown near the top of this page.
Looking for October 2025 LLM API pricing?
Some readers still reach this page through searches for October 2025 OpenAI, Grok, and broader LLM API pricing.
Those searches refer to an earlier version of the market. The current comparison intentionally prioritizes today’s available models and rates rather than presenting expired 2025 pricing as current.
If you are researching a historical price change, use the model name and date together. For current budgeting or vendor selection, use the latest rates in the comparison table above.
10. How We Verify LLM API Pricing
The pricing data on this page is maintained using first-party provider documentation rather than relying on third-party price aggregators.
Current sources include:
- OpenAI API pricing and model documentation
- Anthropic / Claude API pricing
- Google Gemini Developer API pricing
- xAI developer pricing
- Kimi API pricing
- DeepSeek Models & Pricing
- Alibaba Cloud Model Studio pricing and context-cache documentation
During each pricing refresh, we check the current model name, standard input price, cached-input rate where available, output price, context-dependent pricing, and relevant processing-tier changes.
Temporary promotional discounts are excluded from the main comparison when a stable list price is available. Regional pricing is identified when the provider publishes materially different rates by deployment scope.
Last full pricing review: 17 August 2026.
For large production budgets, always confirm the selected model against the linked provider documentation before committing spend.
11. Calculate the Cost of Your Own LLM Workload
The pricing table tells you what each provider charges per token. Your actual application budget depends on how many of those tokens your workload consumes.
Use the LLM Cost Calculator when you want to:
- enter prompt and output tokens;
- model cache usage;
- estimate cost per request;
- estimate cost per 1,000 requests;
- project daily and monthly API spend;
- compare several models using the same workload;
- account for supported processing modes and context-based pricing tiers.
Use this page for LLM pricing comparison and the calculator for workload-specific cost estimation.
What is the cheapest LLM API?
There is no single cheapest LLM API for every workload because input pricing, output pricing, caching, quality, and model behavior all affect total cost.
The lowest published token rate can be useful for creating a shortlist, but production decisions should also consider cost per successful task, latency, retries, output length, and tool usage.
How do I compare LLM API pricing?
Start by comparing input, cached-input, and output prices using the same unit, usually per 1 million tokens.
Then apply the same prompt size, output length, cache assumptions, and request volume to each model. If a provider uses a different context tier or processing mode, account for that before comparing total costs.
How is LLM cost per token calculated?
If a provider charges $2 per 1M input tokens:
$2 ÷ 1,000,000 = $0.000002 per input token
For a complete API request, calculate input and output costs separately because providers generally charge different rates for each.
Is LLM price per token enough to choose a model?
No. Token price tells you the billing rate, but it does not measure task quality, latency, output efficiency, retries, or the number of API calls needed to complete the job.
A stronger comparison considers quality, latency, total token usage, retries, tool calls, and cost per successful task alongside the published API price.
Can I compare LLM pricing and quality side by side?
Yes. Apply the same workload to each candidate model, then evaluate each model on the same representative tasks.
Track quality or task success, API cost, latency, token consumption, retries, and tool calls. This produces a much more useful vendor comparison than ranking models by price per million tokens alone.
Does prompt caching reduce LLM API costs?
It can. Providers that support prompt caching often charge a lower rate when previously processed context is reused.
The economics differ by vendor because cache reads, cache creation, cache storage, and minimum cache requirements can be priced differently.
Why do LLM API prices change?
Providers regularly release new models, retire older API names, change token rates, introduce processing tiers, and update cache or context pricing.
That is why this comparison includes a visible pricing-review date and relies on first-party provider documentation.
Are the 2025 LLM API prices on this page still current?
No. This page was originally published in 2025 but is maintained as a current pricing comparison.
The main table now shows the latest reviewed 2026 rates. Older models are retained only when they remain useful for existing deployments, migration planning, or historical comparison.

Nice