LLMPrice.io

LLM API pricing reference

Published on-demand rates, USD per million tokens, for 40 current first-party text models. Captured 2026-09-21 from public pricing data and committed to our archive.

Cheapest output rate
Ministral 3 3B
$0.100 per 1M out
Widest input to output gap
Gemini 3.5 Flash Lite
output costs 8x its input
Most recent move
DeepSeek Flash
2026-09-14, input, output and cache
Claude Fable 5.1Anthropic$10.00$50.00$0.250
Claude Opus 5Anthropic$5.00$25.00$0.500
Claude Sonnet 5Anthropic$2.00$10.00$0.2002026-07-06input, output, cache
Claude Haiku 4.5Anthropic$1.00$5.00$0.100
DeepSeek V4 ProDeepSeek$1.32$3.96$0.0442026-08-22input, output, cache
DeepSeek FlashDeepSeek$0.300$1.20$0.0062026-09-14input, output, cache
Gemini 2.5 ProGoogle$1.25$10.00$0.1252026-01-19cache
Gemini omni 1.1 FlashGoogle$1.50$9.00
Gemini 3.8 FlashGoogle$0.750$3.75$0.075
Gemini 3.5 Flash LiteGoogle$0.300$2.50$0.030
Gemini 3.1 Flash LiteGoogle$0.250$1.50$0.025
Gemini 2.5 Flash LiteGoogle$0.100$0.400$0.0102026-01-26cache
Gemini 2.0 Flash Lite 001Google$0.075$0.300$0.019
Mistral Medium 3.5Mistral$1.50$7.50$0.150
Magistral Medium 1.2 2509Mistral$2.00$5.00
Devstral 2512Mistral$0.400$2.00
Mistral Large 2512Mistral$0.500$1.50$0.0502026-08-30cache
Codestral 2508Mistral$0.300$0.900$0.0302026-08-30cache
Mistral vibe cli FastMistral$0.150$0.600$0.015
Devstral Small 2512Mistral$0.100$0.300
Ministral 3 14B 2512Mistral$0.200$0.200$0.0202026-08-30cache
Ministral 3 8B 2512Mistral$0.150$0.150$0.0152026-08-30cache
Ministral 3 3B 2512Mistral$0.100$0.100$0.0102026-08-30cache
GPT 6 astraOpenAI$10.00$50.00$1.00
GPT 5.6OpenAI$4.00$20.00$0.4002026-08-23input, output, cache
GPT 5.3OpenAI$1.75$14.00$0.175
GPT 5.6 TerraOpenAI$2.00$12.00$0.2002026-08-01input, output, cache
GPT 5 ChatOpenAI$1.25$10.00$0.125
GPT 5.4 MiniOpenAI$0.750$4.50$0.075
GPT 5 MiniOpenAI$0.250$2.00$0.025
GPT 5.4 NanoOpenAI$0.200$1.25$0.020
GPT 5.6 LunaOpenAI$0.200$1.20$0.0202026-08-01input, output, cache
GPT 5 NanoOpenAI$0.050$0.400$0.005
Sonar ProPerplexity$3.00$15.00
Sonar Reasoning ProPerplexity$2.00$8.00
Sonar ReasoningPerplexity$1.00$5.00
SonarPerplexity$1.00$1.00
Grok 4.6xAI$2.00$6.00$0.500
Grok 4.3xAI$1.25$2.50$0.200
Grok Code FastxAI$1.00$2.00$0.2002026-08-13input, output, cache

Rates captured 2026-09-21. Estimates only; confirm on each provider's pricing page. Tap any column heading to sort.

Last change is the day we observed a rate move in our own archive, and which rate moved. An em dash means every rate has held since that model entered the archive. Only 14 of the 40 models here have repriced at all, and most of those moved the cache-read rate rather than input or output, so a change in this column is not the same as a price cut. Dates are observation dates and can lag a provider's announcement.

Price your own workloadCost on your workload, not just the rateHow to compare cheapness honestly

How to read this table

Every figure is US dollars per million tokens at the on-demand rate, which is the price you pay with no commitment, no volume agreement and no batch discount. It is the rate almost everyone actually pays, and the one providers quote. A token is roughly four characters of English prose, so a million tokens is somewhere near 750,000 words, or about nine copies of a full-length novel.

Input is what you send: your system prompt, your tools and schemas, any documents you supply, and the whole conversation so far, all of it re-sent on every call. Output is only what the model writes back. Cached input is what a repeated prefix costs on the second and later calls, and it is the column most people skim past and should not, because it is where the largest discounts on this page are: 32 of the 40 models publish a cache-read rate, and the median one reads at 90% below its own standard input rate.

The cheapest rate is not the cheapest bill

Sorting this table by input rate answers a question almost nobody has. Input rates across these 40 models span $0.05 to $10 per million, a factor of about 200, but output is priced separately and usually much higher: the median model here charges 5.0 times more for output than for input, and at the extreme the multiple reaches 8. So the model that wins on the input column can lose badly on a workload that writes a great deal, and the reverse is just as true.

What decides your bill is the mix. A retrieval or classification workload reads enormously and writes a sentence, so its cost is nearly all input and caching is worth more to it than any model switch. A drafting or summarizing workload writes far more than it reads, so output rate dominates and caching barely helps. Two teams paying the same provider can rank these models in opposite orders and both be right. That is why the model pages price four different workload shapes rather than publishing one ranking, and why the estimator asks about your work before it names a number.

What is in this table, and what is not

These are first-party text models from 7 providers, priced at the rate the provider that built the model charges to serve it. Resellers, aggregators and hosting marketplaces are left out, not because they are worse but because their prices move for reasons that have nothing to do with the model and everything to do with the intermediary, which would make a price archive meaningless. Fine-tuned, embedding, image generation and audio models are out of scope here too.

Every rate on this page was read from a provider's own published pricing in a dated capture and committed to the archive that night. Nothing is estimated, interpolated or projected, and when a provider publishes no rate for something, this site says so rather than filling the gap with a plausible figure. That is the whole discipline behind the archive, and it is why a cell can be empty here when other comparison sites show a number.

Rates move less often than the coverage suggests

Only 14 of the 40 models here have repriced at all since entering the archive. Providers overwhelmingly prefer to launch a cheaper new model than to cut the price of one already in service, which is why the way costs fall in practice is that a new name appears at a lower rate and you migrate to it, not that the number next to your current model gets smaller while you sleep. The practical consequence is that watching your own model's price is close to useless, and watching what arrives beside it is where the saving is. That is what the monthly index tracks.

Using one of these figures

Rates on this page were captured 2026-09-21 from each provider's published pricing and committed to our archive, so this reading does not change when a provider edits a page. Quote the date with the number.

LLMPrice.io, “LLM API pricing reference”. Rates captured 2026-09-21. https://llmprice.io/pricing

The same rates as JSON · how the archive works · questions about these numbers