LLMPrice.io

LLM API pricing, ranked by cost

Every current first-party text model, priced on the workload you actually run. Cheapest depends on the shape of the work: a model that wins on a retrieval workload can lose badly on a content one. Rates captured 2026-08-10 from public pricing data and committed to our archive.

Workload profile

800k input, 200k output, 70% of input served from cache. Long context read repeatedly.

#
1Mistral Small 3.2 2506Mistral$0.060$0.180$0.084
2Gemini 2.0 Flash Lite 001Google$0.075$0.300$0.019$0.088
3DeepSeek V4 FlashDeepSeek$0.140$0.280$0.003$0.091
4GPT 5 NanoOpenAI$0.050$0.400$0.005$0.095
5Ministral 3 3B 2512Mistral$0.100$0.100$0.100
6Gemini 2.5 Flash LiteGoogle$0.100$0.400$0.010$0.110
7Devstral Small 2512Mistral$0.100$0.300$0.140
8Ministral 3 8B 2512Mistral$0.150$0.150$0.150
9Grok 4.1 FastxAI$0.200$0.500$0.050$0.176
10Ministral 3 14B 2512Mistral$0.200$0.200$0.200
11Grok 3 MinixAI$0.300$0.500$0.075$0.214
12DeepSeek V4 ProDeepSeek$0.435$0.870$0.004$0.280
13GPT 5.6 LunaOpenAI$0.200$1.20$0.020$0.299
14GPT 5.4 NanoOpenAI$0.200$1.25$0.020$0.309
15Grok Code FastxAI$0.200$1.50$0.020$0.359
16Gemini 3.1 Flash LiteGoogle$0.250$1.50$0.025$0.374
17Codestral 2508Mistral$0.300$0.900$0.420
18GPT 5 MiniOpenAI$0.250$2.00$0.025$0.474
19Gemini 3.5 Flash LiteGoogle$0.300$2.50$0.030$0.589
20Mistral Large 2512Mistral$0.500$1.50$0.700
21Mistral Medium 2508Mistral$0.400$2.00$0.720
22Grok 4.3xAI$1.25$2.50$0.200$0.912
23SonarPerplexity$1.00$1.00$1.00
24Grok 3 Mini FastxAI$0.600$4.00$0.150$1.03
25GPT 5.4 MiniOpenAI$0.750$4.50$0.075$1.12
26Claude Haiku 4.5Anthropic$1.00$5.00$0.100$1.30
27Sonar ReasoningPerplexity$1.00$5.00$1.80
28Gemini 3.6 FlashGoogle$1.50$7.50$0.150$1.94
29Grok 4.5xAI$2.00$6.00$0.500$1.96
30Gemini 3.5 FlashGoogle$1.50$9.00$0.150$2.24
31Gemini 2.5 ProGoogle$1.25$10.00$0.125$2.37
32GPT 5 ChatOpenAI$1.25$10.00$0.125$2.37
33Claude Sonnet 5Anthropic$2.00$10.00$0.200$2.59
34Magistral Medium 1.2 2509Mistral$2.00$5.00$2.60
35GPT 5.6 TerraOpenAI$2.00$12.00$0.200$2.99
36Sonar Reasoning ProPerplexity$2.00$8.00$3.20
37GPT 5.3OpenAI$1.75$14.00$0.175$3.32
38Sonar ProPerplexity$3.00$15.00$5.40
39Grok 4xAI$3.00$15.00$5.40
40Claude Opus 5Anthropic$5.00$25.00$0.500$6.48
41GPT 5.6OpenAI$5.00$30.00$0.500$7.48
42Claude Fable 5Anthropic$10.00$50.00$1.00$12.96

42 models priced on the retrieval profile. Models without a published cache-read rate are billed at the standard input rate for the whole input, and their cache column reads em dash.

Waiting for a price cut is not a plan

Every line is a model we track, priced on the workload you picked above. Notice they run flat. A published rate almost never moves once it is set, so the rate you pay today is very likely the rate you will pay next year. The cheaper lines lower down are not price cuts; they are newer models that launched underneath the old ones, and moving to one is the lever you actually have. Hover any line to see what that model charges for your workload, and check back when you next review costs.

Provider
Blended cost per million tokens for every current model, over time
Highlight a model

Embed this chart

Paste this where you want the chart. It stays live: the lines redraw from our archive as rates move, and the reader can switch workload, filter by provider and pick out a model exactly as here. No script of ours runs on your page.

If the model you use is not on this list

This table shows only models a provider will sell you today. If the one you are running is missing, it has been retired or replaced, and a migration is coming whether or not you have planned for one. We keep its final rates and its full recorded history, so you can compare what you are paying now against what is actually on sale, and the ticker at the top of every page carries its last observed change.

Price your own workload Plain rate card Rates as JSON How to compare cheapness honestly

Using one of these figures

Rates on this page were captured 2026-08-10 from each provider's published pricing and committed to our archive, so this reading does not change when a provider edits a page. Quote the date with the number.

LLMPrice.io, “LLM API pricing compared by workload”. Rates captured 2026-08-10. https://llmprice.io/models

The same rates as JSON · how the archive works · questions about these numbers