LLMPrice.io

LLM API pricing, ranked by cost

Every current first-party text model, priced on the workload you actually run. Cheapest depends on the shape of the work: a model that wins on a retrieval workload can lose badly on a content one. Rates captured 2026-09-21 from public pricing data and committed to our archive.

Workload profile

800k input, 200k output, 70% of input served from cache. Long context read repeatedly.

#
1Ministral 3 3B 2512Mistral$0.100$0.100$0.010$0.050
2Ministral 3 8B 2512Mistral$0.150$0.150$0.015$0.074
3Gemini 2.0 Flash Lite 001Google$0.075$0.300$0.019$0.088
4GPT 5 NanoOpenAI$0.050$0.400$0.005$0.095
5Ministral 3 14B 2512Mistral$0.200$0.200$0.020$0.099
6Gemini 2.5 Flash LiteGoogle$0.100$0.400$0.010$0.110
7Devstral Small 2512Mistral$0.100$0.300$0.140
8Mistral vibe cli FastMistral$0.150$0.600$0.015$0.164
9Codestral 2508Mistral$0.300$0.900$0.030$0.269
10GPT 5.6 LunaOpenAI$0.200$1.20$0.020$0.299
11GPT 5.4 NanoOpenAI$0.200$1.25$0.020$0.309
12DeepSeek FlashDeepSeek$0.300$1.20$0.006$0.315
13Gemini 3.1 Flash LiteGoogle$0.250$1.50$0.025$0.374
14Mistral Large 2512Mistral$0.500$1.50$0.050$0.448
15GPT 5 MiniOpenAI$0.250$2.00$0.025$0.474
16Gemini 3.5 Flash LiteGoogle$0.300$2.50$0.030$0.589
17Devstral 2512Mistral$0.400$2.00$0.720
18Grok Code FastxAI$1.00$2.00$0.200$0.752
19Grok 4.3xAI$1.25$2.50$0.200$0.912
20Gemini 3.8 FlashGoogle$0.750$3.75$0.075$0.972
21SonarPerplexity$1.00$1.00$1.00
22GPT 5.4 MiniOpenAI$0.750$4.50$0.075$1.12
23DeepSeek V4 ProDeepSeek$1.32$3.96$0.044$1.13
24Claude Haiku 4.5Anthropic$1.00$5.00$0.100$1.30
25Sonar ReasoningPerplexity$1.00$5.00$1.80
26Mistral Medium 3.5Mistral$1.50$7.50$0.150$1.94
27Grok 4.6xAI$2.00$6.00$0.500$1.96
28Gemini 2.5 ProGoogle$1.25$10.00$0.125$2.37
29GPT 5 ChatOpenAI$1.25$10.00$0.125$2.37
30Claude Sonnet 5Anthropic$2.00$10.00$0.200$2.59
31Magistral Medium 1.2 2509Mistral$2.00$5.00$2.60
32GPT 5.6 TerraOpenAI$2.00$12.00$0.200$2.99
33Gemini omni 1.1 FlashGoogle$1.50$9.00$3.00
34Sonar Reasoning ProPerplexity$2.00$8.00$3.20
35GPT 5.3OpenAI$1.75$14.00$0.175$3.32
36GPT 5.6OpenAI$4.00$20.00$0.400$5.18
37Sonar ProPerplexity$3.00$15.00$5.40
38Claude Opus 5Anthropic$5.00$25.00$0.500$6.48
39Claude Fable 5.1Anthropic$10.00$50.00$0.250$12.54
40GPT 6 astraOpenAI$10.00$50.00$1.00$12.96

40 models priced on the retrieval profile. Models without a published cache-read rate are billed at the standard input rate for the whole input, and their cache column reads em dash.

Waiting for a price cut is not a plan

Every line is a model we track, priced on the workload you picked above. Notice they run flat. A published rate almost never moves once it is set, so the rate you pay today is very likely the rate you will pay next year. The cheaper lines lower down are not price cuts; they are newer models that launched underneath the old ones, and moving to one is the lever you actually have. Hover any line to see what that model charges for your workload, and check back when you next review costs.

Provider
Blended cost per million tokens for every current model, over time
Highlight a model

Embed this chart

Paste this where you want the chart. It stays live: the lines redraw from our archive as rates move, and the reader can switch workload, filter by provider and pick out a model exactly as here. No script of ours runs on your page.

If the model you use is not on this list

This table shows only models a provider will sell you today. If the one you are running is missing, it has been retired or replaced, and a migration is coming whether or not you have planned for one. We keep its final rates and its full recorded history, so you can compare what you are paying now against what is actually on sale, and the ticker at the top of every page carries its last observed change.

Price your own workload Plain rate card Rates as JSON How to compare cheapness honestly

Using one of these figures

Rates on this page were captured 2026-09-21 from each provider's published pricing and committed to our archive, so this reading does not change when a provider edits a page. Quote the date with the number.

LLMPrice.io, “LLM API pricing compared by workload”. Rates captured 2026-09-21. https://llmprice.io/models

The same rates as JSON · how the archive works · questions about these numbers