DeepSeek V4 Pro API pricing
Published on-demand rates for DeepSeek's deepseek-v4-pro, captured 2026-09-21 from public pricing data and committed to our archive. Context window 1M tokens; first observed in our archive 2026-06-22.
| Meter | USD per 1M tokens |
|---|---|
| Input | $1.32 |
| Output | $3.96 |
| Cached input | $0.044 |
Standard on-demand rates. Excludes batch endpoints and negotiated pricing; confirm on DeepSeek's pricing page before committing.
What it costs on a real workload
Per 1M tokens processed, at the rates above. How the four profiles are defined.
| Profile | In / out mix | Cached share | Cost |
|---|---|---|---|
| Retrieval | 800k / 200k | 70% | $1.13 |
| Chat | 500k / 500k | 30% | $2.45 |
| Content | 200k / 800k | 20% | $3.38 |
| Agent | 900k / 100k | 85% | $0.608 |
Observed rate history
| Observed | Input | Output | Cached input | |
|---|---|---|---|---|
| 2026-06-22 | $0.435 | $0.870 | $0.004 | entered our archive |
| 2026-08-22 | $1.32 | $3.96 | $0.044 | repricing observed |
Dates are when we observed the rate in our own snapshots, which can lag the provider's announcement. Rates per 1M tokens.
What the record shows
We have observed one change to DeepSeek V4 Pro's rates since it entered our archive on 2026-06-22. In the most recent, on 2026-08-22, DeepSeek raised the output rate by 355%, from $0.870 to $3.96 per 1M.
Against the 40 models we price, it is the 23rd cheapest for retrieval-shaped work, which reads a great deal and writes little, and the 22nd cheapest for content-shaped work, which does the reverse. The two are close, so it holds its position whichever way your workload leans.
It charges 3.0 times more for output than for input, against a median of 5.0 times across the set. That is narrower than most, so it punishes long answers less than the typical model does.
Inside DeepSeek's own lineup it is the most expensive of the 2 models we track on this workload. That is worth settling before comparing it against another vendor's flagship, because the cheaper alternative is on the same account, behind the same key, with no migration to do.
7 of the 40 models we price offer a larger context window, and 5 match it exactly. Context is a capability limit rather than a cost one: paying more buys no extra room unless the larger window is what you are paying for.
Cached input reads at $0.044 against a standard input rate of $1.32, a 97% discount on the part of your prompt that repeats.
It does not accept image input, so scanned documents have to be turned into text before they reach it.
Its 1M token context window holds roughly 750,000 words of English text in a single request. A workload that does not fit cannot run here at any price, which is a capability limit rather than a cost one.
Priced near this one
Closest to DeepSeek V4 Pro on a retrieval workload, cheaper and more expensive alike. Cost only; we publish no quality score.
| Model | Provider | Input | Output | vs DeepSeek V4 Pro |
|---|---|---|---|---|
| GPT 5.4 Mini | OpenAI | $0.750 | $4.50 | 1% less |
| Sonar | Perplexity | $1.00 | $1.00 | 12% less |
| Gemini 3.8 Flash | $0.750 | $3.75 | 14% less | |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | 14% more |
Retrieval profile: 800k in, 200k out, 70% cached where offered.
No affiliate or referral arrangements: the console link carries no parameters and nothing on this page is paid placement. DeepSeek V4 Pro is a trademark of its owner; LLMPrice.io is independent and unaffiliated.