LLMPrice.io

GPT 5 Nano API pricing

Published on-demand rates for OpenAI's gpt-5-nano, captured 2026-09-21 from public pricing data and committed to our archive. Context window 272k tokens; first observed in our archive 2025-08-11.

MeterUSD per 1M tokens
Input$0.050
Output$0.400
Cached input$0.005

Standard on-demand rates. Excludes batch endpoints and negotiated pricing; confirm on OpenAI's pricing page before committing.

What it costs on a real workload

Per 1M tokens processed, at the rates above. How the four profiles are defined.

ProfileIn / out mixCached shareCost
Retrieval800k / 200k70%$0.095
Chat500k / 500k30%$0.218
Content200k / 800k20%$0.328
Agent900k / 100k85%$0.051

Compare every model on these profiles →

Observed rate history

No repricing observed since this model entered our archive on 2025-08-11. When a rate moves, the change appears here and on the sitewide ticker.

What the record shows

GPT 5 Nano has not repriced once since it entered our archive on 2025-08-11. That is the ordinary case rather than the exception: providers ship new models at lower rates far more often than they cut the price of one already in service, which is why a cheaper option usually arrives as a new name rather than as a smaller number on this page.

Against the 40 models we price, it is the 4th cheapest for retrieval-shaped work, which reads a great deal and writes little, and the 6th cheapest for content-shaped work, which does the reverse. The two are close, so it holds its position whichever way your workload leans.

It charges 8.0 times more for output than for input, against a median of 5.0 times across the set. That is wider than most, so the length of the answers you ask for matters more here than it does elsewhere.

Inside OpenAI's own lineup it is the cheapest of the 10 models we track on this workload, which makes it the floor to measure the rest of the house against before looking outside it at all.

18 of the 40 models we price offer a larger context window, and 3 match it exactly. Context is a capability limit rather than a cost one: paying more buys no extra room unless the larger window is what you are paying for.

Cached input reads at $0.005 against a standard input rate of $0.050, a 90% discount on the part of your prompt that repeats.

It accepts image input and OpenAI publishes a tokenisation rule for it, so a scanned page can be priced honestly: about 3,680 tokens for a US Letter page at 200 DPI, before the instructions you send with it.

Its 272k token context window holds roughly 204,000 words of English text in a single request, or about 73 scanned pages at the per-page figure above. A workload that does not fit cannot run here at any price, which is a capability limit rather than a cost one.

Priced near this one

Closest to GPT 5 Nano on a retrieval workload, cheaper and more expensive alike. Cost only; we publish no quality score.

ModelProviderInputOutputvs GPT 5 Nano
Ministral 3 14BMistral$0.200$0.2005% more
Gemini 2.0 Flash LiteGoogle$0.075$0.3007% less
Gemini 2.5 Flash LiteGoogle$0.100$0.40016% more
Ministral 3 8BMistral$0.150$0.15022% less

Retrieval profile: 800k in, 200k out, 70% cached where offered.

Price your own prompt on it Estimate a whole project Plain rate card OpenAI's console →

No affiliate or referral arrangements: the console link carries no parameters and nothing on this page is paid placement. GPT 5 Nano is a trademark of its owner; LLMPrice.io is independent and unaffiliated.