LLMPrice.io

Gemini 2.5 Pro API pricing

Published on-demand rates for Google's gemini/gemini-2.5-pro, captured 2026-09-21 from public pricing data and committed to our archive. Context window 1.05M tokens; first observed in our archive 2025-06-23.

MeterUSD per 1M tokens
Input$1.25
Output$10.00
Cached input$0.125

Standard on-demand rates. Excludes batch endpoints and negotiated pricing; confirm on Google's pricing page before committing.

What it costs on a real workload

Per 1M tokens processed, at the rates above. How the four profiles are defined.

ProfileIn / out mixCached shareCost
Retrieval800k / 200k70%$2.37
Chat500k / 500k30%$5.46
Content200k / 800k20%$8.21
Agent900k / 100k85%$1.26

Compare every model on these profiles →

Observed rate history

ObservedInputOutputCached input
2025-06-23$1.25$10.00entered our archive
2025-07-14$1.25$10.00$0.313repricing observed
2026-01-19$1.25$10.00$0.125repricing observed

Dates are when we observed the rate in our own snapshots, which can lag the provider's announcement. Rates per 1M tokens.

What the record shows

We have observed 2 changes to Gemini 2.5 Pro's rates since it entered our archive on 2025-06-23. The most recent, on 2026-01-19, left the output rate unchanged. Every change we hold for it has been downward.

Against the 40 models we price, it is the 28th cheapest for retrieval-shaped work, which reads a great deal and writes little, and the 31st cheapest for content-shaped work, which does the reverse. The two are far apart, and that is the useful part: Gemini 2.5 Pro is relatively better value the more your workload reads and the less it writes. A single position in a table sorted one way would have told you the opposite of the truth for the other kind of work.

It charges 8.0 times more for output than for input, against a median of 5.0 times across the set. That is wider than most, so the length of the answers you ask for matters more here than it does elsewhere.

Inside Google's own lineup it is the 6th cheapest of the 7 models we track on this workload, so 5 cheaper and 1 more expensive option sits on the same account behind the same key.

One other model we price charges exactly $1.25 in and $10.00 out: GPT 5 Chat. Where a rate ties exactly, price has stopped being the deciding factor and this site has nothing further to tell you: that choice comes down to your own evaluations, which we do not run and do not score.

Cached input reads at $0.125 against a standard input rate of $1.25, a 90% discount on the part of your prompt that repeats.

It accepts image input and Google publishes a tokenisation rule for it, so a scanned page can be priced honestly: about 1,032 tokens for a US Letter page at 200 DPI, before the instructions you send with it.

Its 1.05M token context window holds roughly 786,000 words of English text in a single request, or about 1,016 scanned pages at the per-page figure above. A workload that does not fit cannot run here at any price, which is a capability limit rather than a cost one.

Priced near this one

Closest to Gemini 2.5 Pro on a retrieval workload, cheaper and more expensive alike. Cost only; we publish no quality score.

ModelProviderInputOutputvs Gemini 2.5 Pro
GPT 5 ChatOpenAI$1.25$10.00about the same
Claude Sonnet 5Anthropic$2.00$10.009% more
Magistral Medium 1.2Mistral$2.00$5.0010% more
Grok 4.6xAI$2.00$6.0017% less

Retrieval profile: 800k in, 200k out, 70% cached where offered.

Price your own prompt on it Estimate a whole project Plain rate card Google's console →

No affiliate or referral arrangements: the console link carries no parameters and nothing on this page is paid placement. Gemini 2.5 Pro is a trademark of its owner; LLMPrice.io is independent and unaffiliated.