Gemini 3.5 Flash Lite API pricing
Published on-demand rates for Google's gemini/gemini-3.5-flash-lite, captured 2026-09-21 from public pricing data and committed to our archive. Context window 1.05M tokens; first observed in our archive 2026-07-27.
| Meter | USD per 1M tokens |
|---|---|
| Input | $0.300 |
| Output | $2.50 |
| Cached input | $0.030 |
Standard on-demand rates. Excludes batch endpoints and negotiated pricing; confirm on Google's pricing page before committing.
What it costs on a real workload
Per 1M tokens processed, at the rates above. How the four profiles are defined.
| Profile | In / out mix | Cached share | Cost |
|---|---|---|---|
| Retrieval | 800k / 200k | 70% | $0.589 |
| Chat | 500k / 500k | 30% | $1.36 |
| Content | 200k / 800k | 20% | $2.05 |
| Agent | 900k / 100k | 85% | $0.313 |
Observed rate history
No repricing observed since this model entered our archive on 2026-07-27. When a rate moves, the change appears here and on the sitewide ticker.
What the record shows
Gemini 3.5 Flash Lite has not repriced once since it entered our archive on 2026-07-27. That is the ordinary case rather than the exception: providers ship new models at lower rates far more often than they cut the price of one already in service, which is why a cheaper option usually arrives as a new name rather than as a smaller number on this page.
Against the 40 models we price, it is the 16th cheapest for retrieval-shaped work, which reads a great deal and writes little, and the 19th cheapest for content-shaped work, which does the reverse. The two are far apart, and that is the useful part: Gemini 3.5 Flash Lite is relatively better value the more your workload reads and the less it writes. A single position in a table sorted one way would have told you the opposite of the truth for the other kind of work.
It charges 8.3 times more for output than for input, against a median of 5.0 times across the set. That is wider than most, so the length of the answers you ask for matters more here than it does elsewhere.
Inside Google's own lineup it is the 4th cheapest of the 7 models we track on this workload, so 3 cheaper and 3 more expensive options sit on the same account behind the same key.
No model we price offers a larger context window, so a request too long for Gemini 3.5 Flash Lite is too long for anything we track, and the answer is to change the shape of the request rather than the model.
Cached input reads at $0.030 against a standard input rate of $0.300, a 90% discount on the part of your prompt that repeats.
It accepts image input and Google publishes a tokenisation rule for it, so a scanned page can be priced honestly: about 1,032 tokens for a US Letter page at 200 DPI, before the instructions you send with it.
Its 1.05M token context window holds roughly 786,000 words of English text in a single request, or about 1,016 scanned pages at the per-page figure above. A workload that does not fit cannot run here at any price, which is a capability limit rather than a cost one.
Priced near this one
Closest to Gemini 3.5 Flash Lite on a retrieval workload, cheaper and more expensive alike. Cost only; we publish no quality score.
| Model | Provider | Input | Output | vs Gemini 3.5 Flash Lite |
|---|---|---|---|---|
| GPT 5 Mini | OpenAI | $0.250 | $2.00 | 19% less |
| Devstral | Mistral | $0.400 | $2.00 | 22% more |
| Mistral Large | Mistral | $0.500 | $1.50 | 24% less |
| Grok Code Fast | xAI | $1.00 | $2.00 | 28% more |
Retrieval profile: 800k in, 200k out, 70% cached where offered.
No affiliate or referral arrangements: the console link carries no parameters and nothing on this page is paid placement. Gemini 3.5 Flash Lite is a trademark of its owner; LLMPrice.io is independent and unaffiliated.