Gemini 3.1 Flash Lite API pricing
Published on-demand rates for Google's gemini/gemini-3.1-flash-lite, captured 2026-08-10 from public pricing data and committed to our archive. Context window 1.05M tokens; first observed in our archive 2026-05-25.
| Meter | USD per 1M tokens |
|---|---|
| Input | $0.250 |
| Output | $1.50 |
| Cached input | $0.025 |
Standard on-demand rates. Excludes batch endpoints and negotiated pricing; confirm on Google's pricing page before committing.
What it costs on a real workload
The four frozen profiles the LLMPrice indices price, per 1M tokens processed, at the rates above. Where no cache-read rate is published, the whole input bills at the standard rate.
| Profile | In / out mix | Cached share | Cost |
|---|---|---|---|
| Retrieval | 800k / 200k | 70% | $0.374 |
| Chat | 500k / 500k | 30% | $0.841 |
| Content | 200k / 800k | 20% | $1.24 |
| Agent | 900k / 100k | 85% | $0.203 |
Profile definitions are frozen and published on the methodology page. Compare every model on these profiles →
Observed rate history
No repricing observed since this model entered our archive on 2026-05-25. When a rate moves, the change appears here and on the sitewide ticker.
Priced near this one
The models closest to Gemini 3.1 Flash Lite on a retrieval workload, cheaper and more expensive alike. We publish no quality score, so this is what they cost and nothing more.
| Model | Provider | Input | Output | vs Gemini 3.1 Flash Lite |
|---|---|---|---|---|
| Grok Code Fast | xAI | $0.200 | $1.50 | 4% less |
| Codestral | Mistral | $0.300 | $0.900 | 12% more |
| GPT 5.4 Nano | OpenAI | $0.200 | $1.25 | 17% less |
| GPT 5.6 Luna | OpenAI | $0.200 | $1.20 | 20% less |
Compared on the retrieval profile: 800k input, 200k output, 70% of input cached where the model offers a cache rate.
No affiliate or referral arrangements: the console link carries no parameters and nothing on this page is paid placement. Gemini 3.1 Flash Lite is a trademark of its owner; LLMPrice.io is independent and unaffiliated.