LLMPrice.io

Ministral 3 3B API pricing 2512

Published on-demand rates for Mistral's mistral/ministral-3-3b-2512, captured 2026-09-21 from public pricing data and committed to our archive. Context window 131k tokens; first observed in our archive 2026-03-02.

MeterUSD per 1M tokens
Input$0.100
Output$0.100
Cached input$0.010

Standard on-demand rates. Excludes batch endpoints and negotiated pricing; confirm on Mistral's pricing page before committing.

What it costs on a real workload

Per 1M tokens processed, at the rates above. How the four profiles are defined.

ProfileIn / out mixCached shareCost
Retrieval800k / 200k70%$0.050
Chat500k / 500k30%$0.086
Content200k / 800k20%$0.096
Agent900k / 100k85%$0.031

Compare every model on these profiles →

Observed rate history

ObservedInputOutputCached input
2026-03-02$0.100$0.100entered our archive
2026-08-30$0.100$0.100$0.010repricing observed

Dates are when we observed the rate in our own snapshots, which can lag the provider's announcement. Rates per 1M tokens.

What the record shows

We have observed one change to Ministral 3 3B's rates since it entered our archive on 2026-03-02. The most recent, on 2026-08-30, left the output rate unchanged.

Against the 40 models we price, it is the 1st cheapest for retrieval-shaped work, which reads a great deal and writes little, and the 1st cheapest for content-shaped work, which does the reverse. The two are close, so it holds its position whichever way your workload leans.

It charges the same rate for output as for input. That is unusual: 4 of the 40 models we price do it, and everywhere else output costs a median of 5.0 times input. On this model the length of the answer costs no more per token than the length of the question, which changes which workloads suit it.

Inside Mistral's own lineup it is the cheapest of the 10 models we track on this workload, which makes it the floor to measure the rest of the house against before looking outside it at all.

32 of the 40 models we price offer a larger context window. Context is a capability limit rather than a cost one: paying more buys no extra room unless the larger window is what you are paying for.

Cached input reads at $0.010 against a standard input rate of $0.100, a 90% discount on the part of your prompt that repeats.

It accepts image input, but Mistral publishes no rule for how images become tokens, so we will not price a scanned page on it. That is a gap in the provider's documentation rather than a limit of the model, and it is the reason this model is excluded from image estimates in our estimator instead of being given a plausible figure.

Its 131k token context window holds roughly 98,000 words of English text in a single request. A workload that does not fit cannot run here at any price, which is a capability limit rather than a cost one.

Priced near this one

Closest to Ministral 3 3B on a retrieval workload, cheaper and more expensive alike. Cost only; we publish no quality score.

ModelProviderInputOutputvs Ministral 3 3B
Ministral 3 8BMistral$0.150$0.15050% more
Gemini 2.0 Flash LiteGoogle$0.075$0.30078% more
GPT 5 NanoOpenAI$0.050$0.40091% more
Ministral 3 14BMistral$0.200$0.200100% more

Retrieval profile: 800k in, 200k out, 70% cached where offered.

Price your own prompt on it Estimate a whole project Plain rate card Mistral's console →

No affiliate or referral arrangements: the console link carries no parameters and nothing on this page is paid placement. Ministral 3 3B is a trademark of its owner; LLMPrice.io is independent and unaffiliated.