See what your AI workload should cost.
Tell us what you're building, or drop last month's usage export. We price it across — models on live rates and show you the cheapest way to run it. Everything is computed in your browser. Tell us what you're building, or drop last month's usage export. Priced across — models on live rates, in your browser.
Not started yet? Estimate it. Already paying? Drop last month's usage export into the auditor and see what the same work costs today.
Built on our own price archive, kept since January 2024. What we hold, and how complete it is → · The LLMPrice indices →
What will it cost to run?
Computed in your browser · nothing uploadedPick what you're building and a number appears. Everything after that just narrows the range; you can stop whenever the answer is good enough.
What are you building?
Pick the closest match. You're choosing a cost shape, not a product: how much text goes in versus how much comes out is what drives the bill.
Not sure which fits? Compare all four cost shapes side by side →
Roughly how many a month?
A rough band is enough. If you don't know yet, keep the default: the table further down prices every level, so you can find your own row later.
How big is each one?
Volume says how many; this says how large each one is, and it moves the bill just as much. You always know this answer.
Refine it (optional)
Each one tightens the estimate. Skip anything you don't know.
Three quick questions (optional)
They sharpen the recommendation. Answer any, or none.
This decides whether prompt caching is within reach. The estimate keeps caching off either way; the breakdown shows what caching would save.
This changes which tier we recommend, not the cost.
The API calls the work itself makes, at published on-demand rates, every month it runs.
Building it, connecting it to your systems, and the person checking its output in the first months. Those usually cost more than this.
Which model, whether prompt caching is on, and what happens to the bill in a month twice as busy as expected.
Nothing on this page is uploaded; every figure is computed in your browser from published on-demand rates. · · Assumption spec, versioned (JSON) →
Advanced calculator
Raw token counts and cache shares · same engine as the estimatorCached input tokens are billed at each provider's discounted cache-read rate. Models without caching support are unaffected.
Monthly cost comparison
Stacked input cost (after caching) and output cost for your current scenario.
Full cost breakdown (cheapest first)
| # | Model | Provider | $ / 1M in | $ / 1M out | Monthly cost |
|---|
Estimates only. Rates sync from public pricing data and may lag official changes; confirm on each provider's pricing page. "Try" links go straight to the provider: we have no affiliate or referral arrangements, and nothing on this page is paid placement.
Compare vs. self-hosting your own GPUs (for CTOs weighing API vs. dedicated hardware) ▾
Open-weight models only (you cannot self-host GPT-5, Claude, or Gemini). Excludes engineering, ops, and reliability, which managed APIs include. Rough estimate; hardware throughput varies widely by model size and batching.
Recent calculations (saved on this device only) ▾
Get on the list for a heads-up when OpenAI, Claude, or DeepSeek slash API prices. No spam, unsubscribe anytime.
Check your inbox and click the confirmation link to finish subscribing.
AI compute got cheaper. Your bill decides whether you noticed.
What the same AI workload costs today versus July 2024 = 100: below 100 means cheaper, above means more expensive.
Two workloads, the same providers, two different outcomes.