LLMPrice.io
Free · nothing uploaded · no signup

Price a prompt on every model.

Paste a prompt, a document, or a code file. The token counter is built in: your text is counted as you type and priced across every current first-party model at your scale, with caching and batch discounts applied. Computed in your browser from published on-demand rates; nothing you paste leaves your device.

Examples:
500 input tokens (sample estimate, paste text for a real count)
Expected response length
Scale
Click to unselect →
Cheapest option
...
...
Best value
...
...
Flagship option
...
...

Monthly cost comparison

Stacked input cost (after caching) and output cost for your current scenario.

Full cost breakdown (cheapest first)

#ModelProvider$ / 1M in$ / 1M outMonthly cost

Estimates only. Rates sync from public pricing data and may lag official changes; confirm on each provider's pricing page. "Try" links go straight to the provider: we have no affiliate or referral arrangements, and nothing on this page is paid placement.

Compare vs. self-hosting your own GPUs (for CTOs weighing API vs. dedicated hardware)

Open-weight models only (you cannot self-host GPT-5, Claude, or Gemini). Excludes engineering, ops, and reliability, which managed APIs include. Rough estimate; hardware throughput varies widely by model size and batching.

Recent calculations (saved on this device only)
History never leaves this device.
Price Drop Alerts

Get on the list for a heads-up when OpenAI, Claude, or DeepSeek slash API prices. No spam, unsubscribe anytime.

How it prices your workload

Token counts are estimated at roughly four characters per token, which lands within about ten to fifteen percent of the providers' own tokenizers for English text. Rates sync on load from public pricing data, and every figure states its snapshot date. The ranking is computed purely from cost: no affiliate deals, no paid placement.

Answering a question about a project rather than a prompt? The project estimator prices a workload from plain answers, no tokens required. Trying to shrink the prompt itself? The prompt optimizer strips filler and reorders it so caching can bite.

Related: Token counter · Flat rate reference · Models compared by workload · How token pricing works