Index methodology
Everything needed to reproduce the indices from published rates. If anything here cannot be checked by a reader, it does not belong in the index.
Construction
For each constituent and each profile, cost = (uncached input × input rate) + (cached input × cache-read rate) + (output × output rate), all rates per million tokens. Constituents are equal-weighted, the mean is taken, and the series is rebased so the base period equals 100.
Where a model publishes no cache-read rate, the entire input bills at the standard rate and the constituent is marked on the issue page. We do not credit a discount a provider does not offer.
The four profiles, frozen
| Index | Input | Output | Cached share |
|---|---|---|---|
| LLMPrice Retrieval Index | 800,000 | 200,000 | 70% |
| LLMPrice Chat Index | 500,000 | 500,000 | 30% |
| LLMPrice Content Index | 200,000 | 800,000 | 20% |
| LLMPrice Agent Index | 900,000 | 100,000 | 85% |
All figures are per one million tokens processed. The profiles are frozen: changing them would break comparability with every prior issue.
Matching a real workload to a profile
The audit tool on the home page benchmarks an uploaded usage export against these indices, which means it has to decide which profile a given line of usage resembles. The rule is published here because an unexplained classification would invite doubt about everything else on this page.
Each row of an export is assigned to the profile whose own output share is nearest to the row's. Output share is output tokens divided by total tokens. The boundaries are the midpoints between adjacent profiles and are derived from the table above, not chosen:
| Profile | Its own output share | A row is assigned to it when output share is |
|---|---|---|
| Agent | 10% | below 15% |
| Retrieval | 20% | 15% to 35% |
| Chat | 50% | 35% to 65% |
| Content | 80% | 65% and above |
Cacheable share is not used to classify. Most usage exports do not report how much input was served from cache, and assuming zero where it is unreported would push almost every row toward the output-heavy profile. Where an export does report cached input tokens, the observed share is shown against the matched profile's assumed share instead, so a reader can see exactly where their workload departs from the index's assumption rather than having that departure silently absorbed into a label.
A profile match is a comparison, not a claim about what the work is. Two teams with the same token shape can be doing entirely different things.
The project estimator's assumptions are not the index's
The estimator on the home page maps its six workload options onto these same four profiles, so what it cites is what this page defines. It differs from the index in one deliberate way: the estimator prices every profile at 0% cached, while the index prices the cached shares of the market it tracks (Retrieval 70%, Agent 85%). Small businesses commonly build through no-code tooling that does not expose prompt caching, so reusing the index's shares would understate their real cost. The estimator states this divergence on its result page wherever it benchmarks against an index.
The estimator's token defaults per workload, its size mappings, and its tool-loop multipliers are editorial judgment about workload shape, not observed prices. They are published in full as a versioned file at /estimator-spec-v1.json, shown with sources on the estimate itself, and user-adjustable. They are structured to be replaced: once enough real workloads have passed through the bill auditor, the defaults will be superseded by empirical medians from anonymised audit data and the version number will change. No index figure depends on any of them.
The basket: the LLMPrice 12
Inclusion rule. Current, generally available, first-party text models from the major providers, excluding previews, deprecated models, and anything without published on-demand pricing.
One stated exception to the preview exclusion. Google ships each new Pro generation under a preview-suffixed key for months while it is the current, and only, Pro tier on Google's own price list. Excluding preview keys there would freeze the Google flagship slot on a superseded model indefinitely, so that slot admits a preview-named Pro when it is the newest Pro on Google's published pricing. The engine has always applied this; it is stated here as of 2026-07-31 because a rule the engine applies and the page does not state is a defect (found by external audit). No other slot admits preview keys.
Six providers contribute two constituents each: the model the provider positions as its frontier offering, and the model it positions as its mid or value offering. The split uses each provider's own product positioning, not our judgement, and never price order.
Constituents are not a list of model identifiers. Each slot names a product family, and the engine selects the newest member of that family present in the snapshot being priced, where newest means the earliest date we observed it in our own archive. A new release therefore enters the basket by itself and is chain-linked like any other substitution. An earlier version of this engine used a fixed list of identifiers and could only ever select a model already typed into it; that fault, and the correction, are described on the 2026-08 issue.
Moving aliases such as -latest are never priced, because they would change what the index measures without producing a substitution anyone could see.
Line assignment, and where each comes from
Every flagship and value assignment below is taken from the provider's own published material, cited here so the split is auditable rather than asserted.
| Provider | Line | Family | Source |
|---|---|---|---|
| OpenAI | flagship | /^gpt-\d+(\.\d+)?(o|-turbo)?$/ | platform.openai.com/docs/pricing, flagship GPT line (verified 2026-07-29) |
| OpenAI | value | /^(gpt-\d+(\.\d+)?o?-mini|gpt-3\.5-turbo)$/ | platform.openai.com/docs/pricing, "mini" tier (verified 2026-07-29) |
| Anthropic | flagship | /^(claude-opus-\d+(-\d+)?|claude-3-opus-\d{8})$/ | anthropic.com/pricing, Opus tier (verified 2026-07-29) |
| Anthropic | value | /^(claude-sonnet-\d+(-\d+)?|claude-3(-5)?-sonnet-\d{8})$/ | anthropic.com/pricing, Sonnet tier (verified 2026-07-29) |
| flagship | /^gemini\/gemini-(\d+(\.\d+)?-)?pro(-preview)?$/ | ai.google.dev/gemini-api/docs/pricing, Pro tier (verified 2026-07-29) | |
| value | /^gemini\/gemini-\d+(\.\d+)?-flash$/ | ai.google.dev/gemini-api/docs/pricing, Flash tier (verified 2026-07-29) | |
| DeepSeek | flagship | /^(deepseek-v\d+-pro|deepseek\/deepseek-(r1|reasoner))$/ | api-docs.deepseek.com/quick_start/pricing, reasoning tier (verified 2026-07-29) |
| DeepSeek | value | /^(deepseek-v\d+-flash|deepseek\/deepseek-(chat|coder))$/ | api-docs.deepseek.com/quick_start/pricing, standard tier (verified 2026-07-29) |
| xAI | flagship | /^xai\/grok-(\d+\.\d+|\d+)$/ | docs.x.ai/docs/models, "most intelligent and fastest model" (verified 2026-07-29) |
| xAI | value | /^xai\/grok-(\d+(\.\d+)?-mini|beta)$/ | docs.x.ai/docs/models, mini tier (verified 2026-07-29) |
| Mistral | flagship | /^mistral\/mistral-large-\d{4}$/ | mistral.ai/pricing/api Large tier; mistral.ai/pricing FAQ names Large best overall performance while Medium 3.5 is a higher-priced coding-specialised line (re-verified 2026-07-31; on the October review list) |
| Mistral | value | /^mistral\/mistral-medium-\d{4}$/ | mistral.ai/pricing/api, Medium tier (re-verified 2026-07-31) |
Where a provider's own naming puts its value tier above its flagship on price, that is reported as found and never corrected away. Mistral currently prices Mistral Medium (2026-04) at $1.50/$7.50 and Mistral Large (2025-12) at $0.50/$1.50, so its value constituent costs more than its flagship. This is Mistral's price list, not an error in ours, and it is one reason the premium is taken as a median of within-provider ratios rather than a mean.
Flagship Premium
The premium is the median of each provider's own flagship-to-value ratio. Ratios are formed within a provider first, where the substitution is one a reader could actually make, and the median is then taken across the six providers so no single provider can swing it.
This definition was revised on 2026-07-29. It was previously a median across all flagships over a median across all values, which compared providers to one another: on the corrected basket that returned 1.2x by measuring one provider's flagship against another's value tier, a number describing no decision anyone can make. The superseded function is retained in the engine so the revision can be checked.
There is no quality or intelligence score anywhere in this index. We publish none, we scrape none, and we republish no third party's benchmark composite. The index measures price, and price only.
Membership is reviewed quarterly, on the first business day of January, April, July and October.
Base period
The base period is 2024-07 = 100, and it will never change silently.
Stage 0's backfill reconstructed verifiable snapshots from LiteLLM's public git history reaching January 2024, and that source is cited here. The series nonetheless begins 2024-07, because before that month fewer than eight of the twelve slots have published rates in the archive, and an index computed on four constituents is a different instrument. No value before 2024-07 is reconstructed, estimated, or inferred.
Chain-linking
A fixed basket goes stale within a year: deprecated models stop being repriced and the series flattens into a line that measures nothing. So the basket is maintained, and every change is chain-linked.
At a change, the index moves by the price relative of the constituents present in both periods. Members that entered or left contribute nothing, so a substitution cannot move the index by itself; only a rate change on a comparable constituent can. Without this, every basket update would write a false cliff into the chart.
Where a boundary has no constituent in common at all, the series cannot be linked across it. That is flagged rather than smoothed over.
Every substitution is recorded with its date, what entered, what left, and the linking factor applied, on the issue where it occurs and in the permanent log on each issue page.
Provisional readings
Between monthly readings the site's ticker shows a provisional now-reading: the latest official level extended by the same matched-model chain link, computed over our archive snapshots as they arrive rather than on the first of the month. It is marked "prov" wherever it appears.
The provisional level is anchored to the published series by construction: it equals the official level times the chained movement observed since that reading, so it collapses onto the official value the day an issue publishes and can never disagree with a published number. It is superseded by each monthly reading, it is never frozen, and it is not citable. Cite the dated issue.
Revision policy
A backdated provider change restates the affected period, and the restatement is flagged in the next issue. Past issues are annotated, never silently edited. A published issue is a frozen artifact: it is not regenerated with later data, so a citation of it stays valid.
Exclusions
The index prices published on-demand rates on the stated token mixes. It excludes batch pricing, negotiated enterprise rates, and reasoning-token overhead. Rate dates are the dates we observed a rate in our own archive, which can lag a provider's announcement.