LLMPrice.io

How to Explain AI Costs to a Non-Technical Client

Your client is not asking what a token is. They are asking whether this thing is going to surprise them on a Tuesday. Answer that question and the rest of the conversation gets easier. Lead with tokens and you will spend twenty minutes teaching a unit of measurement to somebody who will never use it, and they will still not know whether they can afford it.

The comparison that works

AI is metered like electricity, not licensed like software. There is no seat fee, nothing to buy up front, and nobody to negotiate a contract with. You pay for what runs, every time it runs, and the meter reads zero on a quiet month. That framing answers the two questions a business owner actually has, which are what happens if nobody uses it and what happens if everybody does. The honest answers are nothing and it goes up roughly in line with use.

Give them a range and a ceiling

A single number invites a single question, which is whether that number is right. A range invites a better one. Show what the workload costs on a cheap model and on an expensive one, say which you plan to use and why, and name the figure you would be surprised to exceed. Clients do not need precision here. They need to know the worst case is survivable, and they will remember that you told them the ceiling rather than the best case.

Leave them one lever, not five

Do not hand over a menu of optimizations. Give them the one decision that is genuinely theirs: how good the answers need to be. Cheaper models handle routine work well and cost a fraction of the flagship tier, and that single choice moves the bill more than everything else combined. Everything else, the caching and the batching and the prompt trimming, is your job and belongs in your workflow rather than in their inbox.

"Will this get cheaper?" Answer it carefully

This question comes up in every one of these conversations, usually as a reason to defer a decision, and the loose answer creates a problem for you later.

Prices at the frontier really have fallen. Our own index, which prices the same workload month after month, has the cache-heavy agent shape down 48.4 percent since July 2024. But look at how that happened before you promise anything, because it is not what it looks like. Of the 26 monthly links in that series, only 11 carry a price change on a model present in both months. The frontier got cheaper mostly by replacement, not by price cuts. Providers ship new models at lower rates far more often than they reduce the price of a model already in service.

Which means the honest answer is: the market gets cheaper, and your bill does not, until somebody does the work of migrating and re-testing. Waiting for a price cut on the model you are running is not a plan. Say that plainly, and you have both set the expectation correctly and explained why a migration is a piece of work rather than a switch you flip. The output-heavy shape has moved only 13.0 percent over the same period, so if their workload writes more than it reads, temper it further.

What to actually put in front of them

One slide, four lines. What it costs to run at the volume we expect. What it costs at the volume that would worry us. Which of those two the price you are quoting is based on. And what we do if the second one happens.

No token counts, no model names, no per-million rates. Those belong in the appendix for the one client in ten who asks, and having them ready is what makes the four lines credible. If you want the underlying figures to hand, the index is the version of this argument written for someone who does want the detail.

Related