LLMPrice.io

Quoting Fixed-Price AI Work Without Eating the Variance

Fixed-price work and metered dependencies do not mix well. If you quote one number for a system whose running cost rises every time the client uses it, you have written an open cheque and put your name on it. The uncomfortable part is that the better the project goes, the more it costs you. There are three ways out and none of them require you to abandon fixed pricing.

Three structures that hold

The first is a straight pass-through: you quote the build, the client pays the provider bill directly or at cost. Cleanest, and it removes the risk entirely, but some clients do not want a variable line in their accounts. The second is a capped pass-through: you carry the API cost up to an agreed monthly ceiling and anything above it is billed on. The third is fixed price with a usage limit written into the scope, stated in the client's units rather than yours, such as documents processed or conversations handled per month.

Set the cap from the number you do not expect

The cap is the whole game, and the common mistake is setting it near the expected figure. Price the workload on a flagship model, take that figure, and set the ceiling there or a little above. You will almost certainly run well under it, and you will have bought yourself the right to be wrong about volume, wrong about how much context the system ends up needing, and wrong about which model you eventually ship on. A cap set at the expected cost is not a cap; it is a prediction.

Write the limit in their units, and say who reads the meter

A limit expressed in tokens is not a limit, because the client cannot tell whether they are near it and will not know they crossed it until you tell them. Write it in the unit their business already counts: documents processed, conversations handled, invoices read, per month. Then answer the two questions that always follow. Who measures it, and how often do they see the number?

A monthly line in a report is enough. What matters is that the figure arrives before the invoice does, from you, in the same units as the contract. A client who watched a number climb for two months has a very different reaction to an overage than one who learns about both at once.

Agree what happens when it is breached

Write down what happens on the month the ceiling is passed, before it is, because that conversation is easy in advance and unpleasant afterwards. Usually the answer is that the overage is billed at cost and you both look at why. Sometimes it is that the system throttles. What matters is that the client hears the word "cap" while they are signing, and not for the first time in an email about an invoice.

Two costs that eat fixed-price work from outside the API bill

Human review in the first months. Somebody checks the output often enough to trust it, and that is real time that most quotes omit entirely. It usually costs more than the API does at small scale. Price it explicitly and taper it, rather than absorbing it silently and discovering the project was unprofitable in month two.

Migration you did not plan for. Model prices at the frontier fall largely by replacement rather than by cuts to models already in service: across the 26 monthly links in our index, only 11 carry a price change on a model present in both periods. So the cheaper option arrives as a new model name, and capturing it means re-testing and re-deploying. If your fixed price covers a year, decide up front whether migrations are in scope, because "we will just move to whatever is cheapest" is unpaid work with your name on it.

Why a cap set on today's flagship stays safe

The reassuring corollary of all this: rates for a given model very rarely rise. Set your ceiling from a flagship price today and time is mostly on your side, because the risk that the same model costs more next year is small, while the chance that a cheaper equivalent appears underneath it is high. That is what makes the flagship-priced cap a conservative instrument rather than an expensive one. It buys you the right to be wrong about volume, wrong about context growth, and wrong about which model you ship on, for a premium you will probably never pay.

Related