Quoting Fixed-Price AI Work Without Eating the Variance
Fixed-price work and metered dependencies do not mix well. If you quote one number for a system whose running cost rises every time the client uses it, you have written an open cheque and put your name on it. The uncomfortable part is that the better the project goes, the more it costs you. There are three ways out and none of them require you to abandon fixed pricing.
Three structures that hold
The first is a straight pass-through: you quote the build, the client pays the provider bill directly or at cost. Cleanest, and it removes the risk entirely, but some clients do not want a variable line in their accounts. The second is a capped pass-through: you carry the API cost up to an agreed monthly ceiling and anything above it is billed on. The third is fixed price with a usage limit written into the scope, stated in the client's units rather than yours, such as documents processed or conversations handled per month.
Set the cap from the number you do not expect
The cap is the whole game, and the common mistake is setting it near the expected figure. Price the workload on a flagship model, take that figure, and set the ceiling there or a little above. You will almost certainly run well under it, and you will have bought yourself the right to be wrong about volume, wrong about how much context the system ends up needing, and wrong about which model you eventually ship on. A cap set at the expected cost is not a cap; it is a prediction.
Agree what happens when it is breached
Write down what happens on the month the ceiling is passed, before it is, because that conversation is easy in advance and unpleasant afterwards. Usually the answer is that the overage is billed at cost and you both look at why. Sometimes it is that the system throttles. What matters is that the client hears the word "cap" while they are signing, and not for the first time in an email about an invoice.