LLMPrice.io

How Much to Mark Up API Costs When You Resell AI to Clients

Once an AI feature is live, you are reselling a metered product with a bill that moves every month. The build fee is a number you already know how to set. The markup on the running cost is a different decision, made once and then lived with for the life of the contract, and most consultants pick it by guessing rather than by working out what it actually has to cover.

Decide what the markup pays for before you pick a number

A markup on API cost is not profit alone. It has to cover three things that never show up on the provider's invoice: the time you spend watching the account for anomalies, the risk you carry when a client's usage runs past what you quoted, and the eventual afternoon you lose migrating the account to a cheaper model. Price those three separately in your own head even if you charge the client one blended number, because a client who questions the markup is really asking whether any of the three is real. "It covers my margin" is a weaker answer than naming what the margin buys.

Choose how the client sees the number

Two structures show up in practice, and they are not about how much you charge but about how visible the charging is. The first passes the provider's charge through at cost and adds a stated percentage on top, which is transparent and easy to defend in writing but turns every invoice into a conversation about a number the client did not choose. The second folds the running cost into one flat monthly fee set from your own worst-case estimate, so the client sees a single line and you carry the risk that usage runs lighter or heavier than planned. An itemized markup suits a client who already reads their bills line by line. A flat fee suits one who wants a single number and no meter to watch. Neither is more honest than the other, as long as the contract, not just the invoice, says which one the client bought. See the guide on quoting fixed-price AI work for how to set the cap that keeps a flat fee from eating your margin.

Your margin moves when the model does, even if you never touch the price

Providers cut prices mostly by shipping a new, cheaper model rather than repricing the one already in service. Our own index shows the cache-heavy agent workload down 48.4 percent since July 2024, and of the 26 monthly links in that series only 11 carry a price change on a model present in both months. If you charge a percentage of the raw cost, that means your fee falls mainly when you do the work of migrating the client to the newer model, so an account you rarely touch tends to overpay you relative to what a new client would sign at today's rate. If you charge a flat fee instead, the saving runs the other way: you keep all of it until you choose to pass some back. Either way, put a review date in the contract rather than leaving the number to drift on its own, so the renegotiation is scheduled rather than a conversation the client has to start.

Before setting the number, price the actual workload rather than guessing at the raw cost you are marking up. The same support conversation can run $6.00 a month on the cheapest model we price and $1,500 on the most expensive, covered in the chatbot cost guide, so a markup set against the wrong model is wrong by multiples before you have added a cent. Start from the guide on pricing an AI project to size the running cost, then decide the percentage on top of a number you trust.

Related