What AI automation actually returns: building an ROI number you can defend
The 838%-ROI claims in vendor pitches are built on assumptions that don't survive scrutiny. A grounded formula for AI automation ROI, and the costs that get left out.
Where the money goes when you build an AI product MVP — engineering time, model inference, data preparation, evaluation — and which scope decisions move the total most.
Most quotes for an "AI MVP" arrive as a single number with no breakdown. That makes them impossible to compare, and impossible to argue with. A more useful approach is to understand the cost structure first, then judge any quote against it.
There are four cost centres in an AI product build, and they behave very differently from each other.
This dominates nearly every budget, and it is the one most sensitive to scope rather than to rates. An assistant that answers questions over your existing documentation is a fundamentally different build from one that takes actions in your systems — the second needs authentication, permissions, audit trails, and a rollback story for when the model gets it wrong.
The single most effective cost control is narrowing to one workflow and shipping it properly, rather than covering five workflows shallowly.
Per-token cost is easy to model once you know two numbers: how many model calls a typical user session makes, and how much context each call carries.
Teams tend to argue about which model to use, when the larger lever is usually context size. A long system prompt plus a generous retrieval window, sent on every turn, costs more than a more capable model used with disciplined context. Measure tokens per session before optimising anything else.
Inference also has a shape most infrastructure does not: it scales linearly with usage and never amortises. Unlike a server you have already paid for, every additional user costs real money on every interaction — which is why the pricing model for your product needs to be designed alongside the architecture, not after it.
The unglamorous line item. If your product reads contracts, invoices, claims, or support tickets, someone has to gather a representative sample, clean it, and label enough of it to know whether the system is working.
Teams routinely underestimate this because it looks like it is not engineering. It is, and skipping it does not remove the cost — it defers it into rework after launch, which is more expensive.
Hosting, vector storage, queues, monitoring — and the piece most first-time AI builds omit entirely: an evaluation harness.
Without a test set and a way to score outputs against it, every prompt change is a guess. You cannot tell whether today's tweak fixed the complaint from last week or quietly broke three things that were working. Evaluation is not a nice-to-have you add at maturity; it is the thing that makes iteration cheaper than starting over.
| Assumption | What actually happens |
|---|---|
| "We will fine-tune a model" | Retrieval plus a good prompt gets there faster and cheaper in most cases |
| "Evaluation can come later" | Every later change costs more, because nothing is measurable |
| "The demo is 80% of the product" | The demo is the easy 80% of the happy path, which is maybe 30% of the work |
| "Accuracy will improve with a better model" | Data quality and retrieval design usually matter more than model choice |
| "We can add permissions afterwards" | Retrofitting access control into an AI feature often means rebuilding it |
If you are comparing proposals, these five questions separate a considered estimate from a guess:
A vendor who answers these precisely has thought about your problem. A vendor who answers only with a timeline and a total has not.
Committing the whole budget before you know whether the approach works is the most common way to waste it. A phasing that reduces that risk:
| Phase | Goal | Spend shape |
|---|---|---|
| Discovery | Confirm the workflow, gather a sample, define what correct means | Small, days to two weeks |
| Thin slice | One workflow end to end, instrumented, with an eval set | The bulk of the initial budget |
| Hardening | Permissions, error paths, escalation, retention | Moderate, and often underestimated |
| Iterate | Improve against measured failures | Ongoing, tied to usage |
The discovery phase is the one teams skip and the one that most reduces total cost, because it is where you find out that the real problem is narrower — or different — than the brief assumed.
A useful rule: if the thin slice cannot be defined in a sentence that names one user, one workflow and one measurable outcome, the scope is not ready to be priced.
Pick one workflow with a measurable outcome. Instrument it so you can tell whether it works. Ship it to real users early enough that their behaviour, not your roadmap, decides what gets built second.
That approach costs less than the alternative — not because it cuts corners, but because it finds out what is worth building before you have paid to build the wrong thing.
The 838%-ROI claims in vendor pitches are built on assumptions that don't survive scrutiny. A grounded formula for AI automation ROI, and the costs that get left out.
A build vs buy framework for AI features — where buying wins, where a thin wrapper is genuinely correct, and the three conditions that justify building AI infrastructure yourself.
Most SaaS products need one or two AI features done well, not nine done shallowly. A filter for telling which pitched features are load-bearing and which are decoration nobody uses.