What AI automation actually returns: building an ROI number you can defend
The 838%-ROI claims in vendor pitches are built on assumptions that don't survive scrutiny. A grounded formula for AI automation ROI, and the costs that get left out.
Most SaaS products need one or two AI features done well, not nine done shallowly. A filter for telling which pitched features are load-bearing and which are decoration nobody uses.
Every "AI features for SaaS in 2026" post reads the same: predictive churn alerts, AI-personalised dashboards, anomaly detection, a chat assistant, smart recommendations, automated reporting — eight or nine features presented as a menu, as if a serious product needs to check every box. That's listicle padding, not product strategy, and following it produces a roadmap with nine shallow AI features and zero deep ones.
The better question isn't "which of these should we add." It's "how many of these does our product actually need, and which ones." For most SaaS products, done well, the answer is one or two.
Before any AI feature goes on a roadmap, run it through one question: is this replacing something the user currently does by hand and actively dislikes doing, or is it AI because the category expects it?
The first kind gets used every session, because it removes real friction — users notice immediately when it's gone. The second kind gets a launch announcement, a screenshot in the sales deck, and a usage graph that flatlines within a month, because nobody's workflow actually required it. The tell is in your own usage data, if you're honest about reading it: a load-bearing feature has a usage curve that looks like a core workflow. Decoration has a usage curve that looks like a feature nobody remembers exists.
Predictive churn alerts. Genuinely load-bearing when a customer success team currently spends hours manually reviewing account health signals across scattered dashboards to guess who's at risk — the alert replaces a task someone dislikes and does badly under time pressure. Decoration when there's no CS team acting on the signal, or when the underlying usage data is too thin to predict anything reliably and the "prediction" is closer to a coin flip dressed up in a percentage. Build this only after confirming someone will act differently because of the alert — if the answer to "what would you do differently" is vague, the feature won't get used regardless of accuracy.
AI-personalised dashboards. This one is genuinely mixed, and worth being honest about rather than pitching flatly either way. Users who check the same three numbers every morning build muscle memory around where those numbers sit — an algorithm that reshuffles the layout breaks that habit more than it helps, and usage data across the category reflects it: adoption is often lower than the pitch suggests. It works when personalisation operates within a structure the user still controls — surfacing a relevant metric in a slot they've left open — and works poorly when it reorders the whole layout on their behalf. If you build this, build the narrow version.
Anomaly detection. Load-bearing in exactly the domains where a human scanning a report would actually catch the anomaly too, just slower — fraud patterns, unusual spend, a metric that broke its normal range. The value is speed, not capability the human lacks. It becomes noise when it's applied to data too sparse or too volatile to have a meaningful "normal" in the first place, producing a stream of false positives that trains users to ignore the feature entirely within a few weeks — which is worse than not having it, because it also erodes trust in whatever alerts you add later.
A chat assistant over your product's data. Load-bearing when users currently have to dig through documentation, filters, or support tickets to answer questions your product's own data could answer directly, and when your underlying data is clean enough to answer well. Decoration when it's bolted onto a product where the assistant just paraphrases the same help docs a search bar would surface — which is a common outcome when a team builds the chat interface before checking whether their data layer can actually support good answers. If you're not sure which one you'd get, that uncertainty is itself the answer: fix the data layer first.
Across all four, the same pattern repeats: the load-bearing version answers a question the user was already asking, and the decoration version answers a question nobody was asking on the vendor's behalf. Before greenlighting an AI feature, ask your own team the same question you'd ask a vendor pitching one: what does the user currently do instead, how much do they dislike doing it, and would they notice immediately if the feature disappeared tomorrow. If nobody on the call can answer the third question with a concrete guess, the feature is going on the roadmap because it's on everyone else's roadmap, not because it solves anything.
This also changes how you should read a competitor's feature list. A competitor shipping nine AI features is not necessarily nine features ahead — it's often one team spread thin across nine shallow builds, none of which get the depth that makes a feature actually reliable. The product that ships two features people trust enough to depend on daily beats the product that ships nine people tried once during onboarding and never opened again.
Pick the one or two AI features that pass the filter for your specific product, and build those properly — with the evaluation, the edge-case handling, and the iteration that turns a demo into something people trust enough to rely on. That's a materially better use of a roadmap quarter than shipping five features at a shallow depth that each get tried once. The products that make "AI-powered" a real claim rather than a marketing line are usually the ones doing fewer things, better — not the ones with the longest feature list.
This is the same discipline as the build-vs-buy decision one level up: once you've decided a feature is worth building at all, build vs buy for AI features covers whether it's worth building yourself or buying the commodity version and spending your engineering time on the one that's actually differentiating.
If you're scoping which one or two features would move the needle for your product, that's the first conversation in our AI SaaS work.
The 838%-ROI claims in vendor pitches are built on assumptions that don't survive scrutiny. A grounded formula for AI automation ROI, and the costs that get left out.
A build vs buy framework for AI features — where buying wins, where a thin wrapper is genuinely correct, and the three conditions that justify building AI infrastructure yourself.
Where the money goes when you build an AI product MVP — engineering time, model inference, data preparation, evaluation — and which scope decisions move the total most.