The real return on AI, and why we will not give you a percentage
The real ROI of AI implementation. Three kinds of return exist, only one is countable before you build, and we commit to an outcome instead of a percentage.
Published: 2026-08-19 · Author: Ahmed Heshmat · 9 min read
In short: Every ROI figure in this industry is back-calculated fiction. Here are the three kinds of return that actually exist, only one of which you can count, and what we commit to instead of a number we would be inventing.
Key takeaways
- There are three kinds of return. Money that stops leaving, money that was already leaking, and hours given back. They are not equally real, and only the first is straightforwardly countable.
- The largest return in most small business deployments is money that was already leaking, which nobody was measuring. The system's first job is to make the leak visible.
- Hours saved multiplied by an hourly rate is arithmetic, not return. Recovered hours convert to money only if the person spends them on something that earns.
- A vendor who quotes you a percentage before the build is guessing. Ask for a committed observable outcome instead, and a date.
The number we refuse to give you
Somewhere in this industry there is a slide with "340% ROI" on it. Nobody computed that. Someone estimated hours saved, multiplied by a rate they chose, divided by a fee, and rounded to something that felt persuasive. We have seen versions of that slide in proposals we were bidding against, and we have been asked more than once why ours does not have one.
Because if we sell you 340% and deliver 40%, we have lost you in month five. If we tell you that you will stop losing leads after 6pm, and then you stop losing leads after 6pm, you renew. The fake number is not just dishonest. It is bad business, and it is the reason so many of these engagements end in a quiet non renewal that both sides pretend was about budget. It is also most of [what the AI consulting boom gets wrong](/blog/what-the-ai-consulting-boom-gets-wrong).
Here is what we think is actually true about return, based on the builds we have run.
Return type one: money that stops leaving
The countable one. Something was going out of the bank account and now it is not.
An answering service you cancel. A software seat nobody needed once two systems started talking. A contractor invoice for work the system now does. This is the cleanest kind of return, it shows up in your own accounting without anyone helping you interpret it, and in small business AI work it is usually the smallest of the three.
If a vendor's entire case rests on this category, the case is easy to verify and you do not need us to explain it. Pull the two bills and compare.
Return type two: money that was already leaking
The biggest one, and the one nobody can quantify before the build, because the leak is not being measured.
We instrumented the phone lines for a property management operation and its brokerage arm. Clean sample, 277 real calls, deduplicated. Twenty seven percent came in outside business hours or on the weekend. On the brokerage line, a routable callback was captured from roughly seven of every ten calls. Two thirds of the property management calls were resolved in the conversation without a staff member touching them. Zero calls missed.
The important thing about those figures is not how good they are. It is that before the system existed, that business could not have produced a single one of them. There was no record of the weekend calls that went to voicemail and never got returned. Nobody was counting the leads that dissolved on a Saturday, because a lead that dissolves does not generate paperwork.
So when someone asks us for the ROI of a voice deployment before it is built, the honest answer is that we do not know and neither do they, because the denominator has never been measured. What we can tell you is that when we do measure it, the number is consistently larger than the operator guessed. Every time so far.
This has a practical consequence. If the return case for a project depends entirely on stopping a leak nobody has sized, the correct first project is often instrumentation rather than automation. Spend a month counting. Then decide. That is a smaller invoice and a better decision, and we would rather sell it to you than sell you a build against a number we both made up.
Return type three: hours given back
Where all the fake math lives.
The formula is always the same. Twenty five hours a week saved, times fifty dollars an hour, times fifty two weeks, equals sixty five thousand dollars a year. It is arithmetic, and it is not wrong exactly, but it quietly assumes something enormous: that the recovered hour gets spent on something that earns.
It often does not. If your operations manager gets six hours back and spends them answering more email, the return is zero. Real, and zero. The hours only convert when there is a specific higher value thing the person has been unable to get to, and someone has decided in advance that this is what the time is for.
So the question we ask before a build is not how many hours will this save. It is: when this work disappears, what will that person do instead, specifically? If the answer is a named activity that makes money or prevents a loss, the hours are real return. If the answer is a shrug or the word "capacity," the hours are a feeling.
This is also the whole reason we do not build systems designed to remove people. The return we are actually selling is a person freed to do judgment work, relationships, and the things that were being crowded out. If the plan is to delete the person instead, you are not buying return. You are buying a severance calculation, and you do not need a consultant for that.
The three questions to run before you spend anything
Run these on us, and on anyone else you are considering.
1. What is this costing us right now, and do we know or are we guessing? If guessing, the first build is measurement. Anyone who lets you skip this step is happy for you to be unable to evaluate them later, which is convenient for them.
2. If this work disappeared tomorrow, what would the person doing it do instead? Name the activity. If nobody can, the hours will not convert and the project is smaller than it looks.
3. What has to be true in ninety days for us to call this a success? Written down before the build starts. Not a percentage. A sentence a bookkeeper could verify.
That third question is the one that separates a real engagement from an expensive experiment. If nobody can write the sentence in advance, then in ninety days the evaluation becomes a vibe, and vibes renew badly.
What we commit to instead
We do not put a return percentage in a proposal. We put in an observable outcome and a date.
Something along the lines of: within ninety days, every inbound call is answered and logged, no call reaches voicemail, and your team sees a written summary of each one within a minute of it ending. That is checkable by anyone in the business without a spreadsheet or an interpretation. It is either happening or it is not.
Then the arithmetic is yours to run, on your own numbers, which are the only ones that matter. You know what a closed deal is worth to you. We do not, and neither does the vendor who put 340% on a slide.
If the committed outcome does not happen, you should fire us. We would rather say that in writing than defend an invented percentage in month five.
Where this goes wrong on our side
Honesty requires the other half. The engagements of ours that have underdelivered share one cause, and it is not the technology.
They were scoped without full visibility into the operation. Someone described a process, we built against the description, and then the real version turned out to contain four exceptions that everyone in the business handles by reflex and nobody thought to mention. Those exceptions are not edge cases to the operation. They are ordinary Tuesdays. And when they surface mid build, the project gets longer, the return arrives later than anyone planned, and the client is right to be annoyed.
The fix is not a better estimate. It is refusing to price a build against a conversation, and paying proper attention to how the work actually happens before committing to what the system will do. That is the entire argument for [mapping the operation first](/blog/audit-build-operate-model), and it is also why the cases a system gets wrong deserve more design attention than the cases it gets right, which is [where most pilots die](/blog/why-most-ai-pilots-die-after-the-demo).
What to do with this
- Ask any vendor which of the three return types their case rests on. If they cannot separate them, they have not done the thinking.
- Refuse percentages offered before the build. They are derived from assumptions you have not agreed to.
- If the leak has never been measured, buy measurement first. It is cheaper and it makes the next decision real.
- Name what the recovered hours are for, in advance and in writing. Unassigned time does not convert.
- Get one observable outcome and a date into the agreement. Then you can evaluate the work without needing anyone's spreadsheet.