Short answer: Measure chatbot ROI by comparing the monthly cost of the tool against two things you can count yourself: the value of extra enquiries it captures, and the hours of staff time it removes. Both need your own numbers. Any vendor who tells you the return before seeing your traffic, close rate and margin is guessing.
Why should you distrust published ROI figures?
Because they are averages of businesses that are nothing like yours, usually collected from customers who volunteered a testimonial. A plumber with four hundred monthly visitors and a two thousand dollar job value lives in a completely different arithmetic than a shop with forty thousand visitors and a thirty dollar basket.
There is no honest industry number. There is only your number. The good news is that your number is not hard to calculate, and you can estimate it before you spend anything.
What inputs do you need before you start?
Gather six figures. Five come from tools you already have, and one is a judgement call.
Monthly website visitors. From your analytics. Use unique visitors, not pageviews.
Current enquiry rate. How many people contact you per month through the site, by any route: form, email, phone call attributed to the site. Divide by visitors to get a percentage.
Close rate. Of the enquiries you get, what share become paying customers? Your CRM knows, or your gut does.
Average order value or job value. What a customer is worth on the first transaction.
Gross margin. The share of that value you actually keep after delivering the work. This is the number people skip, and skipping it inflates everything.
Hours spent on repetitive enquiries. Estimate honestly, in hours per week, then multiply by your loaded hourly cost.
Write these six down before you look at any tool. They are your baseline, and the baseline is the whole point.
How do you calculate the revenue side?
The revenue side answers one question: how many additional qualified enquiries would the widget need to produce to pay for itself?
Work backwards. Take the monthly cost of the tool. Divide it by the gross margin on one average customer. That gives you the number of extra customers needed to break even. Then divide that by your close rate to get the number of extra enquiries needed.
For example, if a plan costs thirty-five dollars, your average job is worth four hundred dollars, your gross margin is forty percent, and you close one in four enquiries, then one extra customer produces one hundred sixty dollars of margin. You need roughly a quarter of a customer per month to break even, which is about one extra enquiry per month. That is a low bar, and you should say so out loud, because the interesting question is not whether it breaks even. It is how far past break-even it goes.
Run the same arithmetic with your own six numbers. If the break-even enquiry count comes out higher than the total enquiries you get today, the honest conclusion is that this is not a revenue play for you, and you should evaluate it purely on time saved. How much does an AI receptionist cost covers the pricing side in more detail.
How do you calculate the time side?
This half is more reliable because it does not depend on forecasting behaviour.
Take your hours per week on repetitive enquiries. Estimate what share of those questions have a fixed, published answer. Multiply the hours by that share, then by four point three to get monthly hours, then by your loaded hourly cost.
Two cautions. First, saved time only becomes money if it gets redeployed into something that earns or if it prevents a hire. An owner who saves five hours and spends them on email has saved nothing measurable. Second, use loaded cost, not salary divided by hours. Include the overheads.
How do you set up measurement before you launch?
This is where most evaluations fall apart. People install the tool, feel busier, and declare victory based on vibes. Do this instead.
Freeze a baseline period. Record your six inputs for the four weeks before launch. Screenshot the analytics. Export the enquiry count.
Tag the source. Every enquiry that arrives through chat should be identifiable as such. In Clerkzo, chat-captured leads land in the leads inbox with a New to Contacted to Qualified to Won or Lost pipeline, so the source and the outcome stay attached to each other. That link between source and outcome is the whole measurement.
Decide your window in advance. Sixty to ninety days. Shorter windows are dominated by seasonality and by the fact that you were still tuning the thing in week one.
Watch for cannibalisation. Some chat enquiries would have arrived by form anyway. Compare total enquiries across the whole period, not just chat enquiries. If chat is up two hundred and forms are down one hundred and eighty, the net is twenty, and twenty is your real number.
That last point is the single most common way businesses fool themselves. Be strict about it.
What should you actually track month to month?
Five metrics, no more.
Conversations started. Volume of engagement. Tells you whether the launcher is visible and inviting.
Answer rate. Share of conversations resolved without a handoff. Tells you whether your content is good enough.
Leads captured. Conversations that produced contact details and intent.
Qualified and won. From the pipeline. This is the only revenue-relevant number.
Total enquiries, all channels. Your cannibalisation check.
Add thumbs up and thumbs down ratings as a quality signal. A falling rating usually predicts a falling lead rate a few weeks later.
What are the costs people forget?
Be complete or the calculation is worthless.
The subscription itself, which for Clerkzo is flat monthly rather than per conversation, so it does not spike when volume does. Setup time, which is short but not zero, though done-for-you setup shifts that cost. Ongoing maintenance, realistically an hour or two a month reviewing transcripts and correcting answers. Content work, if your site turns out to be too vague to train on. And the opportunity cost of the follow-up: captured leads that nobody contacts have negative ROI, because you paid to generate them and then let them go cold.
If you are comparing options, the comparison page lays out how the alternatives are priced, which changes the cost side of this equation more than most people expect.
When does the calculation say no?
It says no when your traffic is too low for a percentage improvement to mean anything, when your margin per customer is thin enough that break-even requires implausible volume, or when your enquiries are genuinely all bespoke and none of them have publishable answers. Those are real cases. Run the arithmetic and let it tell you.
Frequently asked questions
How long before I can judge the result?
Sixty days minimum, ninety is better. The first two weeks are configuration, not performance.
Should I count time saved as revenue?
Only if the time was redeployed to revenue work or it avoided a hire. Otherwise track it separately as a quality-of-life benefit, which is real but not financial.
What if my close rate is different for chat leads?
Track it separately once you have thirty or so. Chat leads often arrive earlier in the buying process, which can mean a lower close rate and a longer cycle. That is not a failure, but your model needs to reflect it.
Is a free plan enough to measure with?
It is enough to test answer quality and see the questions people ask. For revenue measurement you need the pipeline attached, so run the free plan first and measure properly on a paid one. See pricing.