Every chatbot dashboard shows conversations, and conversations always go up. That number tells you the widget is visible. It tells you nothing about whether it is doing the job, which for a website chatbot is answering questions correctly so they never become emails.
Here are the five numbers that do tell you, how to get them without a data team, and what to do when each one moves.
1. Ticket volume, before and after
The only outcome metric. Count support emails or tickets per week for the four weeks before launch and every week after. If the chatbot works, the line drops for the question types it covers and stays flat for the rest.
How to collect it. Your inbox or help desk already has it. Tag tickets by question type for a month before launch so you can see which types fell.
When it does not move. Either the chatbot is not on the pages where questions arise, or its sources do not cover the top question types. Check placement first, then sources. The ticket reduction guide covers both.
2. Coverage
The share of questions asked that the chatbot could answer from your sources, as opposed to saying the topic is not covered.
How to collect it. Read the conversation log weekly. Mark each question answered, thin, or not covered. Thirty minutes a week is enough for most small sites.
What good looks like. High for the question types you targeted, honest about everything else. A chatbot that never says "not covered" is more worrying than one that does, because it means it is guessing.
When it is low. Add the source that answers the missing questions. Coverage is the metric most directly under your control, and it is improved by content, not settings.
3. Handoff rate
The share of conversations that end with the chatbot offering your email, phone, or contact page.
How to collect it. Same weekly read of the log.
What good looks like. Not zero. A handoff is the chatbot doing its job on a question it should not answer. Very high means the sources are thin; zero means the handoff is misconfigured or the chatbot is improvising instead of admitting gaps.
When it moves. A sudden rise usually means a new type of question has appeared, often because you launched something and did not add its page. A fall to zero after a change means the support contacts got removed; check the configuration.
4. Correctness on a fixed test set
Twenty real questions, in customer phrasing, that you re-ask every month and score right, thin, or wrong.
How to collect it. Write the twenty once. Re-run them after any source change, model tier change, or policy update.
Why it matters. It catches regressions. A returns policy changed on the site but not re-added to the chatbot shows up here as a wrong answer before a customer finds it.
When it drops. Something changed in your sources or your site. Re-add the changed page; re-test. If a tier change caused it, change back. The model choice guide uses this same test for tier decisions.
5. The unanswered-questions list
Not a rate but a list: every question the chatbot could not answer, deduplicated, ordered by frequency.
How to collect it. From the weekly log read. Keep it in a plain document.
Why it is the most valuable output. It is a ranked list of what your customers want to know that your site does not tell them. Every entry is a page to write or a source to add, and writing those pages helps the chatbot, your search rankings, and the customers who never open the widget. The FAQ writing guide is where those pages usually end up.
Metrics to ignore
Total conversations. Visibility, not value.
Average conversation length. Longer is not better; a good answer ends a conversation in one turn.
Satisfaction thumbs. Too few people click them to mean anything on a small site, and the ones who do are the extremes.
Deflection percentage promised by a vendor. It depends on your question mix. Measure your own.
A monthly routine
- Pull ticket volume by type; compare with last month.
- Read the conversation log; score coverage and handoff rate.
- Re-run the twenty-question test.
- Update the unanswered-questions list; add the top three sources or write the top three pages.
- Re-test.
Two hours a month. Most of the gain arrives in the first two months, after which the routine is maintenance.
Frequently asked questions
Do I need analytics tooling for this? No. A conversation log, your inbox, and a document with twenty questions are enough for a site with a few hundred conversations a month.
What deflection rate should I expect? It depends entirely on how many of your questions have documented answers. Measure ticket volume by type before launch and the drop tells you your own number.
How do I know the chatbot is not making things up? The fixed test set, plus the habit of asking it something it cannot know once a week. It should say so. If it does not, review the sources for stale or contradictory pages before anything else.
