Measuring deflection honestly
What "resolved" means on our dashboard, why the recommendation lift is hidden for a while, and why we lead with deflection anyway.
Every conversation ends in one of five outcomes, and the dashboard says which and why, next to the number:
- Resolved: ended without an unanswered signal on its last turn: no error, no offer to escalate, the visitor's last message answered. A proxy for deflection, not a survey.
- Unresolved: the last turn ended with an escalation offer, an error, an exhausted tool budget, or the visitor left without a reply.
- Escalated: a person was asked for.
- Capped: the session hit its turn or token ceiling.
- Empty: opened, nothing asked.
Resolution rate is resolved divided by everything that ended and was not empty. Escalation rate uses the same denominator. Actions completed are tool calls that succeeded; blocked ones are counted separately and shown, because a rising blocked rate is the earliest sign someone is trying something.
We lead with these for a reason. A product with 800 end users produces enough conversations to measure resolution, escalation and actions completed within two weeks. Recommendation lift against a 5 % holdout needs far more volume and can take a quarter to reach significance, so the dashboard hides the headline number until it does and shows the confidence interval instead. A fabricated-looking uplift figure in front of a technical founder costs the account; an honest "not yet significant, here is the interval" does not.
Two more things we do not do. We do not count an opened widget as a conversation, because it inflates every rate. And we do not call a conversation resolved because the model said "I hope that helps". The outcome is derived from what happened on the last turn, in code, the same way every time.