Resolution QA
A regular sample of closed support conversations gets graded against a clear rubric — AI-handled and human-handled alike — so a closed ticket and a good answer stop being treated as the same thing.
Think of a teacher who does not just check that every student handed something in, but actually marks a sample of the work — and keeps a list of the mistakes the whole class keeps making. Resolution QA does that for your closed tickets: closed is attendance; graded is learning.
The problem, in plain words
A ticket marked resolved just means somebody closed it. Maybe the answer was right. Maybe it was polite, wrong, and the customer gave up rather than argue. Maybe an agent — human or AI — closed it to hit a number. Nobody has time to re-read closed tickets, so quality gets measured by the survey lottery: the handful of customers annoyed or delighted enough to answer. Meanwhile the actual quality of your resolutions — the thing all your metrics claim to track — goes unmeasured.
What we set up
We sample your resolved threads — not all of them, but a deliberate slice, weighted so high-volume topics and high-risk cases (refunds, cancellations, anything touching money or accounts) get checked more often. Each sampled thread is scored against a rubric (a written marking sheet: was the answer correct, complete, properly sourced, handled the way your standards require) by an LLM-as-judge (an AI grader that follows that marking sheet). And because a grader can drift too, humans regularly re-score a portion of the same threads and compare notes; how closely the AI grader agrees with human graders is itself a tracked number. Low-scoring resolutions go to a review queue for a person to look at. High-scoring ones become candidate gold cases — model answers worth adding to your test set.
How it works, step by step
- Threads are marked resolved
Handled by your AI agent or by your team — both go into the same sampling pool.
- A weighted sample is drawn
High-volume topics and high-risk cases get sampled more heavily. Nobody reads everything; the point is a fair, honest slice.
- Each sample is graded
An AI grader scores the thread against a written rubric: correct, complete, sourced, and handled to your standards.
- The grader itself gets checked
Humans periodically re-score the same threads. Agreement between AI and human graders is tracked — if it slips, the grader gets recalibrated.
- Low scores go to review
Weak resolutions surface for a person to examine — including tickets everyone thought were fine.
- High scores become examples
The best resolutions are promoted as candidate gold cases for your test set, so future versions are held to them.
What changes for you
Before: quality is whatever the closure rate and a thin trickle of survey responses say it is, and bad answers hide behind the word resolved. After: resolution quality is a calibrated, tracked number; weak answers surface even when the customer never complained; and the class's common mistakes — the patterns the teacher keeps marking — show your team exactly what to coach. What it won't do: it won't re-check every ticket — it grades a weighted sample, which is enough to see the patterns but will always let some individual misses through.