Use case

Resolution QA

A regular sample of closed support conversations gets graded against a clear rubric — AI-handled and human-handled alike — so a closed ticket and a good answer stop being treated as the same thing.

The short version

Think of a teacher who does not just check that every student handed something in, but actually marks a sample of the work — and keeps a list of the mistakes the whole class keeps making. Resolution QA does that for your closed tickets: closed is attendance; graded is learning.

How it flows
Threads marked resolvedWeighted sample drawnGraded against a rubricHumans check the graderWeak flagged, best promoted

The problem, in plain words

A ticket marked resolved just means somebody closed it. Maybe the answer was right. Maybe it was polite, wrong, and the customer gave up rather than argue. Maybe an agent — human or AI — closed it to hit a number. Nobody has time to re-read closed tickets, so quality gets measured by the survey lottery: the handful of customers annoyed or delighted enough to answer. Meanwhile the actual quality of your resolutions — the thing all your metrics claim to track — goes unmeasured.

What we set up

We sample your resolved threads — not all of them, but a deliberate slice, weighted so high-volume topics and high-risk cases (refunds, cancellations, anything touching money or accounts) get checked more often. Each sampled thread is scored against a rubric (a written marking sheet: was the answer correct, complete, properly sourced, handled the way your standards require) by an LLM-as-judge (an AI grader that follows that marking sheet). And because a grader can drift too, humans regularly re-score a portion of the same threads and compare notes; how closely the AI grader agrees with human graders is itself a tracked number. Low-scoring resolutions go to a review queue for a person to look at. High-scoring ones become candidate gold cases — model answers worth adding to your test set.

How it works, step by step

  1. Threads are marked resolved

    Handled by your AI agent or by your team — both go into the same sampling pool.

  2. A weighted sample is drawn

    High-volume topics and high-risk cases get sampled more heavily. Nobody reads everything; the point is a fair, honest slice.

  3. Each sample is graded

    An AI grader scores the thread against a written rubric: correct, complete, sourced, and handled to your standards.

  4. The grader itself gets checked

    Humans periodically re-score the same threads. Agreement between AI and human graders is tracked — if it slips, the grader gets recalibrated.

  5. Low scores go to review

    Weak resolutions surface for a person to examine — including tickets everyone thought were fine.

  6. High scores become examples

    The best resolutions are promoted as candidate gold cases for your test set, so future versions are held to them.

What changes for you

Before: quality is whatever the closure rate and a thin trickle of survey responses say it is, and bad answers hide behind the word resolved. After: resolution quality is a calibrated, tracked number; weak answers surface even when the customer never complained; and the class's common mistakes — the patterns the teacher keeps marking — show your team exactly what to coach. What it won't do: it won't re-check every ticket — it grades a weighted sample, which is enough to see the patterns but will always let some individual misses through.