The Problem
Abandonment rate is traditionally measured with a rule: did the caller hang up before being connected to an agent, or before some duration threshold elapsed? That rule is easy to compute and almost always wrong at the margins. A caller who hangs up one second after being connected because they got cut off isn't the same as a caller who waited 20 minutes and gave up — but a duration-based rule can label them identically. A chat customer who gets their answer and simply closes the tab without clicking "end chat" looks abandoned to a rule, even though the contact was a complete success.
You could keep patching the rule — add a minimum-connected-duration exception, add a "last agent message contained a resolution" exception, add a hang-up-reason-code exception — but every patch invites the next edge case. Rules are binary and context-free; genuine abandonment is a judgement call about what actually happened in the conversation.
The AI Use Case
Instead of relying solely on a fixed rule, pass the conversation transcript and the event timeline for each call or chat to an LLM and ask it to judge whether the contact was genuinely abandoned. This is LLM-as-a-judge applied to an operational metric rather than to AI agent quality — the model acts as a consistent, scalable analyst reviewing every contact the way a careful human QA reviewer would, if they had time to listen to all of them.
The judge is given:
- The transcript — what the caller/chatter actually said, and what the system or agent said back, up to the point of disconnection.
- The event timeline — queue entry/exit times, hold time, connection events, transfer events, and the disconnect reason code if one exists.
- A clear definition of "genuine abandonment" written in plain language — for example: "the contact ended because the customer gave up waiting or disengaged before their need was addressed, not because they were satisfied, resolved their own query, or were cut off by a technical fault."
The judge returns a verdict — genuinely abandoned, not abandoned (resolved/self-resolved/technical disconnect), or uncertain — plus a short reason, so every classification is auditable rather than a black-box label.
Worked Examples: Where the Rule and the Judge Disagree
- The "got what they needed" chat. A customer asks a self-service bot a question, gets a complete answer, and closes the browser tab without clicking "end chat." A duration rule logs this as abandoned (no formal close event). The judge reads the transcript, sees the question was answered and the customer said "thanks, that's exactly what I needed," and correctly classifies it as resolved, not abandoned.
- The split-second disconnect. A caller connects to an agent and the line drops after two seconds, before either party speaks. A rule using "connected = not abandoned" would miss this entirely; a rule using "under 10 seconds connected = abandoned" might wrongly blame the agent. The judge, seeing no meaningful exchange occurred and no technical fault code, flags it as uncertain and worth a human spot-check — more honest than a confident wrong guess either way.
- The patient-then-frustrated caller. A caller waits 12 minutes, finally connects, says one sentence, and hangs up abruptly mid-agent-response. A simple rule sees "connected" and stops counting it as abandonment — but the judge, reading the transcript, can recognise frustration and an incomplete resolution, and correctly flag this as a near-abandonment worth separate tracking even though it technically connected.
Call/chat ends
│
▼
┌─────────────────────────────┐
│ Gather transcript + event │
│ timeline for the contact │
└──────────────┬──────────────┘
▼
┌─────────────────────────────┐
│ LLM judge, given a plain- │
│ language abandonment │
│ definition, returns: │
│ verdict + reason │
└──────────────┬──────────────┘
▼
┌──────────┴──────────┐
▼ ▼
Genuine abandon Not abandoned / Uncertain
→ counted in rate → excluded, or routed to
human spot-check
The Value
- A more trustworthy metric — abandonment rate reflects what actually happened, not an approximation that both over- and under-counts at the edges.
- Fewer false "failures" in reporting — resolved self-service contacts stop dragging down a metric meant to measure genuine customer frustration.
- Richer insight than a single number — the judge's reasons can be aggregated to show why people genuinely abandon (too long a wait, wrong queue, frustration mid-call), which a rule-based count can never tell you.
- No endless rule patching — new edge cases are handled by the judge's reasoning rather than requiring a new exception clause every time one appears.
- An audit trail — every verdict comes with a reason, so a disputed classification can be reviewed rather than argued over blindly.
Keep a human in the loop at the edges. Don't let the judge silently overwrite the metric with no oversight. Sample its verdicts regularly against human judgement, route "uncertain" cases to a reviewer rather than guessing, and treat the plain-language definition of abandonment as a living document that gets refined as new edge cases are found — the same discipline as any eval set.
What You Need to Build It
- Access to the full transcript and event timeline per contact, tied together by a single contact ID.
- A written, plain-language definition of genuine abandonment that your operation agrees on — this is the rubric the judge grades against.
- An LLM judging step that returns a structured verdict and reason, run on a sample or on 100% of ended contacts depending on volume and cost.
- Calibration against human review — spot-check the judge's verdicts periodically, especially on "uncertain" cases, and refine the rubric when it's wrong.
- A reporting layer that separates true abandonment, resolved-but-uncleanly-closed, technical disconnects, and uncertain, rather than collapsing everything back into one number.
Explore More AI Use Cases
Part of a growing series of practical AI use cases for contact centres. Browse the full set on The AI Use Cases hub, and see the related containment-candidate detection use case, which applies the same LLM-as-a-judge pattern to a different metric. Got one you'd like covered? Get in touch.