AI Use Cases • Use Case 09 • October 2026

LLM-as-a-Judge for Genuine Call Abandonment

Abandonment rate is one of the most quoted contact centre metrics — and one of the least trustworthy, because it's almost always calculated from a blunt rule. This use case replaces the rule with an AI judge that reads what actually happened in the call or chat and decides whether it was a genuinely abandoned contact.

Problem
Fixed rules misclassify edge-case abandonment
AI capability
LLM-as-a-judge over conversation + events
Outcome
A trustworthy, defensible abandonment rate

The Problem

Abandonment rate is traditionally measured with a rule: did the caller hang up before being connected to an agent, or before some duration threshold elapsed? That rule is easy to compute and almost always wrong at the margins. A caller who hangs up one second after being connected because they got cut off isn't the same as a caller who waited 20 minutes and gave up — but a duration-based rule can label them identically. A chat customer who gets their answer and simply closes the tab without clicking "end chat" looks abandoned to a rule, even though the contact was a complete success.

You could keep patching the rule — add a minimum-connected-duration exception, add a "last agent message contained a resolution" exception, add a hang-up-reason-code exception — but every patch invites the next edge case. Rules are binary and context-free; genuine abandonment is a judgement call about what actually happened in the conversation.

The AI Use Case

Instead of relying solely on a fixed rule, pass the conversation transcript and the event timeline for each call or chat to an LLM and ask it to judge whether the contact was genuinely abandoned. This is LLM-as-a-judge applied to an operational metric rather than to AI agent quality — the model acts as a consistent, scalable analyst reviewing every contact the way a careful human QA reviewer would, if they had time to listen to all of them.

The judge is given:

The judge returns a verdict — genuinely abandoned, not abandoned (resolved/self-resolved/technical disconnect), or uncertain — plus a short reason, so every classification is auditable rather than a black-box label.

Worked Examples: Where the Rule and the Judge Disagree

  1. The "got what they needed" chat. A customer asks a self-service bot a question, gets a complete answer, and closes the browser tab without clicking "end chat." A duration rule logs this as abandoned (no formal close event). The judge reads the transcript, sees the question was answered and the customer said "thanks, that's exactly what I needed," and correctly classifies it as resolved, not abandoned.
  2. The split-second disconnect. A caller connects to an agent and the line drops after two seconds, before either party speaks. A rule using "connected = not abandoned" would miss this entirely; a rule using "under 10 seconds connected = abandoned" might wrongly blame the agent. The judge, seeing no meaningful exchange occurred and no technical fault code, flags it as uncertain and worth a human spot-check — more honest than a confident wrong guess either way.
  3. The patient-then-frustrated caller. A caller waits 12 minutes, finally connects, says one sentence, and hangs up abruptly mid-agent-response. A simple rule sees "connected" and stops counting it as abandonment — but the judge, reading the transcript, can recognise frustration and an incomplete resolution, and correctly flag this as a near-abandonment worth separate tracking even though it technically connected.
  Call/chat ends
        │
        ▼
  ┌─────────────────────────────┐
  │  Gather transcript + event   │
  │  timeline for the contact    │
  └──────────────┬──────────────┘
                 ▼
  ┌─────────────────────────────┐
  │  LLM judge, given a plain-    │
  │  language abandonment         │
  │  definition, returns:         │
  │  verdict + reason              │
  └──────────────┬──────────────┘
                 ▼
      ┌──────────┴──────────┐
      ▼                     ▼
  Genuine abandon      Not abandoned / Uncertain
  → counted in rate    → excluded, or routed to
                          human spot-check

The Value

Keep a human in the loop at the edges. Don't let the judge silently overwrite the metric with no oversight. Sample its verdicts regularly against human judgement, route "uncertain" cases to a reviewer rather than guessing, and treat the plain-language definition of abandonment as a living document that gets refined as new edge cases are found — the same discipline as any eval set.

What You Need to Build It

Explore More AI Use Cases

Part of a growing series of practical AI use cases for contact centres. Browse the full set on The AI Use Cases hub, and see the related containment-candidate detection use case, which applies the same LLM-as-a-judge pattern to a different metric. Got one you'd like covered? Get in touch.