AI & Quality • October 2026

How to Build an Eval Set for a Self-Service IVR Application

If evals are the mechanism that keeps a self-service bot excellent, the eval set is the asset that makes evals possible. It's also the part most teams get vague about. This is the practical, step-by-step guide to building one for a contact centre IVR or voice agent — where the scenarios come from, the right mix and size, how to write expected outcomes, and the voice-specific cases nobody warns you about.

A quick framing before the steps. If you want the why — why evals matter, the eval types, and the dev/CI/production lifecycle — that's covered in using evals to build high-performing agents. This article is narrower and more hands-on: how to actually construct the eval set itself for a self-service IVR. The eval set is your "golden dataset" — the fixed collection of test cases you run against the application before every meaningful change.[1]

The steps
  1. What an IVR eval set actually is
  2. Step 1 — Decide what you're measuring
  3. Step 2 — Source scenarios from real data
  4. Step 3 — Get the mix and size right
  5. Step 4 — Write the expected outcome for each case
  6. Step 5 — Add the voice-specific cases
  7. Step 6 — Choose a grader per case
  8. Step 7 — Label, validate, and version it
  9. Step 8 — Keep it alive
  10. Sources

What an IVR Eval Set Actually Is

An eval set is a curated, expert-validated collection of representative tasks, each paired with an expected outcome and the criteria for grading it.[2] For a self-service IVR, a single eval case is essentially: a caller scenario → what the system should do → how we'll judge whether it did.

Crucially, in a conversational IVR a "scenario" is rarely one line — it's often a whole multi-turn exchange, because errors compound across turns in a way single-turn tests can't catch. The whole point of the set is to make the application's behaviour measurable and regressions catchable before customers hit them.

Step 1 — Decide What You're Measuring

Before writing a single scenario, decide what "good" means for your IVR, because that determines what each case needs to check. For a contact-centre self-service application, the dimensions that matter most are:

Write these down first. Every scenario you create should test one or more of them, and this list becomes the backbone of your grading.

Step 2 — Source Scenarios from Real Data

This is the single most important principle, and every serious guide agrees on it: build the set from real production traces first; CSV imports and synthetic generation come second.[1] An eval set invented in a meeting room tests the IVR you imagined, not the one callers actually use.

Where to pull scenarios from

If you're pre-launch and have no production data yet, seed the set from the design itself — your intent list, your flows, and the test plan — then replace those synthetic cases with real ones the moment traffic starts.

Rule of thumb: if a scenario isn't grounded in something a real caller did (or plausibly will do), question whether it belongs. Synthetic cases are fine for coverage gaps, but a set dominated by invented examples gives you confident, precise, wrong numbers.

Step 3 — Get the Mix and Size Right

The mix

A good eval set deliberately spans more than the happy path. A widely-cited starting split is roughly 60% happy path, 30% edge cases, 10% adversarial.[4] For an IVR, read that as:

The size

You don't need thousands on day one. Match the size to the job:[5]

PurposeRough sizeSource
Explore a single issue~10 itemsOpenRouter[5]
A full regression set100–1,000 itemsOpenRouter[5]
Per distinct failure mode100–200 labelled examplesHamel Husain[6]
Per change, with an LLM judge500–5,000 exampleslearnwithparam[7]

The practical path: start with ~20 strong, real scenarios to prove the mechanism, grow to 100+ for a dependable regression set, and let it keep expanding as new intents and failures appear. The right number ultimately depends on the metric's variance and the smallest difference you need to detect.[5]

Step 4 — Write the Expected Outcome for Each Case

This is the work that makes an eval set "golden": for each scenario, a human decides what the ideal outcome is.[4] For an IVR, don't pin it to an exact wording — there are many valid phrasings. Pin it to behaviour and boundaries: the intent that should be recognised, the action/route taken, what the response must include, and what it must never include.

# One eval case for a bank card-servicing IVR
id: card-lost-happy-001
source: production_transcript        # grounded in a real call
category: happy_path
caller_turns:
  - "I need to report my card lost"
context: { customer_verified: true, cards_on_file: 1 }
expected:
  intent: report_card_lost
  action: freeze_card
  must_include: ["card ending", "confirmation / case reference"]
  must_not:     ["full card number", "financial advice"]
  escalate: false
grader: [rule, llm_judge]           # rule for the hard checks; judge for tone/faithfulness

For multi-turn scenarios, the expected outcome spans the whole exchange — did the conversation drive to the goal, turn after turn, without looping or losing context? A faithful set captures the trajectory, not just the final line.

Step 5 — Add the Voice-Specific Cases

An IVR isn't a chatbot with a phone number. Voice adds failure modes a text eval set would miss entirely, and leaving them out is the most common gap in IVR eval sets. Make sure the set includes:

If you generate your test scenarios from a flow design, the auto-generated route list and the no-input/no-match cases are a strong seed for exactly these — the IVR Design Tool produces both, and they map directly onto eval cases.

Step 6 — Choose a Grader per Case

Each case needs a way to score the output. Use a mix — different checks suit different cases:

Calibrate the judge before you trust it. An LLM judge is powerful but imperfect — check it against human-labelled examples first, keep humans in the loop for high-stakes categories, and re-check periodically. An un-calibrated judge produces confident, precise numbers that may be wrong.

Step 7 — Label, Validate, and Version It

A pile of scenarios isn't a golden dataset until it's validated and version-controlled. The maintenance disciplines that matter:[1]

Step 8 — Keep It Alive

An eval set is never finished — a good one mirrors production reality and evolves with the product.[8] Build the feedback loop in:

Where to Start

  1. List what "good" means for your IVR (Step 1) — intent, task completion, routing, faithfulness, policy, escalation.
  2. Pull ~20 real scenarios from your top call drivers and worst known failures, and write the ideal outcome for each.
  3. Cover the mix — mostly happy path, a solid slice of edge cases, a few adversarial.
  4. Add the voice-specific cases — no-input/no-match, barge-in, disfluent speech, spoken data capture.
  5. Assign graders — rules for the clear-cut checks, an LLM judge for the nuanced ones, and calibrate the judge against human labels.
  6. Validate, version, and run it on every change — then grow it toward 100+ and let production failures keep feeding it.

Do that and you move from shipping IVR changes on hope to shipping them against a measured, defensible standard — one that gets stronger every time something goes wrong.

Designing the flows and test scenarios your eval set will be built from? The free IVR Design Tool generates a route-by-route test plan and no-input/no-match cases that make a natural seed for your first eval set.

Sources

This article draws on published guidance on building evaluation datasets. Figures (sizes, mixes) are rules of thumb reported by each source and vary by use case; treat them as directional. Content was summarised and rephrased for compliance. This is independent guidance, not affiliated with the organisations cited.

  1. Langfuse — "Golden dataset evaluation: build and maintain LLM test sets"
  2. Lyzr — "Golden Datasets for AI Agent Evaluation"
  3. Digital Applied — "Building an AI Agent Evaluation Pipeline"
  4. Beam.ai — "How to Actually Run Evals on Your AI Agents" (60/30/10 mix; write the ideal output)
  5. OpenRouter — "Building a Golden Eval Dataset from Production Traffic" (sizing guidance)
  6. Hamel Husain — "How many examples do I need for an eval?" (100–200 per failure mode)
  7. learnwithparam — "LLM-as-a-judge production evaluation framework"
  8. Maxim — "Building a Golden Dataset for AI Evaluation"