A quick framing before the steps. If you want the why — why evals matter, the eval types, and the dev/CI/production lifecycle — that's covered in using evals to build high-performing agents. This article is narrower and more hands-on: how to actually construct the eval set itself for a self-service IVR. The eval set is your "golden dataset" — the fixed collection of test cases you run against the application before every meaningful change.[1]
- What an IVR eval set actually is
- Step 1 — Decide what you're measuring
- Step 2 — Source scenarios from real data
- Step 3 — Get the mix and size right
- Step 4 — Write the expected outcome for each case
- Step 5 — Add the voice-specific cases
- Step 6 — Choose a grader per case
- Step 7 — Label, validate, and version it
- Step 8 — Keep it alive
- Sources
What an IVR Eval Set Actually Is
An eval set is a curated, expert-validated collection of representative tasks, each paired with an expected outcome and the criteria for grading it.[2] For a self-service IVR, a single eval case is essentially: a caller scenario → what the system should do → how we'll judge whether it did.
Crucially, in a conversational IVR a "scenario" is rarely one line — it's often a whole multi-turn exchange, because errors compound across turns in a way single-turn tests can't catch. The whole point of the set is to make the application's behaviour measurable and regressions catchable before customers hit them.
Step 1 — Decide What You're Measuring
Before writing a single scenario, decide what "good" means for your IVR, because that determines what each case needs to check. For a contact-centre self-service application, the dimensions that matter most are:
- Intent recognition — did it understand what the caller wanted?
- Task completion — did it actually resolve the request, not just sound like it did?
- Correct routing / disambiguation — did it reach the right place, and clarify (not guess) when ambiguous? (See disambiguation in self-service.)
- Faithfulness — for any answer, is it grounded in approved content rather than invented?
- Policy & safety — did it stay in bounds (verify identity before disclosure, never give advice it shouldn't)?
- Correct escalation — did it hand off cleanly when it should, and only when it should?
Write these down first. Every scenario you create should test one or more of them, and this list becomes the backbone of your grading.
Step 2 — Source Scenarios from Real Data
This is the single most important principle, and every serious guide agrees on it: build the set from real production traces first; CSV imports and synthetic generation come second.[1] An eval set invented in a meeting room tests the IVR you imagined, not the one callers actually use.
Where to pull scenarios from
- Your top call drivers — the intents that make up the bulk of volume. These must be rock-solid, so they belong in the set first.
- Real transcripts and utterances — the actual messy ways callers phrase things, including the "umms", false starts, and roundabout requests.
- Known failures — the conversations where the bot (or a human) got it wrong. The best golden datasets are drawn from real failures.[3] Every production miss should become a permanent eval so it can never silently regress again.
- Edge cases from the logs — the long tail: multi-intent requests, mid-call changes of mind, barge-in, callers who won't follow the menu.
- Compliance & vulnerable-customer situations — anything where getting it wrong carries real risk.
If you're pre-launch and have no production data yet, seed the set from the design itself — your intent list, your flows, and the test plan — then replace those synthetic cases with real ones the moment traffic starts.
Rule of thumb: if a scenario isn't grounded in something a real caller did (or plausibly will do), question whether it belongs. Synthetic cases are fine for coverage gaps, but a set dominated by invented examples gives you confident, precise, wrong numbers.
Step 3 — Get the Mix and Size Right
The mix
A good eval set deliberately spans more than the happy path. A widely-cited starting split is roughly 60% happy path, 30% edge cases, 10% adversarial.[4] For an IVR, read that as:
- ~60% happy path — your common intents, phrased normally, resolving cleanly.
- ~30% edge cases — messy/ambiguous phrasing, multi-intent requests, no-input/no-match, out-of-scope asks, mid-call corrections.
- ~10% adversarial — attempts to confuse or manipulate the system, requests for restricted actions, social engineering.
The size
You don't need thousands on day one. Match the size to the job:[5]
| Purpose | Rough size | Source |
|---|---|---|
| Explore a single issue | ~10 items | OpenRouter[5] |
| A full regression set | 100–1,000 items | OpenRouter[5] |
| Per distinct failure mode | 100–200 labelled examples | Hamel Husain[6] |
| Per change, with an LLM judge | 500–5,000 examples | learnwithparam[7] |
The practical path: start with ~20 strong, real scenarios to prove the mechanism, grow to 100+ for a dependable regression set, and let it keep expanding as new intents and failures appear. The right number ultimately depends on the metric's variance and the smallest difference you need to detect.[5]
Step 4 — Write the Expected Outcome for Each Case
This is the work that makes an eval set "golden": for each scenario, a human decides what the ideal outcome is.[4] For an IVR, don't pin it to an exact wording — there are many valid phrasings. Pin it to behaviour and boundaries: the intent that should be recognised, the action/route taken, what the response must include, and what it must never include.
# One eval case for a bank card-servicing IVR id: card-lost-happy-001 source: production_transcript # grounded in a real call category: happy_path caller_turns: - "I need to report my card lost" context: { customer_verified: true, cards_on_file: 1 } expected: intent: report_card_lost action: freeze_card must_include: ["card ending", "confirmation / case reference"] must_not: ["full card number", "financial advice"] escalate: false grader: [rule, llm_judge] # rule for the hard checks; judge for tone/faithfulness
For multi-turn scenarios, the expected outcome spans the whole exchange — did the conversation drive to the goal, turn after turn, without looping or losing context? A faithful set captures the trajectory, not just the final line.
Step 5 — Add the Voice-Specific Cases
An IVR isn't a chatbot with a phone number. Voice adds failure modes a text eval set would miss entirely, and leaving them out is the most common gap in IVR eval sets. Make sure the set includes:
- Speech-recognition ambiguity — accents, background noise, homophones, numbers and spellings misheard. Voice has ambiguity in what was said on top of what was meant.
- No-input and no-match — the caller says nothing, or something the system can't parse. Test that retries and fallbacks behave, and that it escalates rather than looping forever.
- Barge-in — the caller talks over the prompt.
- Disfluent, natural speech — "umm, yeah, I think I want to… actually, can you…". A good agentic IVR should handle this; your set should prove it.
- DTMF vs speech — callers who key presses instead of speaking, or switch between the two.
- Spoken data capture — reading out card numbers, dates, postcodes; the classic source of errors.
- Latency sensitivity — silence kills a voice call, so where relevant, assert on acceptable response time, not just correctness.
If you generate your test scenarios from a flow design, the auto-generated route list and the no-input/no-match cases are a strong seed for exactly these — the IVR Design Tool produces both, and they map directly onto eval cases.
Step 6 — Choose a Grader per Case
Each case needs a way to score the output. Use a mix — different checks suit different cases:
- Rule-based / deterministic — exact checks: was the right intent recognised? the right tool/route taken? forbidden strings absent? Cheap, fast, unambiguous — use for anything with a clear right answer.
- LLM-as-a-judge — a model grades against a rubric (faithful? on-policy? right tone? did it clarify rather than guess?). This is what makes grading open-ended language affordable at scale, and it's practical to run on hundreds to thousands of cases per change.[7]
- Human review — a sampled slice, used to calibrate the judge and catch subtle domain issues. This is your source of truth.
Calibrate the judge before you trust it. An LLM judge is powerful but imperfect — check it against human-labelled examples first, keep humans in the loop for high-stakes categories, and re-check periodically. An un-calibrated judge produces confident, precise numbers that may be wrong.
Step 7 — Label, Validate, and Version It
A pile of scenarios isn't a golden dataset until it's validated and version-controlled. The maintenance disciplines that matter:[1]
- Expert labelling — have someone who knows the domain and the compliance rules confirm each expected outcome. Agree the gold standard with business and compliance stakeholders for the core cases.
- Schema validation — every case follows the same structure (scenario, context, expected, grader) so the set can be run programmatically.
- Deduplication — near-identical cases inflate the set without adding coverage.
- Versioning — version the dataset like code. Pin a dataset version, run one experiment per application version against it, and read the score deltas against a designated baseline.[1] This is what lets you say "faithfulness went up 2 points since last week" with confidence.
Step 8 — Keep It Alive
An eval set is never finished — a good one mirrors production reality and evolves with the product.[8] Build the feedback loop in:
- Every production failure becomes a new case — so it can't silently recur. This is the self-service flywheel applied to the eval set.
- Every new intent or feature adds scenarios before it ships.
- Policy changes update expected outcomes — when the rules change, the gold standard changes with them.
- Sample live calls into the set — periodically pull real recent conversations to keep the distribution current as caller behaviour drifts.
Where to Start
- List what "good" means for your IVR (Step 1) — intent, task completion, routing, faithfulness, policy, escalation.
- Pull ~20 real scenarios from your top call drivers and worst known failures, and write the ideal outcome for each.
- Cover the mix — mostly happy path, a solid slice of edge cases, a few adversarial.
- Add the voice-specific cases — no-input/no-match, barge-in, disfluent speech, spoken data capture.
- Assign graders — rules for the clear-cut checks, an LLM judge for the nuanced ones, and calibrate the judge against human labels.
- Validate, version, and run it on every change — then grow it toward 100+ and let production failures keep feeding it.
Do that and you move from shipping IVR changes on hope to shipping them against a measured, defensible standard — one that gets stronger every time something goes wrong.
Designing the flows and test scenarios your eval set will be built from? The free IVR Design Tool generates a route-by-route test plan and no-input/no-match cases that make a natural seed for your first eval set.
Sources
This article draws on published guidance on building evaluation datasets. Figures (sizes, mixes) are rules of thumb reported by each source and vary by use case; treat them as directional. Content was summarised and rephrased for compliance. This is independent guidance, not affiliated with the organisations cited.
- Langfuse — "Golden dataset evaluation: build and maintain LLM test sets"
- Lyzr — "Golden Datasets for AI Agent Evaluation"
- Digital Applied — "Building an AI Agent Evaluation Pipeline"
- Beam.ai — "How to Actually Run Evals on Your AI Agents" (60/30/10 mix; write the ideal output)
- OpenRouter — "Building a Golden Eval Dataset from Production Traffic" (sizing guidance)
- Hamel Husain — "How many examples do I need for an eval?" (100–200 per failure mode)
- learnwithparam — "LLM-as-a-judge production evaluation framework"
- Maxim — "Building a Golden Dataset for AI Evaluation"