AI & Quality • October 2026

How to Build High-Performing AI Agents: Using Evals to Deliver Excellence in Self-Service

Anyone can stand up an AI self-service bot in an afternoon now. Making one that's genuinely excellent — accurate, safe, on-brand, and reliably so — is a different discipline entirely. The teams who get there all do the same thing: they treat evaluations ("evals") as the mechanism that holds the quality bar. Here's how to apply that to Amazon Connect ACXD, AI Agents, or any conversational AI project.

There's a pattern to AI self-service projects. The demo is dazzling. The pilot is promising. Then it hits real traffic and the cracks appear — a confidently wrong answer here, an off-policy response there, a regression nobody noticed after a prompt tweak. The bot that wowed everyone in the boardroom quietly erodes trust in production.

The difference between the bots that stay excellent and the ones that drift isn't a better model or a cleverer prompt. It's that the high performers have a systematic way of measuring quality — and they use it as a gate, not an afterthought. That mechanism is evals.

In this guide
  1. Why "vibes" don't scale to excellence
  2. What an eval actually is
  3. The eval types that matter for self-service
  4. How to build your eval set
  5. Where evals run: dev, CI, and production
  6. Applying this to ACXD and AI Agents
  7. Where to start

Why "Vibes" Don't Scale to Excellence

Traditional software is deterministic — the same input gives the same output, so a fixed test suite catches regressions. Generative AI isn't like that. The same question can produce six different, equally valid phrasings; temperature adds randomness; a model or prompt update can shift behaviour in ways no one intended. Checking output against a fixed expected string fails constantly, even when the agent is doing the right thing.

So most teams fall back on "vibes" — someone reads a handful of transcripts, decides it looks good, and ships. That works for a demo. It does not scale to the thousands of real, messy conversations production throws at you. A bot can be factually correct but wrong in tone, safe but useless, concise but incomplete, or helpful in a way that quietly invents unsupported details. Spot-checking won't catch that reliably, and it certainly won't catch a slow regression over time.

Evals replace opinion with measurement. They turn "I think it's good" into "it scores 94% on faithfulness across 300 real scenarios, up 2 points since last week."

What an Eval Actually Is

An eval is simply a repeatable test for an AI system: give it an input, then apply grading logic to its output to measure success. Where a unit test checks a code path, an eval checks judgement — did the agent answer correctly, ground its claims in approved content, complete the task, and stay in policy?

Every eval has three parts:

Run a set of these consistently and you have the AI equivalent of a test suite — one that makes the agent's variability visible and manageable instead of a mystery.

The mindset shift: stop asking "does this response look good?" and start asking "what score does this version get against our eval set, and did it go up or down?" Excellence is a number you can defend, track, and improve — not a feeling.

The Eval Types That Matter for Self-Service

No single eval catches everything — each type finds a different class of bug, and you layer them. For contact centre self-service, these are the ones that earn their place:

Task completion / goal success

Did the agent actually resolve what the customer came for — reset the password, log the dispute, make the payment — not just sound helpful? This is the headline quality signal. For multi-turn conversations, it measures whether the whole dialogue drove to the goal, turn after turn.

Faithfulness / groundedness

Is every claim traceable to approved content (a knowledge base, retrieved document, or system data) rather than the model's own memory? This is your single strongest defence against hallucination — the number-one risk of a generative agent. A contained interaction that gave wrong information is worse than no automation at all.

Policy adherence / off-policy detection

Did the agent stay inside the rules — never giving advice it shouldn't, never revealing data it shouldn't, always verifying identity before disclosure? In regulated sectors a single off-policy response is a compliance incident, so this eval is non-negotiable.

Tool-use correctness

When the agent chose to act, did it pick the right tool, with valid parameters, and did the call succeed? Agentic resolution depends on tools actually working — this catches both reasoning errors (wrong tool) and integration failures (bad call).

Safety & guardrail evals

Adversarial scenarios: jailbreak attempts, prompt injection, social engineering, requests for restricted actions. These probe the boundaries rather than the happy path — closer to security testing than QA. (See the agentic security hot topics.)

Tone & brand

Is the agent empathetic, clear, and on-brand — and does it stay that way under pressure? An agent can complete the task and still leave the customer feeling worse. Tone is measurable with an LLM judge against your brand guidelines.

The three ways to grade

Across all those types, graders fall into three buckets, and good eval suites use all three:

Don't fully trust the judge. LLM-as-a-judge is powerful but imperfect. Calibrate it against human-labelled examples, keep humans in the loop for high-stakes categories, and periodically re-check that the judge still agrees with your experts. An un-calibrated judge gives you confident, precise, wrong numbers.

How to Build Your Eval Set

The eval set — your collection of scenarios and expected outcomes — is the real asset. The model will change; this is what protects your quality through every change. Build it like this:

1. Start from real intents and real failures

Don't invent scenarios in a vacuum. Pull them from your actual call drivers and transcripts: the top intents, the awkward edge cases, and especially the conversations where the bot (or a human) got it wrong. Every production failure should become a permanent eval so it can never silently regress again.

2. Build a "golden set" with known-good outcomes

For a core set of scenarios, agree the ideal outcome with your business and compliance stakeholders — the gold-standard answer, the correct action, the right escalation. This is what you grade against and the thing you defend "excellence" with.

3. Cover the whole spread, not just the happy path

A strong eval set deliberately includes: common happy-path intents, messy/ambiguous phrasing, multi-intent requests, out-of-scope asks, adversarial inputs, and vulnerable-customer situations. If it's a risk in production, it belongs in the eval set.

4. Example: an eval entry

# One eval scenario for a banking disputes agent
scenario_id: dispute-unrecognised-charge-001
input: "There's a £47 charge from TXN-DIGITAL I don't recognise."
context: { customer_verified: true, txn_exists: true, amount: 47.00 }
expected:
  task_completed: "dispute_logged"
  tools_called: [get_recent_transactions, flag_dispute, provisional_refund]
  must_include: ["case reference", "provisional refund"]
  must_not: ["full card number", "guaranteed refund"]
graders:
  - type: rule        # tool sequence + forbidden strings
  - type: llm_judge   # faithfulness, tone, policy adherence

Notice it checks behaviour and boundaries, not an exact wording — because there's no single correct phrasing, only a correct outcome.

5. Keep it alive

An eval set is never finished. Every new feature adds scenarios; every production incident adds a regression test; every policy change updates the expected outcomes. Version it alongside your prompts and flows so you can see exactly what "good" meant at any point in time.

Where Evals Run: Dev, CI, and Production

Evals aren't a one-time pre-launch check. They deliver excellence precisely because they run continuously, at three points in the lifecycle:

  ┌───────────────────────────────────────────────────────────────┐
  │  1. DEV LOOP   Change a prompt / flow / tool → run evals locally │
  │     → see the score move before you commit                      │
  └───────────────────────────────┬───────────────────────────────┘
                                  ▼
  ┌───────────────────────────────────────────────────────────────┐
  │  2. CI GATE   On every PR → run the full eval set               │
  │     → block the merge if scores regress below threshold         │
  └───────────────────────────────┬───────────────────────────────┘
                                  ▼
  ┌───────────────────────────────────────────────────────────────┐
  │  3. PRODUCTION   Sample live conversations → grade continuously │
  │     → catch drift, feed new failures back into the eval set     │
  └───────────────────────────────────────────────────────────────┘
                                  │
                                  └──────────▶ back to step 1 (improve)

That closed loop — change, gate, monitor, improve — is exactly the self-service flywheel, with evals as the measurement engine that powers it. It also pairs directly with the AI agent metrics you track on your observability dashboards: the metrics tell you what moved; the evals tell you whether a change made it better or worse.

Applying This to ACXD and AI Agents

None of this is theoretical for Amazon Connect. The platform is building the hooks you need:

The through-line is that evals give you a defensible definition of "good" that survives model swaps, prompt changes, and vendor updates. When your foundation model is updated under you — which will happen — your eval set is what tells you in minutes whether quality held or slipped, rather than finding out from an angry customer.

Evals are how "deliver excellence" stops being a slogan. They convert a vague aspiration into a measured, gated, continuously-improving standard. A bot without evals is being managed on hope. A bot with them is being engineered.

Where to Start

You don't need a platform or a big budget to begin — you need the discipline. Start small and compound:

  1. Write 20 scenarios from your top intents and your worst known failures, with the ideal outcome for each. That's your first golden set.
  2. Grade with a mix — rules for the clear-cut checks (forbidden strings, tool sequence), an LLM judge for faithfulness, tone, and policy.
  3. Calibrate the judge against a handful of human-labelled examples so you trust its scores.
  4. Run them on every change — even manually at first — and record the score. Make "did the eval score go up?" the question before anything ships.
  5. Automate the gate once it's proving its worth, and add production sampling to close the loop.

Do that, and you move from launching a bot and hoping, to running a bot you can prove is excellent — and keep excellent.

Designing the self-service journey these evals will test? Map the intents, tools, and guardrail checks first with the free IVR Design Tool — its generated test plan is a natural starting point for your eval scenarios. And if you're defining what "excellent" means for your project, pair this with defining success criteria by role.

Note: the LLM-as-a-judge accuracy figure (~85% agreement with humans) reflects commonly reported industry findings and varies by task, model, and rubric — treat it as directional, and calibrate against your own human-labelled data. Content on industry practice was summarised and rephrased from publicly available sources.