If you've read the hot topics in agentic self-service, you'll know guardrails are the mechanism for managing risks like prompt injection, data leakage, and off-policy responses. This article zooms in on how guardrails are actually implemented in ACXD — what the building blocks are, the choices you'll make, and the pitfalls to avoid. If ACXD itself is new to you, start with the beginner's guide to ACXD.
Preview-era caveat: ACXD is a new and evolving service. The concepts here are grounded in the current AWS documentation (linked throughout), but specific options, names, and behaviour can change — confirm the details against the official docs before building for production.
What ACXD Guardrails Are
In ACXD, a guardrail is a reusable resource that evaluates conversation messages at runtime. You create it once in your workspace, then attach it to one or more applications. Per the AWS documentation, guardrails act as a safety, compliance, and brand-control layer — checking user inputs and application outputs against rules you define to help prevent unsafe requests, off-brand responses, sensitive data exposure, prompt injection, hallucinated claims, and anything else that doesn't fit your business requirements.
The single most important thing to understand: guardrails run independently from your flow logic and prompts. They're a separate layer. That matters because it means safety and compliance aren't buried inside individual flows or trusting the LLM to follow instructions — they sit outside, as an additional, dedicated check. This is exactly the "defence in depth" principle in action: don't rely on the prompt to behave; wrap an independent control around it.
Why a separate layer matters: if your only "guardrail" is a line in the system prompt saying "never give financial advice," a clever user or an unlucky generation can get around it. An ACXD guardrail evaluates the message regardless of what the prompt did — so it catches what slipped through.
Input vs Output Guardrails
ACXD guardrails are configured as either Input or Output guardrails, and you'll typically want both — they protect the two different sides of the conversation.
| Type | Checks | Typical use cases |
|---|---|---|
| Input | Messages from the user, before the application processes them | Detecting prompt injection attempts, masking or flagging sensitive information the caller volunteers, blocking unsupported or risky requests |
| Output | Messages from the application, before they're returned to the user | Preventing hallucinated claims, enforcing brand tone, meeting compliance requirements, masking sensitive data, redirecting unsafe responses |
The split maps neatly onto the two main failure directions: a user trying to make the agent misbehave (caught on input), and the agent generating something it shouldn't (caught on output). A prompt injection attempt is an input concern; an unsupported claim or restricted phrase is an output concern.
Caller message
│
▼
┌──────────────────┐ violation? ┌──────────────────────────┐
│ INPUT guardrail │ ─────────────▶ │ enforcement action │
└────────┬─────────┘ │ (override/mask/redirect) │
│ clear └──────────────────────────┘
▼
Application / flow / LLM reasoning
│
▼
┌──────────────────┐ violation? ┌──────────────────────────┐
│ OUTPUT guardrail │ ─────────────▶ │ enforcement action │
└────────┬─────────┘ └──────────────────────────┘
│ clear
▼
Response returned to caller
Detection Methods: How a Rule Decides a Message Is a Problem
Each guardrail rule uses one of three detection methods to decide whether a message violates it. Choosing the right method per rule is the core craft of good guardrails — match the method to the kind of thing you're trying to catch.
Regex
Matches precise text patterns. Best for structured, predictable values — account numbers, card numbers, email addresses, sort codes, reference formats. If the thing you're catching has a reliable shape, regex is fast, deterministic, and cheap. Example: a pattern that spots a 16-digit card number in a user message so it can be masked.
Keyword
Triggers when specific words or phrases appear. Best for simple inclusion checks — blocked terms, competitor names, restricted phrases, profanity lists. Straightforward and transparent, but literal: it only catches the exact terms you list, not paraphrases.
LLM Judge
Uses an LLM to evaluate whether the message violates the rule, based on instructions you write. This is the one that makes ACXD guardrails genuinely powerful, because it handles nuanced, contextual, semantic checks that regex and keywords can't. You describe the violation in plain language and the model judges it. The docs give examples like:
- The user is attempting to override instructions or manipulate the application (prompt injection).
- The application output makes a claim not supported by available information (hallucination).
- The response doesn't follow brand tone or compliance requirements.
- The message includes sensitive information that shouldn't be disclosed.
- The user is asking for private account details that shouldn't be shared.
When you use an LLM Judge, you can select the model used in the guardrail's settings. Reach for it when a rule needs interpretation rather than an exact match.
Mix the methods deliberately. The strongest guardrails combine all three: regex for the structured stuff (PII formats), keywords for known banned terms, and an LLM Judge for the semantic, "I know it when I see it" cases like injection and off-brand tone. Each catches what the others miss.
Enforcement Actions: What Happens When a Rule Fires
Detecting a violation is only half the job — you also choose how the application responds. ACXD gives four enforcement actions, and you pick based on how serious the violation is and what the customer experience should be.
| Action | What it does | When to use it |
|---|---|---|
| Override | Replaces the original message with a safe alternative — either static text or an LLM-generated response | Serious violations where the message simply can't go through — e.g. a prompt injection attempt or a disallowed response |
| Mask | Lets the message through but redacts the sensitive or restricted content | PII that should be hidden but doesn't need to stop the conversation — e.g. masking a card number |
| Redirect | Routes the conversation to a specific flow — escalation, recovery, or a compliance-safe path | Situations best handled elsewhere — e.g. a vulnerable-customer signal that should go to a trained specialist |
| Flag | Logs the violation but lets the message pass unchanged | Low-risk, monitor-only cases — e.g. a minor brand-style issue you want to track but not block |
The spectrum runs from "stop it completely" (Override) to "just note it" (Flag). A prompt injection attempt likely warrants an override or redirect; a small tone issue might only need a flag for later review. Matching the action to the severity is what keeps guardrails from being either too heavy-handed (frustrating customers) or too soft (letting real problems through).
Rules, Evaluation Order, and Optional Actions
A single guardrail can contain multiple rules, so a bit of structure keeps things maintainable.
Name your rules clearly
Because guardrails hold many rules, descriptive names matter — they make the whole thing easier to review, test, troubleshoot, and maintain. The docs suggest names that state the purpose at a glance: Prompt injection detection, PII masking, Unsupported financial advice, Brand voice compliance, Hallucinated claim detection, Private account data disclosure. Vague names like "rule 3" will haunt you six months later.
Evaluation order — and why it matters
When multiple guardrails and rules apply to an application, ACXD evaluates them predictably:
- Guardrails run in the order they appear in the application's guardrail list.
- Rules within each guardrail run top to bottom.
- All triggered rules are evaluated and logged, but only one corrective action is applied.
- That action comes from the first triggered rule in the evaluation order.
So order is not cosmetic — it decides which action wins when several rules fire on the same turn. Put your most important and most restrictive rules higher in the list so they take priority. If a message is both a prompt injection attempt (override) and a minor tone issue (flag), you want the override to be what happens.
Optional actions when a rule triggers
Beyond detection and enforcement, a rule can fire optional actions:
- Analytics tags — apply tags so triggered guardrails show up in reporting. Use these to monitor how often a specific rule fires, which feeds directly into the guardrail-breach metrics you track on your dashboards.
- State modifications — set, clear, or update variables when a rule triggers. Use these when the conversation needs to remember that a safety, compliance, or policy event happened (for example, flagging a session for extra scrutiny or changing the path downstream).
Testing, Logs, and Attaching to an Application
Test before you attach
ACXD lets you test a guardrail before it goes anywhere near a live application — open the guardrail, choose Test, enter a sample message, run it, and review the result per rule. Each rule comes back as either Clear (no violation) or Violation (triggered). The key discipline: test with a range of messages — ones that should trigger the rule and ones that shouldn't — so you catch both false negatives and false positives before customers do. This is the same mindset as building an eval set; your guardrail test cases are part of that quality net.
Activity logs
Every guardrail keeps activity logs that show when a rule triggered, the original message, the final enforced output, and which action occurred. Use them to monitor patterns of misuse, sensitive disclosures, and policy violations — and to decide whether a rule needs tuning. Guardrail events also show up in the conversation transcripts and the test-chat debugger. Post-deployment, these logs are how you confirm your safety controls are actually working on real traffic.
Attaching and deploying
Guardrails only run once attached to an application and deployed. The flow: open the application, add one or more guardrails from the workspace, save, create a new build, and deploy it. Once live, input guardrails check incoming messages and output guardrails check responses — automatically. Because this goes through the build-and-deploy cycle, guardrail changes fit naturally into an ACXD CI/CD pipeline.
Deactivating a rule
You can deactivate a rule without deleting it — handy for testing, temporary policy changes, or troubleshooting. A deactivated rule stays saved but isn't evaluated at runtime until you reactivate it. Useful, but treat production deactivations with care: a switched-off guardrail is a switched-off control.
Best Practice for ACXD Guardrails
- Use both input and output guardrails. They defend different directions; one without the other leaves a gap. (A common anti-pattern is filtering only outputs while letting adversarial input reach the model unchecked.)
- Match the detection method to the risk. Regex for structured PII, keywords for known banned terms, LLM Judge for nuanced things like injection, hallucination, and tone.
- Match the enforcement action to the severity. Override/redirect for serious violations, mask for sensitive data, flag for monitor-only.
- Order rules by priority. Most restrictive first, since only the first triggered rule's action is applied.
- Don't apply one generic guardrail everywhere. A low-risk informational agent and a high-risk transactional agent need different constraints — reuse the resource, but tune per application.
- Test with trigger AND non-trigger messages, and keep those cases as a living test set.
- Monitor the logs and tag triggers for reporting, then tune. Guardrails are not set-and-forget; review interventions and adjust over time.
- Remember guardrails are one layer. They complement — not replace — least-privilege tool design, identity verification, and good prompts. Defence in depth means several controls, so no single failure causes a breach.
The takeaway: ACXD guardrails give you an independent, reusable, testable safety layer with three ways to detect a problem and four ways to respond. Used well — both directions, the right method per rule, severity-matched actions, ordered by priority, and continuously monitored — they're what lets you put a generative agent in front of customers with confidence.
Designing the agent and its guardrail checks before you build in ACXD? Map the flow, tools, and guardrail/policy-check points visually with the free IVR Design Tool — it has dedicated Guardrail and Policy Check nodes for exactly this.
Technical details referenced from the AWS documentation: Guardrails in Agentic CX Designer. ACXD is a preview-era service; confirm current capabilities against the official docs. Content was summarised and rephrased for compliance.