Design & Strategy • October 2026

How to Document an Agentic Self-Service Experience

For decades, "the documentation" for a self-service experience was the flow diagram — every box, every branch, every prompt. Agentic AI breaks that model. When the system reasons rather than follows a script, a flowchart no longer describes what it does. So what replaces it? This is a research-backed look at how to document an agentic self-service experience, drawing on how Anthropic, OpenAI, and Google document their own agents.

If you've moved from deterministic to agentic self-service, you've hit this problem: your old design artifacts don't fit. A flow diagram captures a scripted IVR perfectly, because the IVR only does what the diagram says. An agentic experience is different — it interprets, decides, and acts within boundaries you set. You're no longer documenting every path; you're documenting intent, boundaries, and behaviour, and trusting the agent to navigate between them.

Rather than invent a format, it's worth looking at how the organisations building the underlying models document agents — because they've converged on a strikingly consistent set of artifacts. This article synthesises that into a practical documentation framework for a contact-centre self-service agent.

In this guide
  1. Why documentation changes with agentic
  2. What Anthropic, OpenAI, and Google actually document
  3. The 7 artifacts that document an agentic experience
  4. Keep it living, versioned, and testable
  5. Where to start

Why Documentation Changes with Agentic

The core shift, well put by Anthropic, is that tools (and by extension agents) are "a new kind of software which reflects a contract between deterministic systems and non-deterministic agents." A traditional function always behaves the same way; an agent, given the same input, might take different valid paths. You cannot enumerate every path, so documenting one is futile.

What you document instead is the contract: what the agent is trying to achieve, what it may and may not do, what it can call, what it knows, and how you'll know it's behaving. The artifacts below are that contract. They're also what makes the experience reviewable by compliance, repeatable across environments, and improvable over time — none of which a prose "it uses AI to help customers" description delivers.

What Anthropic, OpenAI, and Google Actually Document

Three different organisations, three stacks — but a remarkably consistent view of what defines (and therefore documents) an agent.

OpenAI: an agent is a bundle of named parts

OpenAI's own definition is useful because it's a checklist in disguise. In their framing, an agent "packages a model, instructions, and optional runtime behavior such as tools, guardrails, MCP servers, handoffs, and structured outputs." That single sentence names most of what you need to document: the model, the instructions, the tools, the guardrails, how it hands off, and the shape of its outputs. Their governance guidance adds a telling principle — that policies should travel with the code, so review becomes approval rather than interrogation.

Anthropic: document tools like you're onboarding a new hire

Anthropic's engineering writing on tools for agents is the richest source on documenting the tool layer. Their standout advice: when you write a tool description, "think of how you would describe your tool to a new hire" — make the implicit context explicit, name parameters unambiguously (user_id, not user), and define niche terms and the relationships between resources. They also stress designing tools around workflows rather than wrapping raw API endpoints, and — crucially — pairing every tool with evaluations grounded in real tasks.

Google: instructions must be specific, unambiguous, and organised

Google's guidance for conversational/CX agents zeroes in on the instructions themselves: agent instructions should be "specific and unambiguous," "well organized and grouped by topics" rather than scattered, and "easy for a human to follow as well." That last point matters for documentation — the instructions are read by both the model and your reviewers, so clarity serves two audiences at once.

The consensus: an agent is documented not as a flowchart but as a set of parts — instructions, tools, guardrails, knowledge, context, outputs/handoffs, and evaluations. Written clearly enough that a human reviewer and the model both understand them. That's the framework the rest of this article builds on.

The 7 Artifacts That Document an Agentic Self-Service Experience

Here's the practical set, adapted for a contact-centre self-service agent. Together these replace the flow diagram. Throughout, I'll use a running example: a bank's card-servicing self-service agent.

Artifact 1

Purpose & scope statement

A short, plain statement of what the agent is for, who it serves, which channels it runs on, and — just as important — what is explicitly out of scope. This frames everything else and stops scope creep. Agentic systems are easy to over-ask; naming the boundary up front is a documentation discipline Anthropic echoes ("choose the right tools... and not to").

Example: "Helps verified retail-banking customers with card servicing — balance, recent transactions, disputes, freeze/unfreeze, and replacement — over voice and chat. Out of scope: lending decisions, financial advice, business accounts."

Artifact 2

Instructions / system prompt

The agent's brief — its role, personality, tone, and decision-making guidance. Following Google's rule, keep it specific, unambiguous, and grouped by topic rather than a wall of scattered rules, and readable by a human reviewer. This is the nearest equivalent to the old prompt script, but it defines how to decide rather than what to say at step 3. Treat it as a versioned document, not a field buried in a console.

Document: role & persona, tone, the goals it pursues, how it should handle ambiguity, when to ask vs. act, and when to escalate.

Artifact 3

Tool catalogue (with contracts)

The single most under-documented layer, and the one Anthropic has the most to say about. Each action the agent can take is a tool, and each needs a clear contract: a precise name, a description written "like for a new hire," unambiguously named inputs with types and limits, outputs, and failure behaviour. Design tools around workflows, not raw API endpoints — a dispute_transaction tool that does the whole job beats three thin wrappers the agent has to chain.

# Tool spec — document each tool like this
name: dispute_transaction
description: "Open a dispute for a transaction the customer does not"
             "recognise. Use only after identity is verified and the"
             "customer confirms they did not make the transaction."
inputs:
  transaction_id:  { type: string,  required: true }
  reason_code:     { type: enum,    values: [not_recognised, duplicate, wrong_amount] }
constraints:
  - customer must be fully verified (high trust)
  - provisional refund capped at the documented limit
  - amounts above the cap must escalate to a human
returns:        { case_reference: string, provisional_refund: boolean }
on_error:
  insufficient_permissions: "escalate to human agent"
  transaction_not_found:    "ask the customer to re-confirm the amount/date"

Note the deliberate choices Anthropic recommends: a semantic name (transaction_id, not an opaque field), an enum that constrains inputs, documented limits, and helpful error behaviour that tells the agent what to do next rather than returning an opaque code.

Artifact 4

Guardrail & policy specification

What the agent must never do, and what requires confirmation or a human. OpenAI frames guardrails and human review together as the controls that decide "when a run should continue, pause, or stop" — and their governance guidance is that these policies should travel with the code, not live in a separate wiki nobody reads. Document the hard limits, the actions needing confirmation, and the escalation triggers, in a form your compliance team can sign off.

Example rules: never reveal full card numbers; never give financial advice; always verify identity before disclosure; any refund above the cap escalates; a vulnerability signal routes to a specialist. (See the dedicated piece on guardrails and the security hot topics.)

Artifact 5

Knowledge sources

Where the agent gets grounded answers from — the knowledge bases, documents, and systems it draws on, and which questions each covers. Document the source, its owner, how it's kept current, and what's deliberately not in scope for it. Grounding answers in named, approved sources is the strongest defence against hallucination, so documenting those sources is a quality control, not an afterthought.

Document: each source, owner, refresh cadence, and the topics it's authoritative for.

Artifact 6

Context & handoff design

What the agent carries in its working memory (verified identity, intent, values gathered), what it does not hold, and how it hands off — to another flow, another agent, or a human. Anthropic treats this as "context engineering": be deliberate about what's in context, because it affects quality, cost, and safety. For self-service specifically, document the handoff cleanly: when it triggers, what context passes with it, and what the human receives. A clean handoff is good design, not a failure.

Document: the state the agent tracks, context limits, handoff triggers, and the summary passed on escalation.

Artifact 7

Evaluation set

The one every lab insists on. Anthropic pairs every tool with evaluations "grounded in real-world uses"; a documented agent isn't complete without a set of scenarios and expected outcomes that prove it behaves. This is both documentation and the test suite — it records what "good" means. (It deserves its own treatment: see using evals to build high-performing agents.)

Document: representative scenarios (happy path, messy, adversarial, out-of-scope), the expected outcome and tools for each, and how it's graded.

The flow diagram isn't entirely dead. For the deterministic parts that remain — identity verification, regulated scripts, fixed sequences — a diagram is still the right tool. Most real self-service is a hybrid: agentic reasoning over a spine of deterministic steps. Document the deterministic bits as flows, and the agentic bits as the seven artifacts above.

Keep It Living, Versioned, and Testable

Documenting an agentic experience is not a one-time deliverable you file away — all three sources are emphatic that this is iterative. A few principles pulled from how the labs work:

The mindset: treat these artifacts as the living source of truth for the experience, not paperwork produced after the build. If the documentation and the deployed agent disagree, that's a defect — the same way a failing test is.

Where to Start

You don't need all seven artifacts polished on day one. A pragmatic order:

  1. Write the purpose & scope statement — one paragraph, including what's out of scope. It frames everything.
  2. Draft the instructions — specific, grouped by topic, human-readable.
  3. Catalogue the tools — one clear contract per action, named and described like you're onboarding a new hire, designed around workflows.
  4. Specify the guardrails — the never-dos, the confirmations, the escalation triggers, in a form compliance can sign off.
  5. List the knowledge sources and the context/handoff design.
  6. Build a starter eval set — 15–20 scenarios with expected outcomes — and let it drive revisions to everything above.

A quick practical note on tooling: a visual canvas can help you capture and share several of these artifacts — the deterministic flow spine, the tool/guardrail touchpoints, and a first cut of test scenarios — in one place a team can review. My free IVR Design Tool is built for exactly that kind of IVR-and-agentic design, and its generated test plan is a reasonable seed for your eval set. But the tool is just a convenience; the best practice here comes from the artifacts and discipline above, not from any particular canvas. Document the contract well in whatever medium suits your team, and the agent — and everyone reviewing it — will be better for it.

Sources and further reading: Anthropic — Writing effective tools for agents and Building effective agents; OpenAI — Agent definitions, Guardrails and human review, and A practical guide to building agents; Google — CX agent best practices. Content was summarised and rephrased from these sources for compliance; consult the originals for full detail.