An equipment case waits beside a closed workshop gate, with a transaction ledger, receipt and key at the checkpoint.
✦

An AI pilot can produce a useful answer and still leave the organisation unable to explain what it has authorised. The important boundary appears when a suggestion changes a record, reserves a resource or creates a promise someone else must honour. I would begin an architecture review at that boundary and work backwards.

This is a proposed review method, developed from my Governed Intelligence approach and Enterprise Agentic Mesh essay. It is a way to organise questions and evidence. It has not been validated as a certification scheme, and completing it does not prove that an agent is suitable for production.

Describe one outcome and its consequences

Choose a case the system is supposed to resolve. Write down the useful outcome, the people affected and the commitments the system might make. “Handle the request” hides too much. “Prepare a permitted equipment loan while protecting the next booking” exposes dependencies, owners and conflicts.

For each proposed action, identify the system that actually changes the world. An agent may recommend the loan while a booking service creates the reservation. Those contributions need separate evidence. The review should be able to distinguish a draft plan, an approved plan, an attempted action and a confirmed result.

Ask who can authorise each change

Describe the agent’s delegated authority in terms of actions and conditions. Which records may it read? Which objects may it change? Which commitments require a person? What happens when a request falls outside the delegation?

Then find the enforcement point. If a payment limit is supplied only as prompt text, the reviewer still needs to know whether the payment service checks it. Open Policy Agent’s documentation describes a separation between policy decisions and the applications that enforce them. The review question is whether that separation is implemented correctly for this action, with current identity, policy and case facts.

A human approval also needs a scope. Record what was approved, by whom, under which conditions and against which version of the proposal. Ask whether changing the recipient or amount invalidates the approval. The useful artifact is a decision-rights map connected to actual action interfaces.

Examine what the agent knows and what it remembers

List the sources required for the decision and the records treated as authoritative. A retrieved document can explain policy. It may be too old to establish today’s equipment availability. Before committing, the system needs a way to check facts whose freshness matters.

Memory deserves the same attention. In my graph-walk experiment, retaining a clue could change a later route; a restrictive context budget could remove it. Those are small authored mechanism tests, not evidence of general retrieval accuracy. They suggest a concrete review question: which unresolved facts survive the budget, and can a discarded source be recovered?

Inspect transitions as well as stored fields. A system may hold an approval without consulting it at execution. A field exists operationally only when its value affects the relevant decision or control.

Give stopping and recovery an owner

Specify the budgets that code enforces: elapsed time, tool calls, branching, cost and repeated attempts. Also specify the conditions for handing work to a person. A model’s confidence score can inform a stopping policy, but it does not establish that the task is complete.

Now follow a partial failure. Suppose the reservation succeeds but its reply is lost. Can the caller check the operation’s status without reserving again? AWS’s engineering account of idempotent APIs explains how explicit request identity can make retries safer. That mechanism must exist in the relevant service contract; adding “retry safely” to an agent instruction cannot supply it.

Some commitments cannot be undone. In those cases, recovery may mean reconciliation, compensation or a responsible person deciding what follows. Identify that person and the evidence they will receive. A stop button should halt future activity without pretending completed actions have disappeared.

End with a decision and a testable handoff

Review the architecture against ordinary cases and cases designed to expose its boundaries: conflicting sources, expired approval, a duplicate request, an unavailable approver and an ambiguous tool result. Inspect the final system state and trace, including failures that a fluent answer might conceal.

The result should connect each material gap to an owner, a proposed change and evidence needed to close it. Decide whether the pilot can continue within its present scope, needs a narrower delegation or requires further engineering before responsibilities expand. Acceptance thresholds belong to the business and system context; a universal pass score would obscure that responsibility.

The useful checklist is therefore short: name the commitment, identify authority, verify current evidence, inspect state changes, bound the work, rehearse failure and decide who carries the remaining uncertainty. If those answers remain vague, expanding autonomy will spread the uncertainty across more systems.

For the design choice behind this review, continue with workflow or adaptive agent. For cross-system failure, read authority and recovery. A focused review can be discussed through Advisory, with scope and outputs agreed together.