Thinking / AI agents

How to build guardrails into an AI agent

A business-led framework for objectives, permissions, human escalation, monitoring and recovery in agentic workflows.

The answer in brief

What it is
A business-led framework for objectives, permissions, human escalation, monitoring and recovery in agentic workflows.
Best suited to
Leaders who need to translate this issue into an investment, workflow, governance or capability decision.
What useful progress looks like
An organisation may have responsible AI principles while an agent still holds excessive access, ambiguous objectives or no practical escalation route. Agentic systems can retrieve, decide and act at speed. A small error can therefore travel through customer, financial or operational systems before periodic review detects it. Guardrails must exist inside the product, workflow and operating model. They should constrain behaviour, create evidence and give accountable people the ability to intervene.

The question leaders ask

What guardrails does an AI agent need?

Every production agent needs a clear objective, bounded information and tool access, prohibited actions, decision thresholds, human escalation, continuous monitoring and a tested recovery route.

Guardrail layers7
Default principleLeast authority
RecoveryVisible and reversible
01

Policy language does not control an action

An organisation may have responsible AI principles while an agent still holds excessive access, ambiguous objectives or no practical escalation route. Agentic systems can retrieve, decide and act at speed. A small error can therefore travel through customer, financial or operational systems before periodic review detects it. Guardrails must exist inside the product, workflow and operating model. They should constrain behaviour, create evidence and give accountable people the ability to intervene.

02

Design seven layers together

No single prompt or filter creates dependable control. The layers should reinforce each other and reflect the consequence of the workflow.

01

Objective

Define the outcome, priority and trade-offs the agent is permitted to optimise.

02

Identity and access

Give the agent its own identity and the minimum data, systems and actions required.

03

Action boundaries

List permitted, prohibited and approval-dependent actions in operational language.

04

Decision thresholds

Set confidence, value, sensitivity and exception conditions that trigger human review.

05

Input and output controls

Validate data, detect sensitive content and constrain unsafe or unsupported responses.

06

Observability

Log sources, decisions, tool calls, outcomes, overrides, cost and changing performance.

07

Recovery

Provide pause, rollback, correction, incident and customer-remedy routes.

03

Put humans where authority matters

Human review has value only when the person receives the right context and can alter the decision. Requiring approval for every low-risk action can create delay and encourage workarounds. Removing people from consequential decisions can create unacceptable exposure. Assign authority according to impact, uncertainty and reversibility. Monitor override patterns, since frequent intervention often signals a weak boundary, poor information or a change in operating conditions.

04

Treat guardrails as a living product

Models, tools, sources, user behaviour and external requirements change. Test the agent before launch using normal cases, edge cases and adversarial scenarios. Continue evaluation in production. Review incidents, near misses, overrides and business outcomes through a named cadence. Expand permissions only when evidence supports greater autonomy. A well-governed agent can gain authority gradually while preserving a clear record of why the decision was made.

FAQ

Questions leaders ask.

Direct answers to the questions that commonly shape an initial conversation.

01Are prompt instructions sufficient guardrails?+

No. Prompts are one layer. Access controls, workflow permissions, validation, monitoring, escalation and recovery provide stronger operational assurance.

02Should every agent action require approval?+

No. Approval should reflect consequence and uncertainty. Low-risk reversible actions can be delegated within clear limits.

03Who is accountable for an agent?+

A named business owner should own the outcome, supported by product, technology, data and risk specialists.

04How often should guardrails be reviewed?+

Monitor continuously and review formally after incidents, material model or workflow changes, and through a regular product-governance cadence.

Latest from the blog

Useful ideas for the decision in front of you.

View the blog

Apply the thinking

Put proportionate control around a production agent

Bring the live decision, workflow or commercial pressure. We will help translate the idea into a focused next step.

Discuss the implication