Thinking / AI agents
How to build guardrails into an AI agent
A business-led framework for objectives, permissions, human escalation, monitoring and recovery in agentic workflows.
The answer in brief
- What it is
- A business-led framework for objectives, permissions, human escalation, monitoring and recovery in agentic workflows.
- Best suited to
- Leaders who need to translate this issue into an investment, workflow, governance or capability decision.
- What useful progress looks like
- An organisation may have responsible AI principles while an agent still holds excessive access, ambiguous objectives or no practical escalation route. Agentic systems can retrieve, decide and act at speed. A small error can therefore travel through customer, financial or operational systems before periodic review detects it. Guardrails must exist inside the product, workflow and operating model. They should constrain behaviour, create evidence and give accountable people the ability to intervene.
The question leaders ask
What guardrails does an AI agent need?
Every production agent needs a clear objective, bounded information and tool access, prohibited actions, decision thresholds, human escalation, continuous monitoring and a tested recovery route.
Policy language does not control an action
An organisation may have responsible AI principles while an agent still holds excessive access, ambiguous objectives or no practical escalation route. Agentic systems can retrieve, decide and act at speed. A small error can therefore travel through customer, financial or operational systems before periodic review detects it. Guardrails must exist inside the product, workflow and operating model. They should constrain behaviour, create evidence and give accountable people the ability to intervene.
Design seven layers together
No single prompt or filter creates dependable control. The layers should reinforce each other and reflect the consequence of the workflow.
Objective
Define the outcome, priority and trade-offs the agent is permitted to optimise.
Identity and access
Give the agent its own identity and the minimum data, systems and actions required.
Action boundaries
List permitted, prohibited and approval-dependent actions in operational language.
Decision thresholds
Set confidence, value, sensitivity and exception conditions that trigger human review.
Input and output controls
Validate data, detect sensitive content and constrain unsafe or unsupported responses.
Observability
Log sources, decisions, tool calls, outcomes, overrides, cost and changing performance.
Recovery
Provide pause, rollback, correction, incident and customer-remedy routes.
Put humans where authority matters
Human review has value only when the person receives the right context and can alter the decision. Requiring approval for every low-risk action can create delay and encourage workarounds. Removing people from consequential decisions can create unacceptable exposure. Assign authority according to impact, uncertainty and reversibility. Monitor override patterns, since frequent intervention often signals a weak boundary, poor information or a change in operating conditions.
Treat guardrails as a living product
Models, tools, sources, user behaviour and external requirements change. Test the agent before launch using normal cases, edge cases and adversarial scenarios. Continue evaluation in production. Review incidents, near misses, overrides and business outcomes through a named cadence. Expand permissions only when evidence supports greater autonomy. A well-governed agent can gain authority gradually while preserving a clear record of why the decision was made.
Questions leaders ask.
Direct answers to the questions that commonly shape an initial conversation.
01Are prompt instructions sufficient guardrails?+
No. Prompts are one layer. Access controls, workflow permissions, validation, monitoring, escalation and recovery provide stronger operational assurance.
02Should every agent action require approval?+
No. Approval should reflect consequence and uncertainty. Low-risk reversible actions can be delegated within clear limits.
03Who is accountable for an agent?+
A named business owner should own the outcome, supported by product, technology, data and risk specialists.
04How often should guardrails be reviewed?+
Monitor continuously and review formally after incidents, material model or workflow changes, and through a regular product-governance cadence.
Latest from the blog
Useful ideas for the decision in front of you.
Learning to disagree with AI
Why a confident, persuasive answer should be the start of leadership judgement, not the end of it.
Read the articleWhat does a Chief of Staff do in an AI business?
The operating role between technical possibility, commercial pressure and executive attention.
Read the articleChief of Staff vs COO: where does the work split?
A practical distinction between enterprise operations and the executive agenda that cuts across them.
Read the articleApply the thinking
Put proportionate control around a production agent
Bring the live decision, workflow or commercial pressure. We will help translate the idea into a focused next step.
Discuss the implication