Guardrails for agents that act
Sort actions by reversibility and set approval, limits, and logging.
An agent differs from an assistant in one consequential way: it takes actions rather than only producing text. Sending an email, booking a slot, updating a record, issuing a refund. That shift changes the risk profile entirely, because a bad paragraph is embarrassing and a bad action is a real event in the world that someone has to undo. Guardrails are therefore designed around reversibility. Sort every candidate action into three tiers. Freely reversible — drafting, tagging, internal notes — can run without approval. Recoverable with effort — sending internal messages, scheduling that can be moved — can run with notification and an easy undo. Irreversible or externally visible — customer emails, payments, refunds, cancellations, public posts, anything touching a third party — requires explicit human approval every time, with no exceptions for convenience. Add hard limits alongside the tiers: maximum actions per run, maximum amounts, allowed recipients, and a kill switch anyone on the team can use. Then keep it observable. Every action logged with what, when, why, and on whose authority. A daily digest someone actually reads. An alert when the agent hits a limit or refuses a task, because refusals are information about where the design is wrong. Roll out in stages — shadow mode where it proposes and you act, then approval mode, then limited autonomy for the lowest tier only — and expand only on evidence. The failure that hurts is never the dramatic one; it is a small wrong action repeated four hundred times before anyone notices. The consequential failure mode with agents is repetition, not severity. One wrong email is recoverable; the same wrong email sent four hundred times before anyone reads the log is a different kind of problem. Refusals are design feedback. Every time an agent hits a limit or declines a task, it is telling you where your specification and reality disagree, which is more useful than a clean log. A rental company ran a scheduling agent in shadow mode for two weeks and reviewed what it would have done. The log revealed a systematic timezone error that would have moved dozens of real appointments incorrectly.
Key takeaways
- Guardrails are therefore designed around reversibility.
- Then keep it observable.
- Refunds are irreversible, external, and financial.
Sign in to complete activities, save progress, and use the AI Coach.
Evergreen instructional material; no time-sensitive claims to source.
Evergreen curriculum written from KleinHub teaching material — no time-sensitive claims.
AI-generated lesson, reviewed by Education Review and approved through the KleinHub approval queue before publication.
Sign in to complete this lesson, track progress, and use the AI Coach.
Sign in to start lessonCreate account