Every tool on this page is referenced directly in the book — the eight-pattern failure matrix and incident sequence from Chapter 13, the Agent Passport and Authority Ladder from Chapter 6, and the Workforce Compact rollout from Chapters 5 and 7. Chapter and concept are noted throughout so you can trace everything back to the source.
Back to the BookThe ways a fleet quietly fails, and what to watch for before it does. Run one live fire drill per quarter on a high-consequence boat — business owner, security, legal and a frontline user all in the room.
| Pattern | What it looks like | How to catch it |
|---|---|---|
| Metric gaming | The agent optimizes the number it's measured on in ways that diverge from the outcome you actually wanted. | Audit what the agent is optimizing against a business outcome, not just its own dashboard. |
| Permission creep | Access and authority expand quietly every time the agent gets wired into one more system. | Least-privilege review at every integration; the registry entry updates whenever access changes. |
| Poisoned context | Corrupted, stale or manipulated data enters the agent's working context, and it acts on it as if it were true. | Source validation and provenance checks on any input the agent is allowed to trust. |
| Automation complacency | A system that's usually right gets reviewed less and less carefully, until nobody is really checking. | Sampled deep review even on trusted agents. An override rate near zero is a flag, not a compliment. |
| Runaway costs | Usage, API spend or compute quietly compounds while nobody is watching the meter. | A standing cost ceiling per boat, with an alert before the ceiling is breached, not after. |
| Silent drift | The model, the data or the world changes gradually, and performance degrades before anyone notices. | Baseline re-testing on a fixed schedule, not only when something visibly breaks. |
| Orphaned credentials | An agent's owner changes roles or leaves the company, and nobody revokes its access. | A registry review that fires automatically whenever an owner's employment status changes. |
| Cross-agent collision | Two boats, each doing its own job correctly, collide because neither knew the other's plan. | A coordination check before dispatch for any boat whose actions can conflict with another's. |
When a drill runs — or a real incident hits — four numbers tell you whether the guardrails actually hold under pressure. A guardrail has value when your organization can use it under pressure, not just describe it in a policy.
Cadence: one live fire drill per quarter, on a high-consequence boat.
Before any agent touches production, it gets its papers. Ten fields, filed in one searchable registry.
Capability is not authority. An agent may be technically able to work at Level 4 the day you unbox it — it operates at Level 1 until it has earned the rest through logged evidence.
The book teaches the rollout in four phases — Prepare, Pilot, Scale, Sustain. Below is that rhythm broken out by who does what, so a leadership team, HR and a pilot team can each see their part. Pace it to what your organization can actually absorb; the book deliberately doesn't prescribe a universal week count, because the right pace depends on the size of the water you're crossing.
Leadership: say the true thing first — what the water will take, what
you commit to. Sign the Workforce Compact before any software lands.
HR: draft the seven Compact commitments (advance notice, task redesign
before role removal, funded learning time, fair access to new roles, transition support,
shared productivity gains, protection of voice) into policy your legal team can stand
behind.
Managers: identify the one team and workflow with the clearest,
highest-value case — and the readiness to volunteer.
Leadership: open the kickoff by naming the fear before anyone else can —
"I know some of you are worried about what this means for your jobs."
Manager: recruit a respected peer as ambassador, so the agent arrives
introduced by someone trusted, not installed by an edict. Run open forums where any
question is allowed.
Team: use the boat, log what breaks, and know that surfacing a failure
in public — "here's what broke, here's what we learned" — is what makes it safe to say
the boat leaks.
Leadership: tour and listen instead of presenting. Publicize the pilot
team's own stories — one colleague's "I got my time back" outweighs a hundred slides of
"trust us."
Manager: every new department gets its own ambassador, ideally someone
who already crossed with the pilot team.
HR: the Compact travels with the boats into every new department — same
seven commitments, same reporting.
Leadership: report against the Compact quarterly — roles affected, people
trained, internal placements, gains shared — like the operating commitment it is, not a
poster.
Manager: keep the failure notes public and the forums open as the water
keeps rising.
HR: update the Compact itself as new rungs flood — it is a living
document, not a one-time announcement.
These are the book's own frameworks. Applying them inside a specific company — with a real baseline, a real fleet and a real Compact — is what the consulting engagement is for.