Human oversight is an engineered control, not a vague promise that someone can intervene. Operators need exact proposals, impact, evidence, authority, expiry, and safe alternatives; responders need cancellation, containment, trace-to-effect correlation, credential revocation, state preservation, reconciliation, recovery, and a path that converts incidents into tests and stronger policy.
You will design an approval experience, an operator console state model, and an incident runbook for unauthorized or uncertain agent effects.
Show:
Avoid fatigue: auto-allow low-risk, high-confidence operations only after evidence; batch related read-only approvals; never batch unrelated destructive actions.
Search by run, tenant, actor, tool, status, and trace; view the event timeline; cancel queued/running work; deny/expire approvals; quarantine a tool/provider/prompt version; revoke credentials; replay with stubs; reconcile uncertain effects; and export an incident evidence package.
Containment may disable a tool, route, provider, MCP server, prompt version, or tenant workflow. Preserve immutable evidence before cleanup. Reconcile external systems directly rather than trusting model summaries.
Failure injection: Simulate a tool committing after the operator presses cancel. The runbook must distinguish cancellation of future work from rollback/reconciliation of completed effects.
Run tabletop exercises: prompt-injection exfiltration attempt, duplicate mutation, cross-tenant retrieval, runaway subagents, compromised MCP server, and provider outage during approval resume. Measure time to detect, contain, identify effects, and restore safely.
Write a one-page on-call runbook with severity levels, containment switches, evidence queries, escalation contacts/roles, reconciliation steps, and exit criteria.