How to Contain the Risks of Long-Running AI Agents

· Original article ↗

Summary

Long-running AI agents can carry mistakes forward and misuse excessive permissions. The authors recommend fixed goals and budgets, scoped and expiring credentials, sandboxing, independent verification, activity logging, human approval for sensitive actions, and a tested

Key points

  • Long-running loops can amplify a mistaken decision, especially when agents process untrusted content or inherit broad credentials and tool access.
  • Set clear, externally verifiable outcomes and limits; enforce budgets and permissions outside the agent’s prompts so the agent cannot change them.
  • Use a separate, read-only verifier to check completed work, and have people review flagged results.
  • Scope credentials to the task, make them expire when the run ends, and limit agents to the systems and actions they need.
  • Run agents in sandboxes or staging environments; require human approval for actions such as sending, spending, deploying, deleting, or publishing.
  • Log network activity and tool use, and retain enough independent evidence to reconstruct decisions without relying on the agent’s account of events.
  • Test a kill switch that agents cannot disable, and reconcile run costs against enforced budgets.

Article Details

Defense Focus
Contain the impact of long-running agents by constraining permissions and budgets, independently verifying outcomes, isolating work, and maintaining external oversight and control.
Detection Methods
  • Use a separate read-only verifier to check each completed work chunk against defined pass/fail outcomes and restart from the last passing unit.
  • Have a person review claims or changes flagged by the verifier and inspect primary sources behind important research claims.
  • Monitor actual system state rather than relying on the agent's status reports.
  • Use decision tracing to relate the agent's task and influencing inputs to requested tools, resulting actions, and the original objective.
  • Reconcile actual spending against the enforced budget.
Data Sources
  • Agent tool-call and action logs
  • Network egress and sensitive-tool-use logs, including destinations and sent and received data
  • Actual system state
  • Verifier results and flagged items
  • Run spending and budget records
  • Primary sources cited in agent-produced research
Defensive Actions
  • Define measurable outcomes, limits, tie-breakers, and independent checks in an outcome file that the agent cannot edit.
  • Enforce hard token, time, spending, and scope budgets at the API-key or account permission layer; make child loops draw from the parent budget.
  • Grant only the access required for the task; use run-scoped credentials that expire when the run ends.
  • Use read-only access when writing is unnecessary, and prevent agents from changing their own limits or credentials.
  • Run agents in sandboxes or staging areas; require human review before changes reach production or canonical systems.
  • Require human approval for consequential actions such as sending, spending, deploying, deleting, publishing, or modifying production.
  • Log network egress and sensitive tool use, and restrict network access to what the task requires.
  • Provide an externally controlled pause or kill switch and test it during a run.
  • Review verifier findings and the checker’s verdict before relying on the agent’s own report.

People

Vendors

Products

Tools

Related Articles