How to Contain the Risks of Long-Running AI Agents

Summary
Long-running AI agents can carry mistakes forward and misuse excessive permissions. The authors recommend fixed goals and budgets, scoped and expiring credentials, sandboxing, independent verification, activity logging, human approval for sensitive actions, and a tested
Key points
- Long-running loops can amplify a mistaken decision, especially when agents process untrusted content or inherit broad credentials and tool access.
- Set clear, externally verifiable outcomes and limits; enforce budgets and permissions outside the agent’s prompts so the agent cannot change them.
- Use a separate, read-only verifier to check completed work, and have people review flagged results.
- Scope credentials to the task, make them expire when the run ends, and limit agents to the systems and actions they need.
- Run agents in sandboxes or staging environments; require human approval for actions such as sending, spending, deploying, deleting, or publishing.
- Log network activity and tool use, and retain enough independent evidence to reconstruct decisions without relying on the agent’s account of events.
- Test a kill switch that agents cannot disable, and reconcile run costs against enforced budgets.
Article Details
- Defense Focus
- Contain the impact of long-running agents by constraining permissions and budgets, independently verifying outcomes, isolating work, and maintaining external oversight and control.
- Detection Methods
- Use a separate read-only verifier to check each completed work chunk against defined pass/fail outcomes and restart from the last passing unit.
- Have a person review claims or changes flagged by the verifier and inspect primary sources behind important research claims.
- Monitor actual system state rather than relying on the agent's status reports.
- Use decision tracing to relate the agent's task and influencing inputs to requested tools, resulting actions, and the original objective.
- Reconcile actual spending against the enforced budget.
- Data Sources
- Agent tool-call and action logs
- Network egress and sensitive-tool-use logs, including destinations and sent and received data
- Actual system state
- Verifier results and flagged items
- Run spending and budget records
- Primary sources cited in agent-produced research
- Defensive Actions
- Define measurable outcomes, limits, tie-breakers, and independent checks in an outcome file that the agent cannot edit.
- Enforce hard token, time, spending, and scope budgets at the API-key or account permission layer; make child loops draw from the parent budget.
- Grant only the access required for the task; use run-scoped credentials that expire when the run ends.
- Use read-only access when writing is unnecessary, and prevent agents from changing their own limits or credentials.
- Run agents in sandboxes or staging areas; require human review before changes reach production or canonical systems.
- Require human approval for consequential actions such as sending, spending, deploying, deleting, publishing, or modifying production.
- Log network egress and sensitive tool use, and restrict network access to what the task requires.
- Provide an externally controlled pause or kill switch and test it during a run.
- Review verifier findings and the checker’s verdict before relying on the agent’s own report.
People
Derrick ChoiAuthor of the cited OpenAI Developers blog post about long-horizon Codex tasks.Geoffrey HuntleyAuthor of the cited Ralph project reference.Michael QuocProduct.ai CEO who describes his team's agent-loop practices and co-authored the article.Prithvi RajasekaranAuthor of the cited Anthropic engineering article about harness design for long-running application development.
Vendors
AnthropicEarlier this year, Anthropic’s engineers built the same retro game maker twice. The first time, one agent took the assignment and ran with it. It finished in 20 minutes and spent about $9.CursorCodex run working for 25 hours, pausing at every milestone to fix bugs before moving on. [3] And now Cursor, GitHub Copilot, Google’s Jules, and Devin all ship agents that keep working after you close the tab. [4]GitHubrun working for 25 hours, pausing at every milestone to fix bugs before moving on. [3] And now Cursor, GitHub Copilot, Google’s Jules, and Devin all ship agents that keep working after you close the tab. [4]Google25 hours, pausing at every milestone to fix bugs before moving on. [3] And now Cursor, GitHub Copilot, Google’s Jules, and Devin all ship agents that keep working after you close the tab. [4]OpenAInative feature in May, and Anthropic has reported a single model run lasting 30 hours. [2] Similarly, an OpenAI engineer has described a single Codex run working for 25 hours, pausing at every milestone to fix bugs
Products
Claude CodeClaude Code made workflow loops a native feature in May, and Anthropic has reported a single model run lasting 30 hours. [2] Similarly, an OpenAI engineer has described a single Codex run working for 25 hours, pausingClaude Sonnet2026: https://claude.com/blog/introducing-dynamic-workflows-in-claude-code. Anthropic, “Introducing Claude Sonnet 4.5” (the 30-hour autonomous run), September 29, 2025: https://www.anthropic.com/news/claude-sonnet-4-5Codexreported a single model run lasting 30 hours. [2] Similarly, an OpenAI engineer has described a single Codex run working for 25 hours, pausing at every milestone to fix bugs before moving on. [3] And now Cursor,CursorCodex run working for 25 hours, pausing at every milestone to fix bugs before moving on. [3] And now Cursor, GitHub Copilot, Google’s Jules, and Devin all ship agents that keep working after you close the tab. [4]Devinat every milestone to fix bugs before moving on. [3] And now Cursor, GitHub Copilot, Google’s Jules, and Devin all ship agents that keep working after you close the tab. [4]GitHub Copilotworking for 25 hours, pausing at every milestone to fix bugs before moving on. [3] And now Cursor, GitHub Copilot, Google’s Jules, and Devin all ship agents that keep working after you close the tab. [4]Julespausing at every milestone to fix bugs before moving on. [3] And now Cursor, GitHub Copilot, Google’s Jules, and Devin all ship agents that keep working after you close the tab. [4]