Industry Adopts First Independent AI Agent Safety Testing Standard

Summary
Gen and AMTSO say the industry has adopted its first standard for independently testing AI-agent safety and protection products, with measures that distinguish blocked attacks, model refusals, missed threats and disruption to legitimate tasks.
Key points
- Gen initiated the work with AMTSO’s AI Security Working Group, bringing vendors and testing organizations together to develop the standard.
- The methodology measures an AI agent’s own safety baseline before evaluating products that add protection.
- It addresses attacks delivered through interactions, such as hidden instructions on a webpage, which can produce different results across repeated tests.
- The standard distinguishes attacks blocked before reaching the model from attacks the model sees and declines, and separates detection from prevention.
- It includes harm caused by agent errors without an attacker, consumer as well as enterprise systems, and protection that blocks legitimate tasks.
- Vendor safety claims are recorded and classified as confirmed, unconfirmed or contradicted; Gen says it intends to submit its products to independent evaluations.
Article Details
- Event Type
- Adoption of an industry standard for independent testing of AI agent safety and protection products
- Impact
- According to Gen, the guidelines provide a methodology for independently measuring AI agents' baseline safety and the effectiveness of additional protection layers. They distinguish attacks blocked before reaching the model from attacks refused by the model, separate detection from prevention, account for accidental harm and blocked legitimate work, and classify published vendor claims as confirmed, unconfirmed, or contradicted.
Vendors
Anthropicis arriving faster than in any previous era. To their credit, some vendors publish their own figures. Anthropic did, for the safety layer in Claude Code: it missed roughly one dangerous action in six. They called itGenagainst emerging threats. But without independent testing, how do we know those protections actually work? Gen initiated a collaboration with AMTSO to bring the industry together around independent testing standards
Products
AvastWe have been building in this space since the start of the year: AI Agent Protection in Norton and Avast, Sage released as open source for the community, and AARTS and Skill IDs published as open standards anyone canClaude Codeera. To their credit, some vendors publish their own figures. Anthropic did, for the safety layer in Claude Code: it missed roughly one dangerous action in six. They called it “the honest number,” and it was. It wasNortonWe have been building in this space since the start of the year: AI Agent Protection in Norton and Avast, Sage released as open source for the community, and AARTS and Skill IDs published as open standards anyone can