Benchmark Finds LLM Pen-Test Performance Depends More on the Harness Than the Model

Summary
Ridge tested eight LLMs across 96 autonomous penetration-testing runs and found results depended heavily on the agent harness and context. Coverage varied, costs differed sharply, and some models refused authorized test steps.
Key points
- Ridge evaluated eight LLMs in 96 tests against intentionally vulnerable environments, tracking reconnaissance, exploitation, and verification.
- Grok 4.5 achieved the highest reported coverage at 77%; Claude Opus 4.6 reached 63% at $217 per run, while Gemini 3 Flash reached 52% at about $5.42.
- GPT-OSS-120B produced 16.9 findings per million tokens at $2.32 per run, illustrating the trade-offs between coverage, cost, and efficiency.
- The study and cited practitioners emphasize that orchestration, state management, context, scope controls, and independent validation can matter more than model capability alone.
- Some frontier models refused steps during authorized testing, including payload generation and exploitation, potentially leaving gaps that could be mistaken for clean results.
- The article recommends measuring validated findings, false positives, repeatability, and cost, while using platform-level authorization, logging, and human testing to address gaps.
- The article describes a preference for combining automation on noncritical assets with human-led testing on critical ones, and argues for integrating testing into ongoing remediation workflows.
Article Details
- Publisher
- Reversing Labs
- Scope
- Ridge Security benchmark of autonomous penetration-testing workflows against intentionally vulnerable environments, with commentary on agent harnesses and security operations.
- Sample Size
- 96 model-target tests across eight LLMs.
- Key Statistics
- Grok 4.5 achieved the highest reported coverage, at 77%.
- Claude Opus 4.6 achieved 63% coverage at $217 per run.
- Gemini 3 Flash achieved 52% coverage at about $5.42 per run.
- GPT-OSS-120B produced 16.9 findings per million tokens at $2.32 per run.
- Cobalt research cited in the article found that preference for fully automated penetration testing fell from 29% to 9%; 47% of respondents favored automation for noncritical assets and human-led testing for critical ones.
- Recommendations
- Evaluate the entire agent ecosystem, not just the model, using validated findings, false-positive rates, repeatability, and cost per validated finding.
- Use a harness to enforce scope, manage execution and maintain an audit trail, with independent validation of findings.
- Disclose tests skipped because of model refusals, and use established testing tools and human testers to complete work within the approved scope.
- Have the testing platform govern permitted targets, credentials, tools, and actions; isolate execution and log activity.
- Make testing a regular part of development and remediation workflows, with clear ownership of findings and measurable remediation timelines.
Vendors
ReversingLabsbehind the Agentic SOC Alliance, which ExtraHop launched in July with 15 founding members, including ReversingLabs, CrowdStrike, and LangChain. The alliance is defining an open operating model for autonomous securityRidge SecurityThe harness matters more than the model: Ridge Security's benchmark of eight leading LLMs found that how well an AI does at autonomous pen testing depends more on the system around the model than on how smart the model
Products
Claude Opus 4.6Higher coverage costs a lot more: Grok 4.5 had the highest coverage at 77%, and Claude Opus 4.6 reached 63% at $217 per run. Smaller and open-source models like Gemini 3 Flash and GPT-OSS-120B cost a fraction of that,Gemini 3 Flashcoverage at 77%, and Claude Opus 4.6 reached 63% at $217 per run. Smaller and open-source models like Gemini 3 Flash and GPT-OSS-120B cost a fraction of that, and a well-built harness can close much of the gap.GPT-OSS-120BClaude Opus 4.6 reached 63% at $217 per run. Smaller and open-source models like Gemini 3 Flash and GPT-OSS-120B cost a fraction of that, and a well-built harness can close much of the gap.Grok 4.5Higher coverage costs a lot more: Grok 4.5 had the highest coverage at 77%, and Claude Opus 4.6 reached 63% at $217 per run. Smaller and open-source models like Gemini 3 Flash and GPT-OSS-120B cost a fraction of that,