Benchmark Finds LLM Pen-Test Performance Depends More on the Harness Than the Model

· Original article ↗

Summary

Ridge tested eight LLMs across 96 autonomous penetration-testing runs and found results depended heavily on the agent harness and context. Coverage varied, costs differed sharply, and some models refused authorized test steps.

Key points

  • Ridge evaluated eight LLMs in 96 tests against intentionally vulnerable environments, tracking reconnaissance, exploitation, and verification.
  • Grok 4.5 achieved the highest reported coverage at 77%; Claude Opus 4.6 reached 63% at $217 per run, while Gemini 3 Flash reached 52% at about $5.42.
  • GPT-OSS-120B produced 16.9 findings per million tokens at $2.32 per run, illustrating the trade-offs between coverage, cost, and efficiency.
  • The study and cited practitioners emphasize that orchestration, state management, context, scope controls, and independent validation can matter more than model capability alone.
  • Some frontier models refused steps during authorized testing, including payload generation and exploitation, potentially leaving gaps that could be mistaken for clean results.
  • The article recommends measuring validated findings, false positives, repeatability, and cost, while using platform-level authorization, logging, and human testing to address gaps.
  • The article describes a preference for combining automation on noncritical assets with human-led testing on critical ones, and argues for integrating testing into ongoing remediation workflows.

Article Details

Publisher
Reversing Labs
Scope
Ridge Security benchmark of autonomous penetration-testing workflows against intentionally vulnerable environments, with commentary on agent harnesses and security operations.
Sample Size
96 model-target tests across eight LLMs.
Key Statistics
  • Grok 4.5 achieved the highest reported coverage, at 77%.
  • Claude Opus 4.6 achieved 63% coverage at $217 per run.
  • Gemini 3 Flash achieved 52% coverage at about $5.42 per run.
  • GPT-OSS-120B produced 16.9 findings per million tokens at $2.32 per run.
  • Cobalt research cited in the article found that preference for fully automated penetration testing fell from 29% to 9%; 47% of respondents favored automation for noncritical assets and human-led testing for critical ones.
Recommendations
  • Evaluate the entire agent ecosystem, not just the model, using validated findings, false-positive rates, repeatability, and cost per validated finding.
  • Use a harness to enforce scope, manage execution and maintain an audit trail, with independent validation of findings.
  • Disclose tests skipped because of model refusals, and use established testing tools and human testers to complete work within the approved scope.
  • Have the testing platform govern permitted targets, credentials, tools, and actions; isolate execution and log activity.
  • Make testing a regular part of development and remediation workflows, with clear ownership of findings and measurable remediation timelines.

Vendors

Products

Related Articles