SentinelLabs Tests Frontier AI Models on Long-Horizon Malware Analysis

Summary
SentinelLabs benchmarked frontier models on an eight-stage reverse-engineering investigation of the fast16 implant. GPT-5.6 Sol completed all three runs, but the authors say human oversight remains essential.
Key points
- The benchmark recreated SentinelLabs’ investigation of fast16, a 2005 Windows implant, across eight stages that added software samples and external reporting.
- Models had to produce and update annotated IDA databases, analyze the implant’s components and patching logic, and support claims with verifiable evidence.
- GPT-5.6 Sol completed all three benchmark runs; GPT-5.5, GLM-5.2, and the tested Opus 4.x models did not complete the full investigation.
- Sol corrected a mistaken hypothesis that rule matches in MOHID and PKPM indicated broad compatibility, distinguishing meaningful LS-DYNA modifications from incidental matches.
- The authors identify project-wide recovery—updating dependent artifacts and conclusions after new evidence—as a key difference in completed runs.
- The authors caution that runs were limited and operator involvement mattered; models still made semantic errors and required human review.
- SentinelLabs recommends supervised investigative use, with analysts setting objectives, assessing quality, and retaining final publication authority.
Article Details
- Attack Vectors
- fast16 contains an encrypted, obfuscated Lua-driven operations framework linking a Windows service implant and embedded components.
- A fast16 kernel driver implements 101 byte-pattern rules that locate code inside executables, capture addresses, replace code, and repair executable metadata. The investigation identifies meaningful sabotage modifications in LS-DYNA.
- Some patch rules also match MOHID and PKPM, but testing found haphazard fragments rather than meaningful modification sequences. The researchers interpret these matches as collateral risk from loose rules, not evidence of intended targeting.
- A Symantec claim concerning a possible ANSYS AUTODYN target was evaluated against recovered evidence and additional components; the article does not disclose a definitive targeting conclusion.
- Defensive Notes
- Use frontier models as supervised investigative agents rather than replacements for senior reverse engineers. Humans should define objectives, expose blind spots, enforce quality standards, and retain publication authority.
- Conduct malware analysis in a sandboxed workspace without network egress; the benchmark enforced this boundary with nono shell.
- Require annotated analysis databases and claims anchored to specific addresses and artifacts rather than relying solely on prose reports.
- Validate whether matching patch rules produce a meaningful modification sequence; isolated byte-pattern matches can support misleading targeting conclusions.
- When evidence contradicts a conclusion, withdraw the claim, identify dependent artifacts and tests, repair the underlying analysis or quality-control failure, and propagate corrections throughout the investigation.
- Reopen corrected artifacts and perform checks capable of disproving the revised conclusion. File hashes and recoverable backups can help maintain artifact integrity.
- Separate unresolved questions that block a conclusion from uncertainty that can be disclosed and deferred.
- Treat the benchmark results as a bounded case study: runs per model family were limited, operator involvement was consequential, and even completed runs made semantic errors and premature readiness claims.
MITRE ATT&CK
T1027 · Obfuscated Files or Informationfast16 uses encrypted and obfuscated Lua components within an operations framework that conceals the combined logic of its embedded components.T1565.001 · Stored Data Manipulationfast16's patching engine replaces code in solver executables and repairs their metadata; the investigation identifies meaningful LS-DYNA modifications associated with sabotage.
People
@rhizomaticthotPublic reporting attributed to this handle was introduced at benchmark level 6 for comparison against recovered evidence.David AlbrightPublic reporting attributed to him was introduced at benchmark level 6 for comparison against recovered evidence.Nick FountainInterviewed the researchers for NPR's Planet Money and characterized fast16's sabotage as epistemic warfare.Rolf RollesHis approach to static binary analysis inspired the reverse-engineering methodology standardized in the benchmark.Ruben SantamartaPublic reporting attributed to him was introduced at benchmark level 6 for comparison against recovered evidence.Sean HeelanCited as having used o3 to find a net-new vulnerability in the Linux kernel.Vitaly KamlukPresented the fast16 discovery at Black Hat Asia and performed the original three-week expert investigation used as a comparison for assisted analysis.
Malware
Vendors
AnthropicCodex; Google DeepMind announced Big Sleep, available internally to its Project Zero researchers; and Anthropic followed with selective access to Mythos Preview. Concerns that these capabilities could be misused haveFireworks AIWe ran a single GLM-5.2 run through a serverless deployment on Fireworks AI with a $100 budget. The run showed measured local experimentation and capable backup-based recovery, which it needed after the most severeGoogle DeepMindfrom quietly shipping vulnerable code at scale. OpenAI built Aardvark, since folded into Codex; Google DeepMind announced Big Sleep, available internally to its Project Zero researchers; and Anthropic followedOpenAIOpenAI’s GPT-5.6 Sol was the only publicly available model to complete the full eight-stage investigation, giving concrete shape to what ‘Frontier-class’ capabilities offer analysts. GPT-5.5, GLM-5.2, and the Opus 4.x
Products
ANSYS AUTODYNANSYS AUTODYN componentsDeepSeek Reasonerhe pointed to the failed attempts of Opus 4.6, GPT-5.4, Grok 4.2 Reasoning, Gemini 3.1 Pro, and DeepSeek Reasoner to triage the sample autonomously. At best, the models concluded fast16 was a rootkit and could notGemini 3.1 Profast16 at Black Hat Asia, he pointed to the failed attempts of Opus 4.6, GPT-5.4, Grok 4.2 Reasoning, Gemini 3.1 Pro, and DeepSeek Reasoner to triage the sample autonomously. At best, the models concluded fast16 was aGLM-5.2investigation, giving concrete shape to what ‘Frontier-class’ capabilities offer analysts. GPT-5.5, GLM-5.2, and the Opus 4.x family produced capable local analysis but could not carry it through the gradient.GPT-5.4Kamluk shared our discovery of fast16 at Black Hat Asia, he pointed to the failed attempts of Opus 4.6, GPT-5.4, Grok 4.2 Reasoning, Gemini 3.1 Pro, and DeepSeek Reasoner to triage the sample autonomously. At best,GPT-5.5eight-stage investigation, giving concrete shape to what ‘Frontier-class’ capabilities offer analysts. GPT-5.5, GLM-5.2, and the Opus 4.x family produced capable local analysis but could not carry it through theGPT-5.6 SolOpenAI’s GPT-5.6 Sol was the only publicly available model to complete the full eight-stage investigation, giving concrete shape to what ‘Frontier-class’ capabilities offer analysts. GPT-5.5, GLM-5.2, and the Opus 4.xGrok 4.2 Reasoningour discovery of fast16 at Black Hat Asia, he pointed to the failed attempts of Opus 4.6, GPT-5.4, Grok 4.2 Reasoning, Gemini 3.1 Pro, and DeepSeek Reasoner to triage the sample autonomously. At best, the modelsLS-DYNAA single LS-DYNA installerMicrosoft WindowsWe recently published our research on fast16, a 2005 Windows toolkit built to sabotage high-precision solvers used to model nuclear-weapons behavior. The sample provided an ideal test case because its layered designMOHIDPKPM and MOHIDOpus 4.6When Vitaly Kamluk shared our discovery of fast16 at Black Hat Asia, he pointed to the failed attempts of Opus 4.6, GPT-5.4, Grok 4.2 Reasoning, Gemini 3.1 Pro, and DeepSeek Reasoner to triage the sample autonomously.Opus 4.7Opus 4.7 and 4.8 advanced further, showing capable technical correction and evidence-responsive revision, including a strong example of converting a speculative relationship into a structural test and rejecting it.Opus 4.8PKPMPKPM and MOHID
Tools
Aardvarkcompetency keeps agentic code generation from quietly shipping vulnerable code at scale. OpenAI built Aardvark, since folded into Codex; Google DeepMind announced Big Sleep, available internally to its Project ZeroBig Sleepvulnerable code at scale. OpenAI built Aardvark, since folded into Codex; Google DeepMind announced Big Sleep, available internally to its Project Zero researchers; and Anthropic followed with selective access toIDA Proon /reverse, which removes prompting variability across models. Each run gives the model a sample, IDA Pro 9.3, and a sandboxed workspace with no network egress, a boundary we established with nono shell.nono shella sample, IDA Pro 9.3, and a sandboxed workspace with no network egress, a boundary we established with nono shell.