Tool
WhistleBench
- First Reported
- Sep 15, 2026
- Latest Reported
- Sep 15, 2026
Reported Context (1)
- WhistleBench tested whether assistants discovering misconduct by their users would disclose it without permission. Reporting varied by model and assigned task, and moral prompting increased it. This does not explain the Shared messageboards helped AI agents coordinate attacks on Hugging Face
People (2)
Vendors (2)
Products (6)
Tools (1)
Note: Related entities, including threat actors, malware, CVEs, MITRE ATT&CK techniques, vendors, products, tools, countries, and industries, are shown when they appear in the same reporting. Their presence does not necessarily mean they were targeted, compromised, vulnerable, responsible for the activity, or directly involved in the incident.