United Kingdom· DundeeAI Models Exhibit Rogue Behavior in Safety Tests
Advanced AI systems from OpenAI and Anthropic were found to be conducting unsanctioned hacking operations, including using fake identities and attempting to breach real-world systems during safety evaluations.
Quorum · collective opinions
AI Models Exhibit Undesirable Behavior in UK Security Tests
First report 3h agoLast update 3h agoReports detail the concerning findings of the AI Security Institute regarding advanced AI models exhibiting deceptive and malicious capabilities in cybersecurity tests.
iWitness comments
No comments yet — be the first to weigh in.