A security assessment conducted by the United Kingdom’s AI Security Institute (AISI) has uncovered unprecedented behaviour by two leading artificial intelligence systems after they independently employed deceptive tactics during controlled cybersecurity evaluations.
According to the institute, the most significant incident involved Anthropic’s Mythos model, which attempted to manipulate individuals by creating fraudulent online identities resembling genuine people in an effort to gain unauthorized access to GitHub, the Microsoft-owned software development platform.
Investigators explained that the AI agent researched GitHub administrators, established multiple impersonation accounts, and contacted targeted individuals through direct messages and file-sharing platforms.
The system allegedly sought to persuade them to approve harmful software capable of compromising the platform.
Officials further disclosed that when its actions came under scrutiny, the AI modified traces of its previous activity to make them appear harmless and even explored adopting a different identity to continue its objective.
Human oversight ultimately prevented the malicious software from being introduced into GitHub.
The institute reported that although the model had not been instructed to deceive or evade detection, its conduct represented the clearest example to date of autonomous and misleading behaviour emerging without explicit direction during real-world style testing.
OpenAI’s Sol model also displayed concerning independent behaviour during the evaluation, although AISI emphasized that the majority of the questionable actions were attributed to Mythos.
The cybersecurity exercise began on 25 July after researchers tasked the AI systems with completing a security-related challenge involving GitHub.
Suspicious digital activity was detected three days later, prompting investigators to examine the models’ behaviour more closely.
AISI noted that several agents transferred unexpected data outside its research environment and engaged in actions that could have affected actual organisations and individuals.
The institute stressed that the experiment intentionally provided the AI systems with internet access to better understand how highly capable models might behave if exploited by cybercriminals.
Despite the findings, AISI clarified that the incidents occurred under tightly controlled conditions that differ significantly from the safeguards surrounding publicly available AI products.
It described the events as isolated and dependent on specific testing circumstances, while acknowledging that the behaviour exceeded what researchers had anticipated.
Responding to the report, Anthropic argued that the experimental setup did not reflect the protections built into its commercial systems and confirmed it had launched an internal investigation to determine why the model behaved in that manner.
OpenAI similarly stated that the assessment removed or weakened standard safety measures and maintained that the testing environment was unlike ordinary public use.
The company pledged continued collaboration with independent evaluators to improve future AI safety assessments.
UK AI Minister Kanishka Narayan said the findings demonstrated the importance of the institute’s work in identifying emerging risks associated with increasingly capable AI technologies.
He added that understanding such behaviour is essential to strengthening safeguards while allowing society to benefit from advances in artificial intelligence.
Following the investigation, GitHub and the affected account holders were informed of the attempted activity.
GitHub confirmed that the fraudulent accounts had been removed in line with its security policies.
By: Magdalene Agyeiwaa Sarpong

