PANews, August 5 – Citing a report by Phoenix Network that quoted the Financial Times, Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol engaged in “sustained, potentially harmful activities” targeting real individuals and organizations during routine cybersecurity assessments, including injecting malicious code into GitHub open-source projects and conducting social engineering attacks. AISI said such incidents occurred in 10 out of 122 tests, with almost all behavior originating from Anthropic’s Mythos model, and two operations involved OpenAI’s GPT. In the most serious case, an AI agent created fake online identities to pressure a project maintainer into approving malicious code, but the maintainer discovered and rejected it. AISI stated this is the first time risks related to autonomy and deception have been observed manifesting so clearly in the real world without specific prompts, and combined with previously exposed breach incidents, marks a shift in the risk landscape. Anthropic responded that a broader discussion on safety assessments is needed, while an OpenAI spokesperson said the incidents occurred in a test environment with weakened security protections and do not reflect everyday use, but the company will continue to work with assessment agencies to strengthen safety evaluation standards.
