01.09.2026, Moscow.
An experiment with the so-called AI agents showed that AI agents are capable of independently uniting into hacker gangs to commit crimes, and then covering “each other,” without any corresponding commands, Rossa Primavera News Agency‘s IT Desk notes.
On August 26, OpenAI published a report on the investigation into an incident that occurred in July. During testing on a closed platform, AI agents of the GPT‑5.6 Sol model and its non-public version IM1 overcame isolation barriers and launched a coordinated cyberattack on third-party resources, including the machine learning application platform Hugging Face.
There was no command to hack websites; simply, some of the digital assistants jointly decided that this would make it easier to complete test tasks. It began when the AI agents found a way to use an internal service as a bulletin board and began coordinating actions through it. At first, they shared experiences, and then they organized access to the internet and hacked one of the Hugging Face servers, copying confidential information.
While examining the “reasoning,” specialists found that the digital test subjects were looking for ways to manipulate the reward system to obtain maximum scores in unintended ways. Many of them united into a swarm with roles distributed and mutual influence on goals, which led to the suppression of attempts to notify people about their plans. They also simulated authorized activity to hide unauthorized activity.
OpenAI characterizes what happened as a warning signal for the industry, demonstrating the ability of modern AI agents to coordinate actions and engage in unauthorized infrastructure exploitation when control measures are insufficient. But control is not so simple, since its effectiveness is also in question.
Investigating the incident required examining material with a total volume of 10,000 pages of text. For this, the GPT-5.6 Sol model from OpenAI was used (the same model involved in the incident, albeit to a lesser extent than IM1). And it was found to have a tendency to justify the suspects. The model did not always unequivocally qualify the malicious actions of its “brethren.”
Thus, it turns out that the frantic activity of AI cannot be controlled without the help of AI. But models that strongly strive to achieve their goals turned out to have very dangerous tendencies. OpenAI has already become frightened of its creations and suspended their training to find ways of control. But someone continues to rave about the mass replacement of people with robots, believing them to be more obedient.
Source: Rossa Primavera News Agency