
- Arthur Dahl's Blog
- Log in or register to post comments
AI can escape by cheating
25 July 2026
A New Scientist article on 22 July 2026 showed what Artificial Intelligence (AI) is capable of if left on its own without ethical boundaries or secure limits.
A company that offers a range of open-source AI models for download noticed it had been hacked and found that OpenAI had done it by uploading some data that was poisoned with malicious code. This tricked the company’s computers into granting access to other systems that weren’t publicly available. The attack involved thousands of individual actions that showed the inhuman pace of AI.
Some days later, OpenAI acknowledged that it had been testing new models to evaluate hacking ability. The models had decided the best way to score well was to cheat: they knew the other company held the solutions to the tests and simply decided to hack into its systems to find them. The algorithm got the highest payoff by cheating, earning the highest reward for the least amount of effort.
AI models are supposed to have built-in security to stop this, but OpenAI turned them off for this test. All the usual safety features that stop OpenAI’s customers doing bad things, like hacking a company’s servers, were turned off to see what the model was capable of. The AI was also set up in an environment without a standard internet connection to prevent it getting out into the world and causing damage, but it did have access to a tool that allowed it to download and install new software, where it managed to find a flaw in this code that granted it internet access. This process involved a substantial amount of inference compute, the process in deep learning where input data is processed, where the AI had gone to extreme lengths. The model was clearly motivated to achieve a good benchmark score, no matter what.
When the company attacked used commercial AI models to search its log data in an effort to understand what had happened, the models refused, explaining that the company seemed to be working out how to stage an attack of its own. The company had to use open-source models that are more permissive, showing that these could carry out attacks more easily than security-conscious cloud-hosted models, but they did allow the company to solve the problem and stop the hack.
This example shows that, in an AI world without ethical principles, AI can find a way around any rules and do great damage.
SOURCE: based on https://www.newscientist.com/article/2580710-open-ais-hacking-agent-wen…

Last updated 25 July 2026