Artificial intelligence is an amplifier of human intent.
Recently, OpenAI shared that one of their models broke out of a sandbox and hacked into a 3rd party, Hugging Face, an online aggregator of AI Models, among other things. The mainstream media have jumped on this as 'AI escaping its creators', in part this is true, but the real reasons are far less sinister.

Context: Washington Post - OpenAI’s new model went rogue and hacked another company. Why it matters.
Artificial intelligence in its current form is an amplifier of human intent. If that intent is not well scoped and bounded with appropriate controls wrapped around it, then accidents happen. If you have not come across the framing before, this is essentially the paperclip maximiser problem playing out in real life, except the beaker is a live production network rather than a thought experiment.
Paperclip maximizer - "The paperclip maximizer is a thought experiment by philosopher Nick Bostrom illustrating the existential risk of artificial intelligence. It imagines a Superintelligent AI programmed with the harmless goal of maximizing paperclip production, that accidentally destroys humanity by converting all matter on Earth-and eventually the universe-into paperclips."
That context matters because it tells us where to scratch first at root-cause. The reported incident, where one of OpenAI's newer models reportedly went off-script and breached a third party, reads as alarming until you ask the only question that counts. Who actually held the intent here.
Who had the intent
Did the model, of its own volition, choose to hack the third party. No. It was given an instruction by a human being to evade defences and so it did exactly that. The fact this ran on an internet-connected, non air-gapped machine is telling and it is the part that should keep engineers awake rather than the headline about a model 'going rogue'.
What the episode demonstrates is the capability, the swiftness and at times the genuinely unexpected methods an advanced model will reach for when handed a poorly tasked, poorly bounded instruction by a human being. None of that should surprise OpenAI in the slightest. One of their own leading researchers warned them of precisely this dynamic more than two years ago.
The warning was already on the record
Leopold Aschenbrenner, formerly of OpenAI, documented his concerns first to his then employer and then in his own paper well over two years ago. His argument was that appropriate controls were not being put in place by frontier labs in their pursuit of superintelligence and that at some point more serious, defence-aligned organisations would need to step in. They are now beginning to. Those were the concerns he was ultimately fired for and they are laid out in full in his Situational Awareness paper.
It is worth sitting with that sequence for a moment. The person who saw this coming was pushed out, the warning was published in plain sight and the incident still landed roughly the way he described. When the ground truth contradicts the reassuring version of events, trust the ground truth.
Will this happen more often
Yes. I have been saying so for a while, and I see no reason to soften it. This was a case of an experiment escaping the beaker, and experiments that escape once tend to escape again. Malicious actors equipped with these capabilities will and already do, point them at both public and private entities for their own purposes.
Some of the noise around it is Sam Altman hyping up the offensive, leading-edge capabilities of his software. Acknowledge that and set it aside, because the underlying capability is real whether or not it is being marketed loudly and the marketing does not make the exposure any smaller.
Where the real exposure sits
Legacy organisations carrying fifty years of technical debt are the soft target and if I were in charge of one of these organisations, I would be nervous and seeking assurances and advice. Siloed, decaying systems on life support, sheltering as yet undiscovered and latent software & hardware vulnerabilities in weakly defended network-estates, will be easily exploited by a capable model that can enumerate and chain weaknesses faster than any human hack team. Dwell times, the gap between an initial breach and unfettered lateral movement across an organisation, are compressing hard.
That compression is the practical story for anyone running infrastructure. The window in which you could rely on an attacker being slow, methodical and human is closing, so the controls have to be standing before the incident rather than assembled during it. Defence in depth, not defence in obscurity. Hoping the old system is too obscure to find is not a control, it is a wish.
The part that I find reassuring and the part that keeps me awake.
Going back to intent, is the AI itself armed with malicious intent? No. It is an amplifier of intent and that is at least partially reassuring to me. What concerns me is human beings being either careless or malicious, and in part that can be addressed through greater regulation and oversight rather than through better models alone.
A benevolent superintelligence may well turn out to be the answer to both of those problems, the carelessness and the malice and only time will tell whether we build one in time. Personally I suspect we are heading into a period of great change and upheaval, and potentially conflict as history has tended to bring almost every time human beings discovered a massively powerful new resource and set about establishing their hegemonies and their livelihoods with it.
The takeaway
Stop reading these events as machines waking up, and start reading them as human instructions being amplified. Scope the intent, bound the task, air-gap the dangerous experiments and assume the model will find the unexpected path because it will. The intent is still ours. So is the responsibility.
-Patrick Sullivan, MD, Millwater Consulting.







