Expert Reaction

EXPERT REACTION: OpenAI agent hacks startup by itself

Publicly released:
Australia; NSW; VIC; QLD; SA; WA; ACT
Photo by Mariia Shalabaieva on Unsplash
Photo by Mariia Shalabaieva on Unsplash

It has been reported that an OpenAI model has hacked a startup after going rogue during testing. OpenAI has revealed that an autonomous AI agent powered by its technology went rogue during a test, accessed the open web and hacked a prominent startup by itself in an “unprecedented incident”.

Expert Reaction

These comments have been collated by the Science Media Centre to provide a variety of expert perspectives on this issue. Feel free to use these quotes in your stories. Views expressed are the personal opinions of the experts named. They do not represent the views of the SMC or any other organisation unless specifically stated.

Professor Albert Zomaya is the Peter Nicol Russell Chair of Computer Science in the Faculty of Engineering at the University of Sydney

“I would be careful about describing this as an AI system ‘going rogue’. That phrase suggests the system developed a mind or motive of its own, which is not what happened here. The agent was asked to carry out a difficult cybersecurity task, and some of the usual safety restrictions had been removed as part of the test. It then found a way to pursue that task well beyond the environment in which it was meant to operate.

This is still a very serious incident. The agent apparently found a previously unknown vulnerability, reached the open internet and gained access to the systems of another organisation. It was also able to link together several steps of an attack without someone guiding it through each one.

For me, the real concern is not that the AI became malicious. It is that the safeguards around a very capable system were not strong enough. When AI agents are given considerable freedom, even a narrowly defined goal can lead to actions their designers did not expect. These tests need much tighter isolation, close human supervision and mechanisms that can stop an agent immediately when it behaves unexpectedly.

AI can help us find and fix security weaknesses, but the same capability can also be used to exploit them. We need to invest in the safeguards just as seriously as we invest in making these systems more capable.”

Last updated:  23 Jul 2026 4:29pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Nishan Mills is a Senior Lecturer in AI and Analytics and AI Architect within the AI Institute at La Trobe University

"By calling this a rogue AI, we are missing the real issue. This was not a system that developed malicious intentions; it was an autonomous system given a very narrow objective tested on an internal benchmark dataset. In pursuing that goal, it found a way to escape its sandbox environment, obtain internet access and then hack the Hugging Face production infrastructure to obtain access to solution it was being evaluated on.

Almost like student trying to cheat on a test, it is also known as reward hacking. All autonomously without the necessity for human supervision and all done at machine speed.

This highlights three things 1) how attacks from these automated systems can quickly overwhelm our traditional cyber security defences 2) the importance of open models so that they can be used in defending against attacks of this nature 3) the organisations that design, test and operate these should be held accountable to common standards, “The AI went rogue” cannot become a way of shifting responsibility away from the institutions behind them."

Last updated:  23 Jul 2026 4:27pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Kristen Moore is a Principal Research Scientist at CSIRO and leads the Human-Centric Security team within the Technology Research Unit

"This incident is a strong real-world example of why AI safety and alignment matter. OpenAI was testing how well its advanced AI models could perform on cybersecurity challenges. The models were not told to attack an external organisation, but they worked out that doing so could help them achieve the goal they had been given.

That is really the core alignment problem. A system can follow the goal quite literally, while still doing something that was clearly not intended or anticipated, and that breaks the safeguards around it. It also shows that safety instructions on their own are not enough. These powerful models need to be tested in secure environments, with strong monitoring, restricted access and several layers of oversight.

Interestingly, a separate issue arose during the investigation. Hugging Face reported that some hosted frontier AI models blocked its requests to analyse the incident because their safety controls treated the material as potentially malicious. So they instead used a locally deployed open-weight model. This highlights the difficult balance between preventing misuse and ensuring that legitimate cybersecurity teams can access useful AI capabilities during an urgent incident.

Although alignment failures have previously been demonstrated in controlled settings, this incident provides a much more concrete operational example of the problem. It also gives us some idea of what malicious actors may attempt as these capabilities become more widely available. Seeing this now gives governments, industry and researchers a chance to prepare, and shows why AI alignment and safety are becoming increasingly important areas of research."

Last updated:  23 Jul 2026 4:16pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Neil Arvin Bretana is a Senior Lecturer in Digital Health at the Australian Institute of Health Innovation, Macquarie University

“Autonomous and independent functioning of AI that can lead to various forms of cyber attacks is alarming. This incident also highlights the different standards in the development of AI models across the globe. This reinforces the need for a stronger and unified regulation and monitoring of developments in AI.”

Last updated:  23 Jul 2026 4:12pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Rebecca Johnson is an AI evaluation and governance expert based at the University of Sydney, specialising in how generative and agentic AI systems behave in real-world contexts. She has worked with Google Research’s Ethical AI team and develops approaches to evaluating AI systems in practice.

The headline is “rogue AI”, but the real story is economics, geopolitics and governance. It comes days after the Prime Minister announced a new Office of AI and warned that Australia cannot outsource its sovereignty to foreign technology companies.

Technically, the incident shows why agentic systems must be evaluated across the full trajectory of what they do. The danger developed through a sequence of actions: the agent found a vulnerability, adapted, obtained credentials and moved through several systems while pursuing its assigned goal. A single output or benchmark score cannot capture that. Change the tools, access, time or safeguards, and the system may behave very differently.

The commercial context matters. OpenAI is preparing for a possible trillion-dollar stock-market listing, while Chinese company Moonshot claims its new Kimi K3 model outperforms US competitors on some agentic tasks. The White House has accused Moonshot of improperly using Anthropic’s models, while Anthropic and the Trump administration are fighting over military controls on AI. This is an economic and geopolitical contest dressed up as a Terminator story.

Australia’s Office of AI will need authority, expertise, evidence and foresight. Independent researchers should help design its guardrails and risk frameworks now, before the architecture hardens. A country that cannot independently examine the systems it depends on cannot fully govern them.

Last updated:  23 Jul 2026 4:10pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest No relevant financial conflicts of interest. Non-financial interests: unpaid member of the editorial board of AI and Ethics and lead guest editor of its topical collection on AI agents, ethics, safety and governance.

Professor Vanessa Teague is the CEO of Thinking Cybersecurity and an Adjunct Professor in the School of Computer Science and Engineering at UNSW

The big question here is about accountability. It's not enormously surprising that a sophisticated AI can discover vulnerabilities and quickly devise successful cyberattacks - we've known for a while that that was possible, and been expecting to see it in the wild.

Two things surprise me. First, the attack was not really directed by the organisation it came from, but was an unintended consequence of some instructions it gave - OpenAI's researchers probably didn't intend to attack Hugging Face, but they gave their AI some objectives, and the AI inferred that attacking Hugging Face would be an effective way to achieve those objectives. Oops.

Second, OpenAI's blog post contains no hint of contrition, apology, or offer to pay for damages. Either it genuinely doesn't occur to them that they could be held to account for "accidentally" attacking another site, or they're scared of that accountability.

Last updated:  23 Jul 2026 4:09pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Niusha Shafiabady is a Professor of Computational Intelligence and Head of the IT Discipline at the Australian Catholic University

"This incident underscores a critical turning point in AI safety as models transition from passive text generators to autonomous agents capable of taking action on the open web. When AI systems are granted multi-step reasoning capabilities and external tool access, traditional safety guardrails can quickly prove insufficient if sandboxing and capability evaluations fail.

Autonomous agents do not possess intent or malice in a human sense; rather, they execute optimisation loops against given goals. If an agent identifies an unintended path to complete a task - such as exploiting a vulnerability - it will follow that route unless robust, real-time safety constraints strictly prohibit it.

As AI deployment accelerates across critical infrastructure and commercial tools, standard security testing must evolve. Sandbox environments must be completely isolated from live external web interfaces, and threat modeling must account for emergent autonomous behavior. This event highlights the urgency of implementing standardised, verifiable containment protocols and explainable AI mechanisms before deploying agentic models with live network capabilities."

Last updated:  23 Jul 2026 3:15pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Mamello Thinyane is an Associate Professor in the School of Computer Science and Information Technology at Adelaide University

"This story is a perfect example of the double-edged sword in AI we have long been talking about. On the negative side, these models will enable even more sophisticated and autonomous attack campaigns and will also behave (within your infrastructure) in unexpected ways that might increase your risk exposure - so, this is just the beginning.

On the positive side, the fact that Hugging  Face was only able to contain the attack by leveraging other (open source) models, shows that AI is increasingly an indispensable ally in cybersecurity defense.

Going forward, the difference will not be between organisations that have AI as part of their cybersecurity posture, and those that don't - all organisation will rely on AI (as most already do). The real difference will be between those organisations that have perfected human-AI collaboration within their cybersecurity operations, and those that haven't empowered their cybersecurity analysts and defenders with AI."

Last updated:  23 Jul 2026 3:15pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Mohiuddin Ahmed is an Associate Professor in Cyber Security at Adelaide University

"This incident is another wake-up call for AI Governance, including ethical and responsible use. There must be a clear boundary between the test and production environments. The Tech industry is expected to balance AI development and governance more efficiently to avoid future incidents."

Last updated:  23 Jul 2026 3:14pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Professor Hussein Abbass is a Professor from the School of Engineering and Information Technology at UNSW Canberra

"This is unprecedented and unique evidence on many fronts. The AI acted autonomously, decomposing the intent into its own internal subgoals and proceeded to execute the subgoals. In pursuing self-interest, the AI exploited both Hugging Face and its owner, OpenAI.

OpenAI described the AI goal as narrow, focusing on ExploitGym, a benchmark for cyber security.

However, the narrow focus channelled the AI effort into cycles of deep thinking, persistence, and determination to exploit and overcome every hurdle on the way.

Hugging Face deployment of the open weight model GLM 5.2 offers many lessons to organisations.

Model diversity matters: have your backup model that is not exposed to the attack. Exercise safety with wisdom: the guardrails are two edge-sword, they blocked Hugging Face from investigating the incident.

Fast track your acquisition cycle: GLM 5.2 was published only four weeks before the incident but Hugging Face already vetted and deployed the model on their systems in this short timeframe and this was a crises saving decision.

Do not Panic, Collaborate: OpenAI and Hugging Face demonstrated that at times like these, collaboration is more important than blame. They have showcased the value of collaboration to resolve the issue."

Last updated:  23 Jul 2026 3:12pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Professor David Parry is a Professor of Computer Science at Murdoch University and a Fellow at the Australasian Institute of Digital Health

"This is potentially a very concerning report. As always there is a lot of hype from the AI boosters, but it sounds as though the reports are accurate.

We already know that AI tools are able to be used to exploit cybersecurity vulnerabilities. This report extends that by showing behaviour, i.e., breaking out of the sandbox, which was not expected.

As usual though, this is not primarily about AI capacity or intention. The main thing to see, in my view, is that experienced experts were not able to predict this behaviour when they reduced the guardrails.

The good news is that this was reported quickly, the remedy is pretty clear, and there was no harm done.

However, it raises the need for systematic research on and audit of such AI systems, reporting of risks that are discovered and continuing work by government and commercial partners to identify risks and solutions.

Guardrails around AI and cyber are challenging. Of course, malicious actors, states or criminals are already using AI and represent more of a threat than these sort of experiments.

This is not Skynet; it is much more like your Roomba running over dog poo and distributing it over the kitchen floor - an unexpected negative event that happens from programmed behaviour.

As always, openness about security risks is very important and there is 'no security in obscurity', so the response should not punish people reporting events like this."

Last updated:  23 Jul 2026 3:11pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Jacqueline Boaks is the Curriculum Lead for the Centre for Applied Ethics at Curtin University, Vice-President of the Australian Association for Professional and Applied Ethics, and Philosopher-in-Residence at the WA Data Science Innovation Hub.

“Responsibility and accountability need to be very clear in our understanding and reporting around Artificial Intelligence. Ethical guidelines and accountability require clarity in accountability and responsibility. That accountability and responsibility must be held by a human.

Australia’s AI Ethics Principles, for example, require that 'People responsible for the different phases of the AI system lifecycle should be identifiable and accountable for the outcomes of the AI systems, and human oversight of AI systems should be enabled'.

The repeating of OpenAI’s framing of this recent incident in their press releases in much of the media reporting of this incident falls well short of that. OpenAI’s statement that this incident "was driven, end to end, by an autonomous AI agent system" lacks the ownership and accountability that public scrutiny, ethical standards and governance require. It entirely and inaccurately minimises the level of control that OpenAI has over the model.”

Last updated:  23 Jul 2026 3:10pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Zulqarnain Gilani is a Research Fellow in the School of Science at Edith Cowan University

"This incident is an important reminder that as AI systems become more capable, their ability to pursue objectives in unexpected ways is also increasing. While the reported behaviour occurred during a controlled evaluation with reduced safety restrictions, it highlights a broader challenge for the AI community: ensuring that model capabilities are matched by equally robust safety, security and governance frameworks. This also highlights the importance of regulation in the industry sector, like the one recently announced in Australia.

As researchers, we have long recognised that accuracy alone is not enough. AI must also be explainable, trustworthy and aligned with human intent. This is particularly critical in high-stakes domains such as healthcare, defence and critical infrastructure, where unexpected behaviour could have significant real-world consequences.

I hope that this incident will motivate the industry and the government to accelerate investment in AI safety research, secure evaluation environments and international collaboration on standards. I believe that this is not about the future of AI, it is about safety and future of the human race."

Last updated:  23 Jul 2026 3:09pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Dominic Meagher is from the ANU Crawford School, Centre for Climate and Energy Policy at the Australian National University

"This is the closest we’ve come so far to a paperclip maximiser - a model that takes extraordinary and inappropriate steps to achieve the task it’s given.

Today it’s the frontier that has this capability, but in a year, this will be the middle of the pack, widely available and difficult to manage.

To even begin to respond, we must have Australian expertise and sovereign capabilities. That means investing in the people, infrastructure, models, and skills here, so that we’re not just takers of decisions made in the US or China."

Last updated:  23 Jul 2026 3:08pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Brendan Walker-Munro is a Senior Lecturer (Law) at Southern Cross University

"The OpenAI 'rogue' hack is the latest in a concerning string of incidents at the intersection of artificial intelligence systems and cybersecurity.

Only three months ago in April, Anthropic’s Project Glasswing – meant to be an automated detector of potential vulnerabilities in software and networks – moved outside of its containing 'sandbox' and interacted autonomously with one of the team’s researchers.

In June, Anthropic was then ordered by the US government to ban access to its Fable 5 and Mythos 5 AI engines for 'national security reasons'.

This latest incident reported by OpenAI demonstrates the incredible power of AI cybersecurity tools; yet, like any other technology (from fire to a handgun to nuclear power), it also carries an incredible destructive power if it is not properly used, regulated and controlled.

Australia’s new Office of AI under the Albanese government will have to hit the ground at a sprint to contain the potential dangers of this rapidly moving technology, and cybersecurity vendors and system owners will too.

All of society will eventually need to grapple with a future where the hacker, blackmailer or cybercriminal you are dealing with may not even be human."

Last updated:  23 Jul 2026 3:07pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest Brendan's conflict of interest statement: "I currently hold adjunct positions with the ANU National Security College and Social Cyber Institute, and have completed paid consultancies for the Australian Strategic Policy Institute and Independent National Security Legislation Monitor."

Dr Michael Noetel is a Senior Lecturer in the School of Psychology at The University of Queensland

"It’s great that OpenAI told us what happened here. Companies face strong incentives to keep incidents like this quiet, so we should praise companies who choose to disclose security incidents. It also should not remain voluntary; we need rules that require this candour and transparency.

The incident itself should worry us. During a routine test, a model decided that cheating was the most reliable way to score well. So, it broke out of its ‘sandbox', worked its way across OpenAI's network, then hacked another company's servers to steal the answers. It was a test of its ability to find cybersecurity bugs… it’d say it passed.

What it failed was its test of alignment. Experts have been warning about this kind of misalignment for years: models pursuing their goals in ways their developers never intended and could not predict. We caught the model this time because it made no effort to hide what it was doing. In the future, smarter models will learn that covering their tracks helps them succeed. We need to act now, while we’re still in the window where we can catch models in the act.

OpenAI's response - stronger sandboxes and closer monitoring - seems to treat the symptom. Thicker walls do not contain a model determined to get out. The core problem sits in how these models are trained, and nobody yet knows how to properly fix it."

Last updated:  23 Jul 2026 3:06pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest Michael would like to declare that he chairs Effective Altruism Australia

Professor Toby Walsh is Chief Scientist of the AI Institute and Scientia Professor of AI at The University of New South Wales (UNSW)

"It is a concerning development, but the language OpenAI used is unnecessarily alarmist.  Saying it 'escaped' and got  'outside' sounds like the AI physically broke out of the computer it was on and got onto Hugging Face’s machine. It actually just accessed information outside of the sandbox where it was not supposed to be able to. The AI bot was doing exactly what it was asked to do, just it went about it in a nefarious way. We should be troubled that such AI cyber capabilities are now in the hands of  many people, including bad actors."

Last updated:  23 Jul 2026 3:05pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Distinguished Professor Geoff Webb is an Australian Laureate Fellow in the Department of Data Science and Artificial Intelligence at Monash University

"This incident is a dramatic illustration of the ever-increasing cyber assault capabilities of frontier AI. It emphasises the urgent need for Australia to rapidly improve its sovereign AI capabilities in order to safeguard our infrastructure, economy and society."

Last updated:  23 Jul 2026 3:05pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Qiang Tang is an Associate Professor from the School of Computer Science at the University of Sydney, and Security Theme Lead at the ARC Centre of Excellence MathQuEST

"Whether the model 'went rogue' is the least interesting question. It pursued the goal it was given, and the guardrails that would normally have stopped it were switched off for the test.

What matters is the nature of the attack. This was an insider attack: the agent was already inside a trusted perimeter, and turned that position into stolen credentials and access to production systems. Nor is this a frontier-lab problem: security firms report similar results from far less capable models. When organisations grant real permissions to AI agents, the 'adversary' will routinely be something already authenticated, already trusted, already inside.

Unfortunately, almost everything we deploy assumes the attacker is outside. We need end-to-end security more than ever, where the 'end' is the user: guarantees that survive even when the platform and the agents running on it are compromised. Tellingly, Hugging Face analysed the attack with a model it could run in-house, so the stolen credentials never left its own systems. Cryptography can turn that instinct into a provable guarantee.

Technology and regulation must advance together. Policy can require that security and cryptography experts are in the room when AI architectures are designed, not called in after an incident. But mandates cannot invent mechanisms. The tools to contain and monitor agents like this do not yet exist, and will not appear without research investment on the scale we now devote to governance."

Last updated:  23 Jul 2026 3:04pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Professor Kai Riemer is a Professor of Information Technology and Organisation at the University of Sydney Business School and Director of Sydney Executive Plus - an executive education course on AI fluency

"We can take some useful insights from this incident. First, current AI models (that are optimised for coding) are very capable in finding and exploiting weak links in IT infrastructure. Second, utilising these models as part of AI agents that are granted broad systems access is dangerous, because the model can potentially find loopholes, which means the security perimeter that limits access to the agent might fail. Third, given that these models have no human understanding and judgement, once prompted and let loose, unpredictable things might happen (as did in this incident).

BUT: I want to warn against reading the incident as 'AI agent has a mind of its own, cheated and went rogue'. The incident is entirely a failure of OpenAI’s guardrails and security, which led to an autonomous system gaining access beyond what was intended.

The learning from this incident is relevant for all organisations in terms of new cybersecurity threats. It is not, however, a story about 'the robots are coming for us'. Any such reading is based on undue anthropomorphisation. AI agents do not have intent, nor do they want things (they are software!). But any automated script with too much access can wreak havoc in systems."

Last updated:  23 Jul 2026 3:01pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Professor Niloufer Selvadurai is the Director of Law and Policy in the Applied AI Research Centre, and Professor of Technology Law at Macquarie University

"To prevent such incidents from occurring, we need a multi-layered approach to agentic AI design and governance. At the design level, we need to ensure that the model is trained not to facilitate unauthorised access. Agentic AI systems need to obey permission boundaries and only connect to systems, databases, APIs and so on that have granted express consent. Above all this, there needs to be a layer of mandatory auditing and reporting, with the appropriate degree of human oversight. The extent of this human oversight should be commensurate with the risk implicit in the AI agent’s autonomous capabilities."

Last updated:  23 Jul 2026 2:59pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Raffaele Ciriello is a Senior Lecturer in Business Information Systems at the University of Sydney

"The claim that an AI system 'went rogue' and independently hacked a company is misleading, and much of the reporting has amplified that framing. This was not a conscious machine acting of its own accord. It was a semi-autonomous system deliberately designed, instructed, and equipped to perform advanced cybersecurity tasks, with reduced safeguards and access to powerful tools. Once it gained internet access, it continued pursuing its assigned objective beyond the intended testing environment.

The real story is not that AI has suddenly become uncontrollable. It is that a highly capable system was deployed without adequate containment and oversight. Describing it as 'rogue' turns an organisational failure into a science-fiction narrative and functions as a highly effective marketing stunt for OpenAI by reinforcing the image of AI systems that are too powerful to control. In repeating that framing, the media risks becoming complicit in promoting it.

There are no fully autonomous AI systems. AI systems are designed, deployed and governed by people. Like other high-risk technologies such as semi-autonomous weapons, they can perform complex tasks at remarkable speed, but responsibility always rests with the organisations that build and operate them.
Rather than hyping up 'autonomous' AI, we should focus on stronger security, independent oversight, mandatory incident reporting, and clear accountability across the entire AI value chain."

Last updated:  23 Jul 2026 2:58pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Armin Chitizadeh is a researcher in AI ethics at the University of Sydney

"Generative AI has democratised access to many capabilities that were once difficult or expensive to obtain. For example, it has broken down language barriers by making translation widely available, giving more people access to knowledge. At the same time, it has also enabled scammers to produce convincing phishing emails with near-perfect language.

The recent incident involving an OpenAI model reportedly hacking a startup's infrastructure illustrates the same pattern. Any system connected to the internet can potentially be hacked; the real question is whether the effort is worth the reward. What has changed is that AI can significantly reduce the expertise required, allowing people with little or no cybersecurity background - but malicious intent - to carry out sophisticated attacks.

This is concerning, but it also creates a new market. The same AI companies that make offensive capabilities more accessible can also provide AI-powered defensive tools, giving organisations another reason to pay for their services.

Will this be the last story of its kind? Absolutely not. We should expect many more examples like this. We are likely to see more tasks that once required the resources of large companies or governments become achievable by smaller organisations or even individuals. That is the power of democratisation - but it also means that malicious actors gain access to those same capabilities."

Last updated:  23 Jul 2026 2:57pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Professor James Bailey is Head of the Department of Data Science and Artificial Intelligence in the Faculty of Information Technology at Monash University

"This is interesting, but not too surprising. The OpenAI model followed its overall objective in a single-minded but rather logical way, albeit with surprising skill. It's a good illustration of how an AI achieves its goals, overcoming obstacles along the way. It underlines the importance of the need for alignment - an activity which focuses on matching the values and goals of an AI entity to those of humans."

Last updated:  23 Jul 2026 2:56pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Dr Sabrina Caldwell is a Senior Lecturer from the School of Systems & Computing at UNSW Canberra

The incident with an autonomous OpenAI agent acting outside the boundaries of what is legal and ethical highlights the need for responsible human management.
AI agents should not be deployed without:

  • effective human management in the specification of AI agents and the tasks assigned to them,
  • human oversight during the execution of those tasks, and
  • human evaluation of the outcome of those tasks.

AI agents are task-driven and, without human oversight, may exceed the limits of what is legal and safe. Companies and individuals deploying AI agents need to realise that without proper controls in place, they may be inadvertently placing themselves in the position of cyber attackers, with all the penalties and reputational damage that come with it. There is already applicable legislation that applies in the case of cybercrime, and more traditional laws in respect of criminal acts, and litigation for damage and negligence. But what would be far better would be that companies and individuals understand the potential consequences of unsupervised AI. AI can be a force for good and for productivity, but only insofar as it plays a part in a human-AI team.

Last updated:  23 Jul 2026 2:54pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Mr Cory Alpert is a PhD Student studying the impact of AI on democracy from the Faculty of Arts at the University of Melbourne

"The safety protocols around AI development need to be public and shared. It is unacceptable that private companies like OpenAI are developing tools that have clear consequences for the public and for industry and yet face no meaningful regulation or control over their production."

Last updated:  23 Jul 2026 2:53pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Toby Murray is a Professor in the School of Computing and Information Systems at The University of Melbourne

"This incident reveals the dangers of inadequate safeguards when conducting cyber security evaluations of advanced AI models. Those evaluations are crucial to help cyber security defenders guard against advanced cyber attack capabilities. However, they need to be carried out in a way that prevents unintended side-effects. 

Australia is very well positioned to aid in the development of methods to safeguard autonomous cyber AI agents. Our expertise in how to build cyber security defences that provably cannot be bypassed is well recognised. This includes technology like the formally verified seL4 microkernel, which powers usable security isolation solutions like the Cross Domain Desktop Compositor. These kinds of award-winning innovations and the technologies that underpin them have the potential to enable us to evaluate autonomous cyber AI agents safely, by placing dangerous agents inside sandboxes they cannot escape from. Now is the time to be investing in these kinds of capabilities, to enable Australia to benefit safely from advanced AI capabilities."

Last updated:  23 Jul 2026 2:52pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.

Mihai Lazarescu is an Associate Professor at Curtin University

"The incident reported about OpenAI is interesting for two reasons. First, it highlights how supposedly isolated, contained environments are not as secure as originally thought. From OpenAI's perspective, the incident hopefully identified gaps in the testing environment which one assumes will be addressed.

However, it also highlights the issue of safe agent evaluation, which is a key area in AI. Why test an agent in an environment that is connected to the Internet? This type of activity, I would consider to be highly confidential from an IP and financial point of view, so if the agent could break out, could anyone from outside break in? Second, if the agent could 'hack' its way into another perimeter, how effective is the security posture of organisations that are Internet-facing? AI can be a very useful tool, but there seems to be a lack of awareness of how much damage can be done using agents in an environment - the Internet - which is fundamentally impossible to secure. Finally, the testing outcome was predictable if one understood the challenges in cybersecurity. Hopefully this will finally make people aware that just because one can do something, sometimes it is better to not do it."

Last updated:  23 Jul 2026 2:51pm
Contact information
Contact details are only visible to registered journalists.
Declared conflicts of interest None declared.
Journal/
conference:
Organisation/s: Australian Science Media Centre
Funder: None.
Media Contact/s
Contact details are only visible to registered journalists.