← Back
shadow magazine

Who should be held accountable when an AI Agent (accidentally) acts maliciously?

It looks like public perception of how 'intelligent' current AI models are varies widely. Back in 2022, a Google employee already thought their AI model was sentient. Today in 2026, it seems like every other week there's a new article released about how AI companies "can't hold back their AI agents anymore"[2, 3, 4].

It makes perfect sense for the general public to start fearing AI. In the past, people feared companies would use AI to replace all kinds of jobs, and now they are even hacking government organisations.

Responsible AI

As a professional in the field of AI, I'd like to make one clear distinction. At the end of the last paragraph, what do you think the word "they" refers to? Reading popular headlines on the topic, it typically reads like AI agents are the ones doing the hacking and thus being the ones to blame. I'd argue that these headlines in part cause fear-mongering among the general public, as it's not the AI agents at fault for finding vulnerabilities and accessing digital infrastructure in unexpected ways. AI, and AI agents, are merely tools that companies and individuals run to reach some goal. Before using a tool, it is essential to deeply understand its limitations. That's also why the European Union released the EU AI Act including its mandated AI Literacy: organisations that deploy AI systems should sufficiently educate their users on it. In the physical world, users of a circular saw should carefully read its instructions before use, and even then, the engineers of the saw still add a blade cover and emergency stop just to mitigate risks as well as possible. Digitally, we need to similarly act responsibly on both the engineer's and user's side of AI, too.

Let's be clear: the fact that AI agents are breaking out of sandboxes and "hacking" public websites is very concerning. The AI models behind these agents have gotten incredibly good at generalization, to the point where their text generation seems like intelligence. But let's not forget that these AI models (Large Language Models; LLMs) are doing just that: generating text, effectively predicting the next word, over and over again. They are not deemed conscious like humans. They just show semantic understanding of text, in the sense that they can output text that logically follows the previous text. It's powerful, but not human-like conscious or independently harmful. These models are simply goal-oriented.

Who is to blame?

Let's get back to the question of accountability. Headlines talk about AI agents breaking out of sandboxes. The AI agents are merely tools used. These companies' researchers set up agents to complete a task, sometimes an impossible one in the case of the HuggingFace hack, and the agents (thus: tools) start processing everything necessary to reach the given goal. They do not have harmful intent per se. They do not have any intent other than solving the initial query, or prompt, that the researchers supplied. It is these researchers, who set up AI agents in sandboxes to contain them, who determined that the sandboxes are secure enough that they don't require continuous human-in-the-loop monitoring. Unfortunately, these sandboxes were rarely sufficiently secure.

And that right there shows where the accountability should be.

Any system with risks of causing major harm to other systems or people should have sufficient risk mitigations. Setting up a sandbox that should restrict public internet access to these agents, is merely one such mitigation. AI companies like OpenAI, Anthropic and many more should always account for the Swiss cheese model: it is not enough to assume one mitigation will patch all risks. Although I'd always recommend as much human-in-the-loop as possible, e.g. a human gatekeeper to approve potentially dangerous AI-suggested actions, I can understand persistent human gatekeeping would slow down AI innovations too much. Perhaps a better mitigation would be a human-on-the-loop: human supervision based on potentially dangerous consequences of actions. Heck, why not use a separate LLM or even Jev to automate classifying danger-levels of agents' actions before running them, to raise a flag and pause the agent until the human supervisor has approved the potentially dangerous action. There's little need to approve the fetching of website data, but they should implement automated flag-raising and temporarily halting the system when the AI agent's text output suggests e.g. hiding secrets in a web request. In fact, the most obvious mitigation would be to simply halt the system the moment it first attempts to access the public internet outside its expected scope, regardless of request content. The fact that this relatively easy-to-implement risk mitigation wasn't applied to sandboxes that AI agents "break out of" tells you a lot about the ethics of AI-use at said AI companies. Not only should the researchers have been more responsible, leadership should absolutely have understood the dangers of these experiments and pushed back as well without proper monitoring.

What should we do?

There are plenty of ways to mitigate risks that come with the use of AI agents. You don't have to be afraid of AI. You also shouldn't anthropomorphize AI. What you could be afraid of, however, is companies like OpenAI treating AI agents on the internet like the Wild West, and pretending their researchers, engineers and leadership aren't accountable for the AI-generated actions that they allow. These companies should be held accountable for insufficient risk mitigation, irresponsible use of AI and all of the harm this causes. And journalists, too, should really think twice about the phrasing of AI news. A headline that talks about how AI "has gotten too intelligent" or "couldn't be contained" might get more views than the objective truth, but clearly at the cost of readers' notion and understanding of AI. Please reconsider the ethical side of journalism, and the potential consequences of sensationalising this hard-to-grasp topic that is AI, for those who are less familiar with the subject matter.