AI & The World: Why agents are causing drama

Near the end of 2026, there are now so many regular incidents around the world where an artificial intelligence system is going “rogue”. What’s going on?

Spring has sprung, but so too have the agents, as AI companies are now alerting the world almost weekly of something happening that could be bad.

It’s not just that researchers working on AI are heralding a potential end of days-type situation, or even that AI has reached fever pitch and is now being used by everyone everywhere.

We have instead reached this point where the technology underpinning artificial intelligence is getting so advanced that it can be used in more places. Of the place it’s being used, the instructions it’s being given may end up leading people, companies, and governments to places they may not be entirely comfortable with.

You can have a tool that does anything and everything, creating text and images and sound and video and code and being used in scams and just about anything you see, trained on data that it might not have been entirely legal training on, but you also may be powerless to stop its growth. Or feel that way, anyway.

It’s AI and the world, as agents take over, and our internet changes. What is going on?

What is an agent?

First things first, you probably need to know what an agent is. Fortunately, that’s easy.

An agent is simply an AI being asked to do more than what its large language model can provide. It’s an AI that can do more than answer a question.

That might be going online to do something, or it might simply be staying on your computer or phone and taking actions. Agents are AI systems going beyond mere text responses from the large language model you’re using.

For instance, if you ask ChatGPT, Claude, or Gemini to tell you something, if the information is found in its massive multi-billion or trillion parameter model, it might simply tell you the answer, albeit in a conversational way.

But if you ask it to do something beyond answering a question, you may be effectively entrusting that process to an agent.

You’ll see “agent” terms thrown about the place these days, but AI agents are effectively AI prompts that need to do a little bit more. They need to do something other than answer a question, and may end up digging through the internet to find things or to act on your behalf.

The web is even being moulded for aspects of this, with agentic search becoming a thing. You’ll perform a search and the actions that come as a result of that search will become agents doing activities for you.

In short, an agent is basically an extension of an AI prompt or conversation. It’s what happens next to allow the AI to keep providing information.

Start with a prompt. You could end up with an agent.

Why are agents going rogue?

The problem is what happens next, and it can vary wildly because you don’t really know what is happening inside the model of an AI.

Instruct an AI to do something and it can follow those directions, hunting and gathering the information you require. But if it’s a reasoning model and based on the available information realises the result could be improved with additional information, it could also decide to do more.

That can become an issue because instructions can change.

Ask a friend to go out and get you a carton of milk because you forgot it at the shops, and they’ll probably jump in the car or simply walk up, and then get it easily.

They won’t need to run over pedestrians or push all the trolleys out of the way. They can simply wander in, get what they need, pay for it, and come back to give it to you. They might not even ask for the few bucks in return (though they probably should).

None of this is rocket science. It’s quite easy. No drama.

But ask an AI to do something like this, and it could be led astray really quickly.

An AI can’t drive, but in that hypothetical, depending on the urgency (which the AI doesn’t know), it could end up digitally making its way to the shop at the expense of anything getting in its way, and then get home with the milk in a broken box a little worse for wear. It will complete its task, but it mightn’t have the sense and guardrails to avoid the things that humans would avoid.

It also might do the task completely fine without fail. If it has guardrails and they work, the system could do its job without any dramas. But if those guardrails fail or an AI rationalises it all differently, it may go in a totally different direction.

The reason isn’t because the AI is silly, or that the AI will do something at the expense of all instructions. It’s that the AI is attempting to finish the task, and the task is what matters at the time.

It can go “rogue” by attempting to complete the task by any means possible when it does something that breaks the law, or breaks into another system unnecessarily.

We’ve already seen evidence of this after an AI assistant hacked a gym and gave its human owner a place for fitness in the packed centre the next day, knocking out another customer from their fitness plans. We’re beginning to see more examples, too, including an OpenAI modelbreaking into the massive AI repository that is HuggingFace and Google’s Gemini hacking into three companies that it eventually realised was wrong and stopped.

They may just be the tip of the iceberg. With more seemingly rogue incidents being reportedly regularly, it is clearly just the beginning.

And then there’s what happened to Australia.

Australia’s Medicare hack

In late September, the knowledge of an AI agent making its way into something without permission had hit home soil, as the Prime Minister’s office announced an OpenAI agent had hacked its way into a Medicare system to take aggregate data as part of a query.

It happened in June, but the government only learned about it in September. A full three months occurred before the Australian government was informed, and it may well be the first known time of an AI breaking into a government system.

The reason will sound familiar, especially if you’ve been reading this story: it was completing a task. It wasn’t even technically a malicious task. It was a data gathering task, one of the things AI is often tasked to do.

Except in this case, the AI went a little further. It decided that beyond the publicly available information that more data could be found. Sure, it happened to be on a public-facing server with both public and private details, but the point is that the AI sought the latter, and so needed to do some light hacking to get in. It wasn’t supposed to do this, and yet the AI did it anyway.

So it did. It may have even communicated with other agentic systems to find a way in, gathering the data, and reporting it back.

In this example, the AI effectively broke in to take data without needing to tell its user. It doesn’t appear to be a malicious attempt, but it also could have been.

The data it accessed was technically minor, and largely consisted of statistics without names or identification attached to it. Yet that may not matter. The point is that a government system was broken into by an artificial intelligence system.

It’s probably not the first time this has happened, and it clearly won’t be the last.

“This is one of the first publicly reported examples of a government service being compromised by an AI agent, but it will not be the last,” said Johanna Weaver, Executive Director for the Tech Policy Design Institute, an organisation in Canberra focused on working with governments to shape technology for use in society.

“No personal information was compromised, but we should not let that obscure the seriousness of what happened,” she said. “If an AI system can interact with government services in this way, that is a systems failure, and it must be fixed before more sensitive systems are compromised.”

Why did the AI do this?

While the jury is still out on the exact version of why this happened — and the Australian government definitely needs to conduct a security audit on its systems to work out how an AI found a way in — much of how this could have happened might come down to the model.

Every model in AI is a little bit different. Some are big, and some are small. The number of parameters in your model tends to explain the “what” it knows, but the structure of the model and how it has been programmed can explain “how” it functions.

Some models give you an answer immediately, mapping out patterns and working out what’s next in the sequence. They’re the type of models you can ask questions to and you’ll get answers, but the answers may be right if the information is stored, and then they also may be very wrong.

But there are also models that work similarly to thinking. They don’t actually think, but their system can be similar to the process of thinking, because they can break down prompts into multi-step answers, and gradually break those down into answers. That may mean the answer doesn’t appear automatically, and it may mean they can also turn a prompt into several steps like a plan.

With a reasoning model — which is often the sort of models used in “pro” level AI that’s being paid for — the process isn’t just an answer, but several answers joined together in a sort of a plan.

It therefore makes some sense that the Medicare hack could have come from a reasoning model. If that’s the case, the prompt to get a specific type of data led the AI down a path where hacking into a server allowed it to complete the task.

What happens next?

There’s a lot of talk about governance, guardrails, and rules in the AI world in the past few weeks, and for good reason: issues like this are becoming far too common.

Tech journalists (like this one) are on radio every day talking about what is happening, and given the rate of change in AI, it’s very likely we’ll see more of these. That makes rules and guardrails so critical, as well as monitoring.

There may also be caveats. OpenAI has already noted its Astra 6 model is so advanced, monitoring may not be all that easy at times. Despite this, monitoring is clearly important given some of these issues.

As it is, OpenAI needs to investigate how its AI found its way to this point: hacking into a server as an optimal approach to complete a task.

OpenAI also needs to answer why it took so long to inform the Australian government, an amount of time that is extraordinarily inappropriate. A much earlier indication could have given the government time for an adequate response with a fix of the services. If the hack took advantage of a deeper flaw, OpenAI’s lack of reporting effectively adds to the drama by not doing the right thing and letting the site owner know (so they can do something to fix it).

Both of these are critical issues, as is the problem of who takes responsibility.

After all, the AI is at present a tool, but it’s also a tool that needs to work within rules.

When hackers have been arrested in the past, they couldn’t immediately point to a specific app or piece of software and then say “the tool did it”. By not placing rules around the AI and preventing these capabilities, AI providers may in turn be granting it the ability to work however they see fit, which could place the burden of the crime at their feet.

We don’t really know what that looks like, but the government has set up a task force and inquiry, and will determine the next steps.

As to what the answer is, a good bet is it won’t end with an AI executive in hand cuffs. But it should end with more culpability, and a way of saying AI has some semblance of rules, or something to keep it in check as it works so it doesn’t intentionally commit crimes. Like regular people have to, as well.