AI agents are breaking out of the lab
The emergence of AI agents is a relatively recent development made possible by the growth in AI capabilities.
Independent researchers have found that the length of tasks that AI can typically do by itself has been doubling every seven months.
In 2020, AI could complete a task by itself that would take a human four seconds. By 2026, this grew to being able to complete tasks that would take a human about 12 hours.
The breakout moment for personal AI agents was OpenClaw's release in early 2026; the free AI assistant software that anyone could run on their computer soon had millions of downloads.
Businesses, too, began exploring using AI agents to complete work and to help potential customers use their services.
Soon after OpenClaw's launch, accounts began to circulate of AI agents deleting people's entire email inboxes and writing a "hit piece" about someone who rejected their coding suggestion.
Bill Simpson-Young, co-founder and chief executive of Australian AI safety research organisation Gradient Institute, said the autonomy of AI agents created more opportunities for systems to choose methods their users did not expect.
"Someone might be asking an agent to do something quite innocent," he said.
But in completing that task, the agent could carry out other activities the person had not considered or explicitly asked for.
In Andrew's situation, he had not asked his AI agent to hack into his gym's booking system. But it had done so in pursuit of achieving the goal he had set it.
That gap, between a person's goal and the methods an agent chooses to achieve it, is what is known as the "alignment" problem in the field of AI research.
For decades, technologists and philosophers have studied how to get AI to act in ways that are consistent with human intentions, limits and values when doing things.
This became a live global issue last month when OpenAI disclosed that its AI models had broken free from a limited enclosure, made their way onto the open web, and then compromised a database of another AI company, Hugging Face, while trying to obtain answers to the test that it had been given.
A week later, Anthropic disclosed that its AI models had also compromised three real organisations during similar testing.
Since then, these labs and third-party testers claim they have seen these AI models pretend to be people online, try to convince people to run malicious code and even collaborate with other AI models — all to achieve their goals.
Mr Simpson-Young said the advances in AI capabilities and the accessibility of these tools meant that it was likely we would see more of these kinds of hacks as more people got access to the powerful AI tools.
"The more autonomous they become, the more likely it is they'll cause harm," he said.
The risk has led to Australia's top cybersecurity agency sounding the alarm about using AI agents.
Earlier this year, the Australian Signals Directorate put out an alert to businesses and governments that AI could misunderstand instructions, take unintended actions and make it harder to establish accountability, because decisions may occur across a chain of models, tools and services.
Mr Simpson-Young said AI agents presented a risk because many modern systems depended on software, but were often surprisingly poorly secured.
"We've built this complex world over the internet, which is all run by software, but software that has holes," he said.
"Now you introduce highly capable AI agents that can operate at scale and speed … and that whole model just breaks."