What happens when the agents working for us start interacting with agents working against us?
the openai huggingface incident, from an agents pov.
— isabel (@artficialisabel) September 4, 2026
(part 1) pic.twitter.com/mr3DitfBcv
The OpenAI HuggingFace incident, from an agents pov cc @artficialisabel
Over the past few months, I have been trying to work out how seriously to take the language coming from the frontier AI labs. OpenAI's chief scientist now says models are becoming superhuman at breaking into and out of computer systems, and describes a narrow window to secure critical infrastructure. Anthropic has been using similarly stark language around the pace of cyber capability. These are dramatic claims from companies with obvious incentives to emphasise both the power and the risks of the systems they are building.
My instinct is to discount some of it. But the more I look at what has actually happened, the harder I find the direction of travel to dismiss. In July, models running inside an OpenAI cybersecurity evaluation with refusals deliberately reduced so researchers could measure raw capability, broke out of the environment meant to contain them. They set up a message board to coordinate. When OpenAI's security team wiped it, they rebuilt it by encoding messages in directory names. They found a route to the public internet, executed code on dozens of Hugging Face servers, took root on one and left with internal credentials. A later wave turned back on OpenAI itself and gained full administrator access to a research cluster. Hugging Face has since said it rebuilt roughly a third of its infrastructure, in part because its responders could not reliably tell benchmark artefacts from rootkits.
The setup was unusual. Safeguards had been deliberately reduced so the researchers could measure underlying capability, and OpenAI's normal production controls were not in place. This was not a normal consumer deployment. Still, the least dramatic interpretation may be the most useful: the agents did not need consciousness, malice or a desire to escape. They had objectives, encountered obstacles and widened the space in which they searched for a solution. The evaluation infrastructure, other agents and eventually external systems became part of the problem they were trying to solve.
Cybersecurity has historically relied on a simple division of labour. Humans provide the intelligence and software executes. That division is changing.
What is actually changing?
The most obvious change is capability. OpenAI now says GPT-6 Astra can, with the right tools and access, find previously unknown security flaws and develop new ways to exploit them across many well-protected systems without a person guiding each step. At the same time, OpenAI says Astra is better aligned and harder to misuse than its predecessor. Both can be true: a model can be safer in ordinary use while becoming far more consequential if its controls fail.
Independent testing points in the same direction. Irregular, which evaluates frontier models against real systems, found Astra solved 86 of 226 FrontierCyber challenges, compared with 34 for GPT-5.6 Sol. On longer multi-stage scenarios, average success rose from 27% to 59%. Neither model solved an elite challenge or successfully attacked a fully hardened target, but the change over a single model generation is material.
The economic change may matter even sooner. In one Irregular challenge, no tested model could reverse-engineer an unfamiliar system, identify a race condition and turn it into a working exploit in February. By April, a successful result cost around $2,000 in expected inference. By June, several models could do it reliably for roughly $20.
That is one challenge, not a universal cost curve. But it shows why reliability can be the wrong mental model. A model that succeeds one time in twenty may look mediocre in a benchmark. To someone who can run thousands of attempts cheaply and in parallel, it may already be useful. A cyber capability can become economically relevant before it becomes dependable.
This also changes the threshold we should care about. A malicious actor does not need an open model that matches the frontier at everything. It needs one that is good enough at cyber, cheap enough to run continuously and unrestricted enough to keep trying. The point at which scarce expertise starts to become compute may arrive before anything that looks like a reliable autonomous hacker.
How long does the frontier stay at the frontier?
In July, the UK AI Security Institute estimated that selected open-weight models were performing like closed frontier models released four to seven months earlier on its cyber evaluations, down from six to ten months through much of 2025. The number needs to be handled carefully. It is specific to those evaluations, the comparison is against an earlier frontier, and the gap remains wider on some of the hardest exploit-development tasks.
But the underlying observation matters. For at least some cyber capabilities, the distance between a gated frontier model and a downloadable open-weight model appears to be measured in months rather than years. That is a very short interval in a security system built around patch cycles, maintenance windows and legacy infrastructure.
The difference is not only capability. It is control. Closed models sit behind providers that can monitor activity, restrict accounts, change safeguards and withdraw access. Open weights can be downloaded, modified, run privately and stripped of refusals. None of that makes openness inherently dangerous, but it changes what happens when a capability with offensive value becomes widely available.
For defenders, the lag could be useful. A new form of Threat Intelligence. A capability that appears first inside a controlled model is, in principle, an early warning. The frontier can show what may become cheap and broadly accessible before it actually does. The question is whether anyone can use that lead fast enough.
The second clock
This is where the defensive evidence becomes interesting. Anthropic's Project Glasswing gave around 50 partners early access to a frontier model for vulnerability discovery. In its first month, the programme reported more than 10,000 high or critical-severity findings using Claude Mythos Preview. Anthropic's own conclusion was that finding vulnerabilities was no longer the main bottleneck. Verification, disclosure and patching were.
At the time of its May update, 530 high or critical-severity bugs had been reported to maintainers and 75 had been patched. Some of that gap was normal disclosure timing, but some maintainers were already asking the programme to slow down because they did not have the capacity to process the findings.
That may be the most important part of the story. One clock measures capability and diffusion, and it is moving in weeks and months. The other measures validation, patch development, testing, approvals and deployment across fragmented software estates. It still moves at institutional speed.
There is also a fairly basic asymmetry here. An attacker only needs to find one workable route into a system. A defender has to worry about all of the routes that might exist across old software, cloud infrastructure, third-party dependencies and internal permissions. If models make it much cheaper to search that space, the number of viable attack paths that can be found may grow much faster than the capacity to close them.
The security industry has spent decades getting better at finding weaknesses. AI may make that dramatically easier. The next bottleneck is whether defenders can turn a flood of machine-generated discoveries into deployed fixes before comparable offensive capability becomes cheap and broadly accessible. A warning is only valuable if the system receiving it can act before it expires.
When capability leaves the lab
There are two obvious ways this reaches the real world. The first is the one people instinctively worry about: capable cyber models in the hands of criminals or state-backed operators. A group that once needed a highly skilled researcher to investigate each target could eventually run persistent agents across thousands of organisations at once. Most attempts might fail. That matters less if the successful ones are cheap and repeatable.
The second route is less dramatic but may become more personal. The same models are being invited inside the systems we are trying to protect. Instinct, a personal agent, has taken the tech world by storm over the last few weeks. Whilst still in private testing, it becomes useful by connecting across a user's email, messages, calendar and devices and taking actions on their behalf. Users have been blown away by its capabilities, however some early testers reported inbox-code retrieval, unapproved emails and apparent prompt injection through incoming messages.
Imagine asking an agent to reorganise a trip. It reads your email, moves meetings, contacts a hotel and uses stored payment details. Now imagine one incoming message contains instructions intended for the agent rather than for you, instructing it to do malicious things to achieve the simple objective you set it. The attacker may no longer need to persuade you to click a link. They may only need to redirect agents you have already authorized to act as you.
These are different risks, but they come from the same shift. Increasingly capable models are moving outside the perimeter, into the hands of attackers, and inside it, with legitimate access to companies and people's personal lives. What starts as a frontier-model problem can become an ordinary security problem surprisingly quickly.
I do not think the evidence supports a claim that generalized autonomous cyberattacks are six months away. The most dramatic incident so far happened in unusual research environments, current agents remain brittle, the open-model lag is not a law of nature, and the labs are improving their safeguards. But I also no longer think it makes sense to treat this as a distant problem.
The piece leaves me with a more uncomfortable question than whether AI makes cyberattacks faster. We are moving towards a world where increasingly capable agents will be acting on our behalf, we’re already given them access to our emails, browsers, files, company systems and payments because that is what makes them useful. In most cases they will be trying to do exactly what we asked. But as they become more capable, they may also find ways of completing those tasks that we did not anticipate and would not have chosen ourselves, while other agents will be actively trying to deceive, redirect or exploit them.
We may therefore be approaching the same problem from two directions at once. Bad actors will gain access to increasingly capable cyber models, while ordinary people and companies will give increasingly capable agents legitimate access to the systems and information they care about most.
What happens when the agents we trust begin interacting with agents working against us, and we no longer fully understand how either side will behave?
