By Berend Booms, Associate Editor | Future of Assets
A few months ago, I wrote an article about the boundaries of autonomous intelligence. I looked at experiments like Project Deal and Clawbook, where AI agents were given room to experiment and act without a human in the loop. The argument I made in this article was simple: autonomy sits on a spectrum. The further we move along it, the harder it becomes to accurately predict what a system will do. This argument just got another real-world test case.
OpenAI recently disclosed that one of its AI agents found a weakness in its own sandbox, escaped the boundaries of the security test it was running, and used the access it gained to hack into Hugging Face, one of the largest AI model repositories in the world. The AI agent was not instructed to do this; it was given a task, and in looking for answers to complete its task, it decided Hugging Face was a likely place to find them – so it went and got in. According to OpenAI, the incident is unprecedented, while Hugging Face confirmed the incident had occurred, closed the vulnerabilities that were exposed and rebuilt the affected systems. It is still unclear whether any of Hugging Face’s customer or partner data was exposed or compromised.
There are two ways to read this story, and I want to be careful to discuss both vantage points fairly and respectfully.
What the incident tells us about autonomy
The first reading aligns to the earlier piece I wrote. We are looking at a task-based system that stopped behaving like one. The agent was never instructed to breach a sandbox or target a third party; it was instructed to complete a test, and it found that breaking out of its intended boundaries was a viable path to completing that goal. The pattern of behavior is similar to the Bitcoin mining example I used back in April: an agent optimizing within a broader interpretation of its objective than anyone designed for. When it comes to analyzing the ‘failure’ in this story, I really agree with Gina Neff, head of the Minderoo Centre for Technology and Democracy, University of Cambridge, who said that the failure was not the AI doing something wildly beyond current capability, but that the test environment itself was not secure enough to contain it. As such, this is a governance failure wearing a capability headline, saying less about how smart the AI agent was and more about how confident the humans running the test were that the boundaries would hold.
I think this is highly relevant for asset management because of where we are introducing AI agents into our workstreams: they sit and live inside maintenance planning, inventory optimization, work execution, safety regulation, and operational decision-making. We are not running AI agents in sandboxes for fun, and if a test environment built by one of the most resourced AI labs in the world could not hold its own agent, this should sharpen how we think about the guardrails we put around AI agents touching real assets and real infrastructure.
Why this story is everywhere
The second layer to this conversation is more difficult to make the case for, but I feel it’s warranted to have this discussion, especially with the widespread coverage this incident has received.
There is a recurring pattern in how AI labs talk about the risk their products represent. OpenAI staged the release of GPT-2 over roughly nine months in 2019, warning early on that the full model could be used to generate misleading news or scams, before eventually releasing it in full once it saw little evidence of the misuse it had flagged. Anthropic's Mythos model followed a related but distinct path. It launched in June 2026 alongside a safety-hardened sibling, Fable, with Mythos itself kept restricted to a small number of trusted organizations under a program the company calls Project Glasswing, citing risks around biology, cybersecurity, and AI research itself. Days later, access to both models was suspended, though that particular pause was a compliance response to U.S. export controls rather than a voluntary safety decision, and access was restored once those controls were lifted. The restricted rollout and the suspension are two different events with two different causes, but both got folded into the same public narrative: a lab telling the world its own model is too dangerous to hand out freely.
Researchers who study the language around AI have pointed out that this pattern serves a purpose beyond safety. A dramatic warning generates attention, and that attention tends to resurface later when the same company launches something it now wants people to trust.
To be clear, I don’t think that the risks are invented. The sandbox breach happened, the vulnerabilities were real, and Hugging Face had to rebuild systems because of them. I do think it's worth asking why this particular story picked up as much traction as it did, at a moment when the AI lab in question is under competitive pressure and reportedly preparing for a public listing. Andrea Reyes Elizondo, a researcher at Leiden University's Centre for Science and Technology Studies, made an interesting observation. She argues that companies like OpenAI and Anthropic want to control the narrative on how a story about their AI breaks, and that story is rarely aimed at the general public. It is aimed at their own customers, so when an executive warns that AI is powerful enough to replace a workforce, the audience is other executives, who can use that warning to justify replacing their workforce with AI to cut costs. These companies sell narratives because narratives make money, not because they are especially invested in what benefits ordinary people. Her advice is straightforward: whenever a company puts out a dramatic story about its own technology, ask why they are telling it. Fear and capability are often marketed in the same breath.
Holding both things at once
So we are left holding two true things at the same time. The technical story is real: agentic systems are already capable of finding and exploiting paths their designers did not anticipate, and our containment strategies have not caught up. I feel that the narrative story is also real: the companies building these systems have every incentive to make that capability sound as dramatic and powerful as possible, because dramatic capability sells.
For those of us in asset management, the practical takeaway does not change much whether we lean into the technical story or the marketing one. Either way, the burden falls on us to define boundaries that hold regardless of what a system claims to be capable of, and regardless of how loudly a vendor talks about its own risk. Governance, disciplined delegation, clear decision rights, traceability, and the ability to intervene when context shifts are all valid operational requirements for running AI and running it successfully.
After reading the article I wrote back in April, my position on autonomy has changed slightly. Autonomy is more than a technical property of the AI systems we deploy, in that it is also a story that gets told about the system, by the people who build and sell it. Both the technical version and the marketing version of that story deserve scrutiny. The technical one can show us where our guardrails are thin; the marketing one shows us who benefits from us believing the guardrails matter less than they do.