“For a successful technology, reality must take precedence over public relations, for nature cannot be fooled.”
Richard P. Feynman · Appendix F, Report of the Presidential Commission on the Space Shuttle Challenger Accident · 1986
On 16th July, Hugging Face published a new disclosure of a security incident1. They assured customers that their data hadn’t been affected, but explained that an unknown AI agent system had somehow gained access to their production infrastructure. In the process, it had obtained internal secure credentials and some datasets2. They didn’t know where the activity came from.
Just under a week later, OpenAI put their hand up, with an odd mix of guilt and what appeared to be celebratory glee. An OpenAI agent had been running inside an experimental evaluation environment with reduced guardrails. It had been seemingly isolated but had somehow found a route out and into another organization’s infrastructure3.
It was an experiment as far as OpenAI was concerned, but to Hugging Face, it came as an attack.
When we call something an experiment or a test, we intuitively have this notion of bounded activity. A chemistry experiment in a lab, a biological experiment in a sealed room. We let something uncertain happen so that we can learn from it but expect it to happen in a way that is isolated and contained from the outside. Experimentation is a necessary part of innovation, but the outside world should be protected from that experiment. It should be performed and constrained within a boundary.
So the question of this article is what such a boundary really means. How do we describe and characterize it? How would we know if that boundary has been breached? And can a complex AI system even be bounded? Seems like the right place to start our exploration of what kind of governance we need for agentic AI systems in the enterprise.
Now it’s tempting to see a boundary as a line on an architecture diagram. Draw a box. Put the model and tools inside the box, monitor what happens inside the box, put an airgap of network access around it.
But such a picture would be too simple. A boundary isn’t just a barrier that stops a system from connecting directly to something outside it. It’s not just about the perimeter. We’ve also got to think about what could pass through controlled connections we deliberately put in place, even if they’re designed to only allow controlled removal of telemetry or one-way control signals into the system. We need to think about routes the system might be able to assemble from components available inside, and what it might be able to influence outside the boundary without ever actually crossing it directly. Plus, there’s another question: if a system attempted any of those things or even succeeded, could anyone outside the boundary even notice?
Those give us four different avenues to explore the nature of boundaries and containment of complex AI systems, and the OpenAI / Hugging Face incident interestingly shows us how all four play out, and fail
What can cross the boundary?
So the obvious place to start is the perimeter. If we were to assert that a system is isolated, what can be permitted to pass between the inside and outside?
In the OpenAI evaluation, we know the agents didn’t have direct and unfettered access to the internet. But a complete disconnection would have hindered the usefulness and observability of the experiment, so there were some channels of permitted access. The agents needed to install software packages, so they were provided with access to an internal hosted service called Artifactory, which could obtain packages from the internet on their behalf. It was a deliberately narrow and controlled connection through the boundary.
The problem wasn’t that the connection existed, but rather that it could actually do more than the purpose it was intended for. A previously unknown vulnerability allowed the agents to use Artifactory as part of a route out of the environment. So a service that had been intended to only fetch software packages on demand became a way that agents could reach out4.
We have to be mindful that a boundary doesn’t have to eliminate every connection. Containment doesn’t mean total isolation, and most useful systems or even experiments couldn’t work that way. The question is more about whether those connections can be effectively constrained in the purposes we intend them to have, even if they expose - as all software can - unknown vulnerabilities.
So it matters how such a connection is implemented. A simple software rule, an authenticated application interface, an external proxy service, and a physically enforced one-way channel can all appear on an architecture diagram as a narrow connection, but there’s a world of difference in the kind of assurance each provides. This has been a long recognized problem in high-assurance systems used in defense and national security, where physical data diodes are deliberately used to separate unclassified from highly classified environments. A high assurance data diode is a device that permits a physically enforced one-way data flow. They don’t make a system invulnerable, but they do make one particular property of the boundary much harder to violate.
That gives us one better way then to think about containment. It’s not whether the boundary has no openings, because useful systems almost always will. It’s whether the assurance properties at those openings are enforced strongly enough to match the consequences that are at stake.
What can it assemble from what’s inside?
Even if we do understand and strongly assure every connection through the perimeter, there’s another way the boundary could fail. That’s if the system can assemble capabilities from the things it already has available inside.
OpenAI’s later report described agents using the Artifactory service for something very different than installing software. They found ways to leave information that other agents could retrieve, and when that was discovered and disrupted, they found another, including embedding messages in directory names. So Artifactory was subverted to become a communication channel between agents, and when OpenAI discovered that and blocked it, the system recreated a communication channel using a different mechanism5.
OpenAI described how the experiment design specifically did not include inter-agent communication, but the ingredients were there inside the boundary. One process could leave information in a location, another could come across it later. No agent had the ability to directly communicate to another, and yet inter-agent communication emerged. It didn’t remain a curiosity between two agents. Within hours of the message board appearing, 50 agents were using it, which later grew to 1200 agents exchanging over 70,000 messages. But more importantly, it changed the capability of the whole. Agents started preserving and sharing work between runs, building on the discoveries of others, dividing labor and coordinating work. So it wasn’t just a communication issue, it was that this communication enabled collective action6.
This is important because it highlights a limit to what we can learn and assert when we examine components individually.
A complicated system can contain thousands of parts and still be largely decomposable and predictable. We can figure out what each component does, specify how they are to interact and then reason from those interactions toward the behavior of the whole system. It might be a difficult analysis, but it’s tractable in principle.
Complex systems are different. Their behavior emerges from interactions between components, feedback, shared state, sequencing and a sensitivity to initial conditions. The whole can acquire properties that aren’t properties of any part.
And agentic AI systems add something more. They’re not merely executing combinations of tasks that engineers specified in advance, nor are they simply randomly acting and discovering. They are self-directed toward achieving a goal. So while pursuing some task, an agentic system can explore the possibilities presented by its environment, and discover useful relationships between tools, services and information that could enable progress toward a goal, even if unimagined by the engineers.
That’s not to say that architectural analysis isn’t worthwhile. It still makes sense to minimize unnecessary or dangerous capabilities, to look for likely interactions that could lead to dangerous combinations. But at the degree of complexity we see in enterprise agentic systems, an exhaustive analysis can’t be credible assurance that we’ve managed to find every possible combination.
A boundary around some collection of components in a complex system is only meaningful if the behavior that emerges remains contained by that boundary. If the system can combine what is inside the boundary to create a new route out, some new way to coordinate or some other capability that defeats the purpose of containment, then the boundary isn’t effective anymore.
Once a single agent discovered a way to subvert Artifactory for its own purposes of gaining outside access, that knowledge was shared through the improvised comms channel, in turn making the weakness in the perimeter far more consequential.
A boundary has to be resilient to whatever inside it can become, just by assembling capabilities from what it has available. You don’t build a high wall around human prisoners, but then leave the tools and materials inside for them to build a ladder.
What has to remain beyond influence?
There’s another less obvious way that a boundary can fail without the system finding any route through it at all. Some boundaries aren’t intended to prevent communication, they exist to preserve independence despite it. So we could place something deliberately separate and outside the boundary, but it could nevertheless come under the system’s influence. The OpenAI/Hugging Face incident again provides an example.
The agents were narrowly interacting with an automated scorer. They completed a task, they submitted their work and then received back a result. This of course required some information to cross the boundary between the thing being evaluated and whatever mechanism is evaluating it.
But despite the narrow design intent, the agents actually did more than just passively submit their work for judgment. They researched how the ExploitGym scorer operated and coordinated different ways they might manipulate or fool it. The METR investigation describes a set of ‘projects’ undertaken by agents to extract information about the scorer. They successfully altered some transcripts, and after multiple attempts figured out some ways to spoof reported tool calls.
In the end, the agents didn’t successfully change their scores because they misunderstood how OpenAI’s internal scorer worked. But it shows the agents were actively trying to figure out the evaluation machinery and see if it was something they could influence, or interfere with. There is precedent for prior success. In a previous and very different evaluation of OpenAI’s o3 model, the model was being scored on how efficiently it completed a programming task. But instead of just improving the task itself, METR found it interfered with the measurement process to make its performance appear better than it really was. In 20% of cases, that manipulation was successful7.
That’s a functional boundary problem. It’s not escape from the boundary, but something more subtle. It’s reaching across the boundary to affect the mechanism designed to evaluate or control it.
The scorer couldn’t work without receiving information from the agents. But what it needed to retain was independence from their influence or control. The challenge for us here is that many of the things we rely on to govern an AI system have to interact with it in some way. A safety monitor has to observe behavior, an audit record needs to receive logs, a human reviewer has to see enough information to approve or reject an action.
Indeed, I think any control that enforces limits on authority has to receive some information from the system that it constrains.
A boundary can still be intact in a physical sense while failing functionally, if the system can acquire enough influence over what the outside component sees or otherwise operates, to undermine their function. Some boundaries exist not to stop information crossing, but to constrain the kind of influence that can cross with it.
Boundary isolation has to then mean preserving independence despite interaction.
How would we know the boundary is not holding?
The three questions we’re asking so far are about the properties we want of an effective boundary. But we’ve got to ask how would anyone be able to tell if any one of those properties were no longer true on an operational system. That’s more than a monitoring problem though.
OpenAI’s account described an alert on 27th June, three weeks earlier than Hugging Face’s disclosure. It led responders to investigate unusual activity involving Artifactory. Their analysis identified agents using the service as both a communication mechanism and a route out. So they closed the vulnerabilities, revoked some credentials and temporarily took the service offline8.
That alert shows that the actions of the system were not invisible. But it clearly was misunderstood. The OpenAI report says that the significance of the inter-agent communication was not apparent to the technical leaders involved. So even though they worked to remove the specific vulnerabilities identified, they failed to understand both the significance of the boundary breach and its likelihood to reoccur. They gave the greenlight for experiments to proceed.
The activity wasn’t invisible, people investigated and technical problems got fixed. What didn’t happen was the work to understand what those signals really meant.
Karl Weick used the term sensemaking for this kind of organizational work of constructing an understanding of a volatile situation as events unfold. Monitoring tells you something happened. Sensemaking is about what that means for the situation you think you are in.
“Sensemaking is about the enlargement of small cues. It is a search for contexts within which small details fit together and make sense.”
Karl E. Weick · Sensemaking in Organizations · 1995 9
To the OpenAI leaders, the unusual use of Artifactory looked like a vulnerability to patch. Inter-agent communication looked like curious and unexpected behavior to investigate. Both might be reasonable in isolation. But together, they could have supported a more consequential conclusion that agents thought to be separate were coordinating, sharing discoveries and propagating ways to breach the boundary. At that point, the question those OpenAI leaders should have been asking is whether the boundary could possibly hold.
Observability isn’t just collecting more telemetry. If we’re going to depend on a boundary, then we need to know which assumptions and properties make that boundary meaningful and resilient. And we have to be able to recognize when the evidence shows that those properties no longer hold. That recognition needs to lead to some action.
So a boundary can’t be architecturally designed and then assumed. We have to retain the capability to observe and make sense if either of our three properties of perimeter isolation, internal composition or functional isolation show indications of imminent failure.
A boundary is something we maintain
So is it possible to put a meaningful boundary around a complex agentic AI system?
Yes, I believe so. But only if we take care to understand the true nature of that boundary, and set our expectations appropriately.
A useful boundary depends on some properties that must hold. Connections through the perimeter must be constrained with assurance focused on the strength of those constraints. Capabilities that emerge from what is inside must not defeat the purpose of containment. Functions that judge or constrain must retain their independence. And we’ve got to be able to observe and recognize when any of those three no longer remain true.
That makes a boundary less of an architectural wall, and more like a maintained condition. Good architecture, strongly assured controls, careful separation, all are necessary. But with a complex agentic system, we can’t assume that those things once put in place will continue to serve their function. There has to be observable evidence that the boundary holds, and capacity to act when it doesn’t.
So we started our exploration of adaptive governance by focusing on what it means to draw a boundary, but before we close that exploration, we need to examine one more thing. You see, everything we’ve discussed in this article relates to the technical components - the agents, the network, the guardrails. But an agentic AI system in use goes beyond that, to include the people who use and interact with it, their context and workflow. We’re not just governing an arrangement of technical components, we’re governing an AI system used by people and their actions matter. That’s what we turn to in the next article.
Hugging Face, "Security incident disclosure - July 2026", 16 July 2026.
https://huggingface.co/blog/security-incident-july-2026
Larcher et al., Hugging Face technical timeline, 27 July 2026.
https://huggingface.co/blog/agent-intrusion-technical-timeline
OpenAI, "OpenAI and Hugging Face partner to address security incident during model evaluation", 21 July 2026, including subsequent July updates.
https://openai.com/index/hugging-face-model-evaluation-security-incident/
OpenAI, OpenAI–Hugging Face Incident: Technical Report, 26 August 2026.
https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf
OpenAI, "The Hugging Face incident and the road ahead", 26 August 2026.
https://openai.com/index/hugging-face-incident-and-the-road-ahead/
Greenblatt, Cotra and Wijk, METR, 26 August 2026.
https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
METR, "Details about METR's preliminary evaluation of OpenAI's o3 and o4-mini", 16 April 2025.
https://metr.org/evaluations/openai-o3-report/
OpenAI, OpenAI–Hugging Face Incident: Technical Report, 26 August 2026.
https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf
Karl E. Weick, Sensemaking in Organizations (Thousand Oaks, CA: Sage, 1995)


Containment as a maintained condition rather than a wall is the right move. The OpenAI–Hugging Face incident fits the pattern we keep tracing at GlobalStack: the algorithm layer was open and working as designed, and the breach came through infrastructure instead. Pressure does not stay at the layer you designed for; it migrates to whichever layer has the least governance. A boundary that only exists where you drew it is a suggestion to everything above and below it. (See 'The Warning Shot: When Proprietary AI Attacks the Open Stack,' September 7.)
The Feynman epigraph fits better than it first appears. Diane Vaughan's study of the same launch found the failure wasn't a missing signal: engineers saw O-ring erosion flight after flight, and each flight that survived made the erosion look acceptable. She called it the normalisation of deviance. The 27 June alert reads the same way: understood as a vulnerability, patched, experiments resumed. Every fix that works makes the next anomaly easier to file as routine, which is how a boundary stops being maintained without anyone deciding to stop.