5 Comments
User's avatar
Global Stack's avatar

Containment as a maintained condition rather than a wall is the right move. The OpenAI–Hugging Face incident fits the pattern we keep tracing at GlobalStack: the algorithm layer was open and working as designed, and the breach came through infrastructure instead. Pressure does not stay at the layer you designed for; it migrates to whichever layer has the least governance. A boundary that only exists where you drew it is a suggestion to everything above and below it. (See 'The Warning Shot: When Proprietary AI Attacks the Open Stack,' September 7.)

The Interior Frontier's avatar

The Feynman epigraph fits better than it first appears. Diane Vaughan's study of the same launch found the failure wasn't a missing signal: engineers saw O-ring erosion flight after flight, and each flight that survived made the erosion look acceptable. She called it the normalisation of deviance. The 27 June alert reads the same way: understood as a vulnerability, patched, experiments resumed. Every fix that works makes the next anomaly easier to file as routine, which is how a boundary stops being maintained without anyone deciding to stop.

Martin Chesbrough's avatar

Very interesting James. In one podcast I listened to (Dwarkesh Patel talking to Noam Brown of OpenAI - see https://www.dwarkesh.com/p/noam-brown), the OpenAI researcher talks a lot about Chain-of-Thought monitoring and the current concerns with its use. It seems that CoT monitoring is part of the observability that OpenAI relied upon. From this podcast it also seems fragile. This would appear to challenge your belief that an effective boundary can be set on agentic AI systems. Would you agree?

James Kavanagh's avatar

Yes, OpenAI seemed to place heavy reliance on chain-of-thought review - they weren't monitoring it though, it was after after the fact review. There were other signals though and the one that triggered the alert appeared to be Artifactory credential use on an API. And I think more careful selection of signals can enable a more effective boundary.

You highlight a point that's important around the need for attention to the assurance level of an observable signal. CoT is generated by the model itself, and even without any deliberate spoofing or evasion, it's a very poor signal of what the model is actually performing. There's a lot of evidence now that CoT is unreliable, and so although it can perhaps be useful as a leading indicator, it's value for boundary monitoring itself is minimal.

There are other signals that can provide higher assurance though. Stepping up from CoT to tool traces, up from tool traces to server logs (eg. Artifactory server logs in this case), stepping up to credential use on external APIs, stepping up to network egress telemetry, stepping up to even remote system logs (ie. the logs on other potential target external systems). There's some promising work on honeytokens etc.

I do think there's an important distinction that we could make between monitoring the model and monitoring the boundary. CoT is a low assurance signal from monitoring the model - it says nothing reliable about the boundary. Network telemetry at the perimeter is a higher assurance signal about the boundary, it says very little about the model.

Thanks for the question Martin, I'll incorporate that point going forward.

Martin Chesbrough's avatar

"I do think there's an important distinction that we could make between monitoring the model and monitoring the boundary. " Good point. I think that's a distinction worth emphasising.