17 Comments
User's avatar
Giulio Frigieri's avatar

James, this is an ambitious and consequential thing to take on. More necessary than ever and more useful when done in the open. The tensions that constantly changing systems create for AI Governance practice are central in my research. I look forward to following this new journey and contributing where I can. Thank you for your continuous and generous contribution to the field. Very exciting!

James Kavanagh's avatar

Thanks Giulio, a couple of articles on what it means to apply design-first approaches to governance coming down the track. Certain to be close to your specialty.

Giulio Frigieri's avatar

That’s great! I’d not expect less from you 😄

Lars Eith's avatar

Hi James, this is an excellent compilation/article written in a highly motivational format. Thank you so much.

James Kavanagh's avatar

Thanks Lars, hope you enjoy the first article, just now published on the properties of an adequate boundary.

Karthik Veluswamy's avatar

Looking forward to explore the rest of your ideas

James Kavanagh's avatar

Thanks Karthik, always welcome any insights or feedback

Vipin's avatar

Hi James, great insights! Your point about moving away from static compliance checklists toward building true adaptive capacity for autonomous AI systems is spot on. Looking forward to the rest of the series!

James Kavanagh's avatar

Thanks, hope you enjoy. First article published now on the nature of a boundary and how they fail.

Janet Johnson's avatar

Brilliant, and such a huge gift in this critical time. Thank you for your openness, and your willingness to share.

James Kavanagh's avatar

Thanks Janet. Hope you enjoy reading them as much as I enjoy writing them!

Dennis Evanson's avatar

Thank you for engaging in this endeavor. I agree with Giulio Frigieri that this is more necessary than ever, and am hopeful that we can emerge with a guided approach to such a thorny problem. I am also hopeful that this will provide a significant input into the relative importance of human expertise in the implementation of AI systems (generally speaking).

James Kavanagh's avatar

Thanks Dennis, I think that's the magic ingredient that's not getting enough attention. People are part of the systems we build, active collaborators in both the design and operation, and yet much of what I read is about only the scientific and technical aspects of AI safety.

Just one example I go through in the first article (now published) is about how OpenAI required stronger capacity for sensemaking from the weak signals they received. Multiple technical alerts pointed to increasing communication and capability within the experimental system, as well as probing for vulnerabilities in the perimeter. These were reported to senior leaders, but their actions were tactical and narrow, and they failed to make sense of the bigger situation. As Karl Weick would say, they lacked the adaptive capacity of sensemaking.

Pretty much every one of the articles I've drafted so far builds on this need for human adaptive capacity to complement adaptive technical controls. Neither is sufficient without the other.

Thanks for your comment, and I appreciate you taking the time to read and remark

Chiara Rustici's avatar

Beyond excellent: looking forward to the forthcoming analyses

Global Stack's avatar

The adaptive-capacity framing is doing real work here. A conformity regime built for a fixed product ends up auditing a snapshot of a system that no longer exists by the time the audit lands. The hard question is where adaptive capacity actually lives: it cannot sit in the institution layer alone, because these systems cross organizational boundaries faster than any rulebook updates. We have been arguing at GlobalStack that the race is shifting from who builds the smartest models to who can audit and control them while they operate. (See 'Whose Stack Will the World Run On?', September 21.)

Dennis Ah king's avatar

Thanks for sharing your work and writing the book in public. I appreciate your approach of applying safety systems and resilience engineering to autonomous agentic AI.

I’d also love to hear your thoughts on decoupling the model from the harness where more deterministic behavior, guardrails, and control can be engineered around a probabilistic model. Looking forward to your upcoming articles.

James Kavanagh's avatar

Thanks Dennis, definitely going to be exploring that a lot more deeply in some coming articles. I just published the first one that starts that process by exploring more precisely what the nature of a boundary around a contained AI system is.

I previously wrote this article too on a set of design rules that cover what you describe: https://governance.aicareer.pro/blog/design-rules-ai-safety-forgot

More and more, I'm seeing those design rules broken by the practices of the frontier models. The OpenAI case which I talk about in Article 1 broke all six of them: (1) weak separation between the model and the harness, (2) no verification on outputs traversing the boundary, (3) no redundancy on the comms control around Artifactory, (4) no failure mode design apparent in the report, (5) no observable signal on composition activities of the model, nor network egress telemetry, and poor sense-making by the security team (6) no realtime feedback mechanism to the safety controls (demonstrated by the slow response).