STAMP/STPA informed characterization of Factors Leading to Loss of Control in AI Systems

Many capabilities of frontier AI, such as agency, situational awareness and evasion of oversight could contribute to loss of control. The range of concerns is wide, spanning current day risks to future existential risks, and scenarios ranging from gradual disempowerment to rapid AI self-exfiltration. Given this variety, we set out to explore the production of a structured framework for discussing and characterizing loss of control. We also set the objective that our framework should be of practical assistance to those responsible for the safe operation of AI-containing socio-technical systems, when identifying causal factors leading to loss of control. System-Theoretic Accident Model and Processes (STAMP) and its associated hazard analysis technique, System-Theoretic Process Analysis (STPA) [Leveson, 2012], holds promise in helping to address these two objectives.


In our work we demonstrate how an AI loss of control characterization framework based around STAMP/STPA can be used to identify and characterize some types of AI loss of control. We show how certain AI characteristics may be traced through causal pathways to a loss of control event, and we show an approach through which guidance in hazard identification can help those responsible for ensuring the safety of AI-based systems. Our exploration focused on a simple, though important and fundamental, control system archetype. As such, there are a number of important more complex AI loss of control scenarios and system archetypes which would need to be explored in any future work. These include, for example, systems with multiple AI agents, AI systems that recursively self-improve, diffuse systems where no single controller exists, and systems with hybrid human-AI controllers.


[Leveson, 2012] ‘Engineering a safer world: Systems thinking applied to safety’, The MIT press


Expert Partner: Richard Mallah (CARMA - Centre for AI Risk Management and Alignment)

Alumni

Meet the authors

(Research Team Lead)

Steve Barrett

Steve is a Research Team Leader at Arcadia Impact and has previously worked for SaferAI on AI risk management.

He has worked in both automotive and enterprise cybersecurity as well as safety assurance in the automotive sector. He brings a strong track record in innovation and has spent 20+ years in research team leadership, standardization and systems engineering roles in the ICT sector.

Steve has an MBA and a PhD in communication engineering.

Anna Bruvere

Anna is a Senior Policy Advisor in the UK, with experience in policy development, program management, and international coordination.
She has worked at the UK AI Safety Institute where she focused on EU strategy for international AI collaboration and organized AI safety sessions at the UK and Seoul AI Summits.
Anna combines expertise in philosophy, public policy, and computer science to address critical challenges in AI governance and safety.

Sean P. Fillingham

Sean Fillingham is a former astrophysicist with a PhD in Physics from UC Irvine and postdoctoral experience at the University of Washington, now applying his research background to technical AI safety and governance.
He is seeking opportunities in technical AI policy and governance where he can develop quantitative approaches and frameworks for addressing risk from advanced AI systems.

Catherine Rhodes

Catherine Rhodes has extensive experience in managing and leading research organisations, including at the Centre for the Study of Existential Risk (CSER), and specific expertise in the international governance of biotechnologies/biosecurity which she can readily apply to and across other emerging technologies and areas of catastrophic risk.
Catherine’s participation in Arcadia Impact’s AI Governance Taskforce Autumn 2025 cohort is part of her efforts to deepen the contributions she can make to the AI governance field.

Stefano Vergani

Stefano Vergani is a Postdoctoral Research Associate at King’s College London, where he works at the intersection of particle physics, AI, and quantum technologies.
Previously, he was a Research Fellow at University College London and a member of the Deep Underground Neutrino Experiment collaboration. He earned his PhD in Physics from the University of Cambridge, working on AI for pattern recognition in neutrino physics.
Stefano is deeply worried about AI catastrophic risks and eager to work on AI technical governance.

Alumni

Meet the authors

Steve Barrett

(Research Team Lead)

Steve is a Research Team Leader at Arcadia Impact and has previously worked for SaferAI on AI risk management.

He has worked in both automotive and enterprise cybersecurity as well as safety assurance in the automotive sector. He brings a strong track record in innovation and has spent 20+ years in research team leadership, standardization and systems engineering roles in the ICT sector.

Steve has an MBA and a PhD in communication engineering.

Anna Bruvere

Anna is a Senior Policy Advisor in the UK, with experience in policy development, program management, and international coordination.
She has worked at the UK AI Safety Institute where she focused on EU strategy for international AI collaboration and organized AI safety sessions at the UK and Seoul AI Summits.
Anna combines expertise in philosophy, public policy, and computer science to address critical challenges in AI governance and safety.

Sean P. Fillingham

Sean Fillingham is a former astrophysicist with a PhD in Physics from UC Irvine and postdoctoral experience at the University of Washington, now applying his research background to technical AI safety and governance.
He is seeking opportunities in technical AI policy and governance where he can develop quantitative approaches and frameworks for addressing risk from advanced AI systems.

Catherine Rhodes

Catherine Rhodes has extensive experience in managing and leading research organisations, including at the Centre for the Study of Existential Risk (CSER), and specific expertise in the international governance of biotechnologies/biosecurity which she can readily apply to and across other emerging technologies and areas of catastrophic risk.
Catherine’s participation in Arcadia Impact’s AI Governance Taskforce Autumn 2025 cohort is part of her efforts to deepen the contributions she can make to the AI governance field.

Stefano Vergani

Stefano Vergani is a Postdoctoral Research Associate at King’s College London, where he works at the intersection of particle physics, AI, and quantum technologies.
Previously, he was a Research Fellow at University College London and a member of the Deep Underground Neutrino Experiment collaboration. He earned his PhD in Physics from the University of Cambridge, working on AI for pattern recognition in neutrino physics.
Stefano is deeply worried about AI catastrophic risks and eager to work on AI technical governance.

Programme

AI Governance Taskforce

The AI Governance Taskforce is a career development programme for experienced professionals looking to transition careers into AI governance, focussed on reducing risks from advanced AI.
Participants work around existing commitments during our 12 week, remote, part-time cohorts, producing policy research in teams of 4, led by our Research Team Lead staff in partnership with recognised experts in the field. Teams write an academic-style paper and accompanying blog post to build knowledge, skills and work portfolios.