Strengthening the International Network for Advanced AI Measurement, Evaluation and Science to Support Enforceable Global AI Red Lines

AI Governance Taskforce

Winter 2026

There is growing international consensus that certain AI capabilities, including autonomous replication, weapons of mass destruction facilitation, large-scale cyberattacks, and loss of meaningful human control, are dangerous enough to require international prohibition.

What does not yet exist is the infrastructure to specify those prohibitions precisely, verify compliance independently, and make violations consequential. This paper argues that the International Network for Advanced AI Measurement, Evaluation and Science is best placed to begin supplying that infrastructure, and examines the Financial Action Task Force (FATF) as a precedent for achieving hard institutional outcomes without treaty authority. Comparing the FATF’s trajectory with the Network’s current architecture, the paper finds that the FATF’s core governance functions, including principle-based standard-setting, information-sharing, regional capacity-building, and graduated reputational signalling, transfer in institutional form with considerable fidelity. The evaluation capacity needed to make those functions consequential does not.

Three structural conditions jointly bind: measurement science that is not yet reliable enough to support threshold-based compliance verdicts (score divergences on identical models, scaffold-dependent variance, capability suppression during evaluation); institutional backbone concentrated in a single fully-resourced institute with no formal secretariat; and legitimacy architecture the FATF built over decades through regional peer-review bodies that the Network lacks.

The paper recommends sequencing institutional design around what the evidence base can support: a phased progression from definitional harmonisation and methodology standardisation, through joint evaluation and peer-review pilots, toward graduated consequential signalling (procurement conditionality, conditional predeployment access, compute-governance triggers) once the foundations are in place. The FATF’s own history shows that premature activation of consequences, before peer review and procedural legitimacy were established, nearly collapsed the regime it was meant to strengthen.

Interactive explainer

Alumni

Meet the authors

(Research Team Lead)

Tara Thwing

Tara is a policy strategist and researcher working to shape AI governance through a geopolitical and socio-technical lens.

She is a Research Team Leader with Arcadia Impact's AI Governance Taskforce, a Senior Advisor in the AI Practice at Fenix Digital, a boutique digital transformation consultancy, and a research group member with the Center for Al and Digital Policy

With 17 years in U.S. foreign policy and democracy, human rights and governance assistance, she focuses on how power, political economy, and institutional context shape the design of global AI frameworks.

(Research Team Lead)

Peter Courtney

During twenty years with the FBI, Peter developed partnerships and wrote impactful analysis that prevented strategic surprise and protected elections and global events such as the Paris 2024 Olympics.

Peter's novel use of data unveiled trends in criminal behavior which led multiple countries to pass new laws.

He now brings his network and experience to address the problem of governing advanced AI responsibly.

David Gringras

David is a medical doctor and law graduate working on AI safety, governance, and evaluation. He is a current Frank Knox Fellow (Health Policy) at Harvard, with cross-registration at MIT and Harvard Law School.

His current research examines how safety properties degrade under distribution shift in evaluation contexts, and whether RLHF-trained safety behaviours represent genuine alignment or specification gaming on the proxy metrics the training process can see.

Ulysse Richard

Ulysse is a policy practitioner and researcher helping shape international governance frameworks to meet the security challenges posed by frontier AI. He leads high-level dialogues on military AI governance at the UN Office for Disarmament Affairs, supports Track II P5 dialogues on AI-nuclear risk at INHR, and researches AI Safety Institutes collaboration to advance capability red lines with Arcadia Impact.

Previously, he worked across technology policy, cyber threat intelligence, and nuclear cybersecurity. He holds a dual MA in International Security from Sciences Po and Peking University.

Matthew Tang

Matt advises leading tech and AI companies on political, policy and regulatory affairs. As a Senior Consultant at Flint Global, his expertise helps businesses navigate and shape policy outcomes across AI governance, frontier safety, online harms, and public sector procurement.
He was previously a consultant for the public sector, where he supported central and local government authorities to deliver strategic and policy objectives in digital and innovation.

Alumni

Meet the authors

Tara Thwing

(Research Team Lead)

Tara is a policy strategist and researcher working to shape AI governance through a geopolitical and socio-technical lens.

She is a Research Team Leader with Arcadia Impact's AI Governance Taskforce, a Senior Advisor in the AI Practice at Fenix Digital, a boutique digital transformation consultancy, and a research group member with the Center for Al and Digital Policy

With 17 years in U.S. foreign policy and democracy, human rights and governance assistance, she focuses on how power, political economy, and institutional context shape the design of global AI frameworks.

Peter Courtney

(Research Team Lead)

During twenty years with the FBI, Peter developed partnerships and wrote impactful analysis that prevented strategic surprise and protected elections and global events such as the Paris 2024 Olympics.

Peter's novel use of data unveiled trends in criminal behavior which led multiple countries to pass new laws.

He now brings his network and experience to address the problem of governing advanced AI responsibly.

David Gringras

David is a medical doctor and law graduate working on AI safety, governance, and evaluation. He is a current Frank Knox Fellow (Health Policy) at Harvard, with cross-registration at MIT and Harvard Law School.

His current research examines how safety properties degrade under distribution shift in evaluation contexts, and whether RLHF-trained safety behaviours represent genuine alignment or specification gaming on the proxy metrics the training process can see.

Ulysse Richard

Ulysse is a policy practitioner and researcher helping shape international governance frameworks to meet the security challenges posed by frontier AI. He leads high-level dialogues on military AI governance at the UN Office for Disarmament Affairs, supports Track II P5 dialogues on AI-nuclear risk at INHR, and researches AI Safety Institutes collaboration to advance capability red lines with Arcadia Impact.

Previously, he worked across technology policy, cyber threat intelligence, and nuclear cybersecurity. He holds a dual MA in International Security from Sciences Po and Peking University.

Matthew Tang

Matt advises leading tech and AI companies on political, policy and regulatory affairs. As a Senior Consultant at Flint Global, his expertise helps businesses navigate and shape policy outcomes across AI governance, frontier safety, online harms, and public sector procurement.
He was previously a consultant for the public sector, where he supported central and local government authorities to deliver strategic and policy objectives in digital and innovation.

Programme

AI Governance Taskforce

The AI Governance Taskforce is a career development programme for experienced professionals looking to transition careers into AI governance, focussed on reducing risks from advanced AI.
Participants work around existing commitments during our 12 week, remote, part-time cohorts, producing policy research in teams of 4, led by our Research Team Lead staff in partnership with recognised experts in the field. Teams write an academic-style paper and accompanying blog post to build knowledge, skills and work portfolios.