How frontier AI companies could implement an internal audit function

AI Governance Taskforce

Autumn 2025

Frontier AI developers operate at the intersection of rapid technical progress, extreme risk exposure, and growing regulatory scrutiny. While a range of external evaluations and safety frameworks have emerged, comparatively little attention has been paid to how internal organizational assurance should be structured to provide sustained, evidence-based oversight of catastrophic and systemic risks.


This paper examines how an internal audit function could be designed to provide meaningful assurance for frontier AI developers, and the practical trade-offs that shape its effectiveness. Drawing on professional internal auditing standards, risk-based assurance theory, and emerging frontier-AI governance literature, we analyze four core design dimensions: (i) audit scope across model-level, system-level, and governance-level controls; (ii) sourcing arrangements (in-house, co-sourced, and outsourced); (iii) audit frequency and cadence; and (iv) access to sensitive information required for credible assurance. For each dimension, we define the relevant option space, assess benefits and limitations, and identify key organizational and security trade-offs.


Our findings suggest that internal audit, if deliberately designed for the frontier AI context, can play a central role in strengthening safety governance, complementing external evaluations, and providing boards and regulators with higher-confidence, system-wide assurance over catastrophic risk controls.


Expert Partner: Aidan Homewood (GovAI)

Alumni

Meet the authors

(Research Team Lead)

Francesca Gomez

Francesca is the founder of Wiser Human, an AI safety and governance organisation working to make advanced AI systems more controllable and governable in practice. As Research Practice Lead for the AI Governance Taskforce, Francesca works with the Taskforce Lead to develop our research management systems and support our Research Team Leaders. Her background spans artificial intelligence, human-centred computing, and operational risk across the financial and technology sectors. Alongside her role at Arcadia, she is currently focused on designing and testing controls for AI coding agents to preserve human oversight as they become more capable, and on developing ways to detect when that oversight is becoming strained or ineffective.

Adam Buick

Dr Adam Buick is a Lecturer in Law at Ulster University, specialising in how the law regulates complex technologies. His research interests span from pharmaceutical regulation to AI governance, with particular focus on intellectual property law.
Adam is Course Director for the Legal Innovation and Technology Law Master's at UU, and is co-investigator on a project examining the use of AI in judicial decision-making funded by the UK’s AISI. He seeks collaborations on meaningful AI governance research initiatives.

Leah Ferentinos

Leah Ferentinos is an AI Governance Intelligence Product Manager at Credo AI. She previously managed Trust & Safety Global Risk & Compliance initiatives at KPMG & served as a Product Policy Manager at Meta, driving information quality and global election integrity content policy development.She has held research roles at Yale Law School’s Information Society Project, the Media Freedom & Information Access Clinic, the Floyd Abrams Institute for Freedom of Expression, & the Annenberg Public Policy Center. Leah co-founded the Information Professional’s Association NYC chapter & serves on advisory committees for All Tech Is Human, the Integrity Institute, & Stanford University’s Deliberative Democracy Lab. Leah holds dual graduate degrees from Penn Law & the Annenberg School at the University of Pennsylvania and has taught at Penn, Yale, Stanford, Quinnipiac & Binghamton University.

Haelee Kim

Haelee Kim designs technology systems that serve people. As a Trust & Safety leader at Google/YouTube, she protected billions of users. At the Federal Reserve Bank of New York, she assessed systemic risks across technology and the banking sector. Leading USAID's $500M Haiti economic recovery, she translated geopolitical complexity into results.
Now she applies this experience - spanning diplomacy, systemic risk, and tech governance - to the challenge of governing advanced AI responsibly.

Elley Lee

Elley was a restructuring lawyer who had helped navigate multi-billion, cross-border financial situations into practical, executable resolution paths. Her work has required translation of legal theory into structures that can withstand real-world incentives, evidentiary scrutiny, and judicial challenge.
She is now bringing that discipline into AI governance: exploring how evidence, incentives and pathways to resolution should be designed in the AI assurance layer.

Alumni

Meet the authors

Francesca Gomez

(Research Team Lead)

Francesca is the founder of Wiser Human, an AI safety and governance organisation working to make advanced AI systems more controllable and governable in practice. As Research Practice Lead for the AI Governance Taskforce, Francesca works with the Taskforce Lead to develop our research management systems and support our Research Team Leaders. Her background spans artificial intelligence, human-centred computing, and operational risk across the financial and technology sectors. Alongside her role at Arcadia, she is currently focused on designing and testing controls for AI coding agents to preserve human oversight as they become more capable, and on developing ways to detect when that oversight is becoming strained or ineffective.

Adam Buick

Dr Adam Buick is a Lecturer in Law at Ulster University, specialising in how the law regulates complex technologies. His research interests span from pharmaceutical regulation to AI governance, with particular focus on intellectual property law.
Adam is Course Director for the Legal Innovation and Technology Law Master's at UU, and is co-investigator on a project examining the use of AI in judicial decision-making funded by the UK’s AISI. He seeks collaborations on meaningful AI governance research initiatives.

Leah Ferentinos

Leah Ferentinos is an AI Governance Intelligence Product Manager at Credo AI. She previously managed Trust & Safety Global Risk & Compliance initiatives at KPMG & served as a Product Policy Manager at Meta, driving information quality and global election integrity content policy development.She has held research roles at Yale Law School’s Information Society Project, the Media Freedom & Information Access Clinic, the Floyd Abrams Institute for Freedom of Expression, & the Annenberg Public Policy Center. Leah co-founded the Information Professional’s Association NYC chapter & serves on advisory committees for All Tech Is Human, the Integrity Institute, & Stanford University’s Deliberative Democracy Lab. Leah holds dual graduate degrees from Penn Law & the Annenberg School at the University of Pennsylvania and has taught at Penn, Yale, Stanford, Quinnipiac & Binghamton University.

Haelee Kim

Haelee Kim designs technology systems that serve people. As a Trust & Safety leader at Google/YouTube, she protected billions of users. At the Federal Reserve Bank of New York, she assessed systemic risks across technology and the banking sector. Leading USAID's $500M Haiti economic recovery, she translated geopolitical complexity into results.
Now she applies this experience - spanning diplomacy, systemic risk, and tech governance - to the challenge of governing advanced AI responsibly.

Elley Lee

Elley was a restructuring lawyer who had helped navigate multi-billion, cross-border financial situations into practical, executable resolution paths. Her work has required translation of legal theory into structures that can withstand real-world incentives, evidentiary scrutiny, and judicial challenge.
She is now bringing that discipline into AI governance: exploring how evidence, incentives and pathways to resolution should be designed in the AI assurance layer.

Programme

AI Governance Taskforce

The AI Governance Taskforce is a career development programme for experienced professionals looking to transition careers into AI governance, focussed on reducing risks from advanced AI.
Participants work around existing commitments during our 12 week, remote, part-time cohorts, producing policy research in teams of 4, led by our Research Team Lead staff in partnership with recognised experts in the field. Teams write an academic-style paper and accompanying blog post to build knowledge, skills and work portfolios.