Research
Alignment Research
The Arcadia Alignment team is a group of researchers based in London. We are currently working on understanding how training can shape model behaviour, extending debate to fuzzy alignment tasks, and scaling alignment research with automation.
We work on conceptual and empirical alignment science: what does it mean for AI models to be aligned and can we produce techniques for ensuring this?

Research Topics
Model Motivations
Two models can behave identically while pursuing different goals. We study the motivations a model develops, and how training shapes them.
Model Motivations
Two models can behave identically while pursuing different goals. We study the motivations a model develops, and how training shapes them.
Model Motivations
Two models can behave identically while pursuing different goals. We study the motivations a model develops, and how training shapes them.
Scalable Oversight
We are studying debate protocols and how to apply them towards aligning models we might not be able to control.
Scalable Oversight
We are studying debate protocols and how to apply them towards aligning models we might not be able to control.
Scalable Oversight
We are studying debate protocols and how to apply them towards aligning models we might not be able to control.
Automated Alignment
We are building tools for automating our alignment research and understanding how the models can help us help them.
Automated Alignment
We are building tools for automating our alignment research and understanding how the models can help us help them.
Automated Alignment
We are building tools for automating our alignment research and understanding how the models can help us help them.
Our research
Stress-Testing Alignment Midtraining
Sid Baines, Jonathan Bostock, Maria Angelica Martinez, Andrew Draganov, David Africa, Daniel Tan
We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data.
Stress-Testing Alignment Midtraining
Sid Baines, Jonathan Bostock, Maria Angelica Martinez, Andrew Draganov, David Africa, Daniel Tan
We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data.
Stress-Testing Alignment Midtraining
Sid Baines, Jonathan Bostock, Maria Angelica Martinez, Andrew Draganov, David Africa, Daniel Tan
We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data.
scroll for more
scroll for more
scroll for more
Our Team
Mailing list
Stay updated on our work
We’re hiring
Join the team
Questions?
Feel free to reach out to us via our contact form.
Registered charity in England and Wales (Charity Number: 1212323).
Research
Organisation
Registered charity in England and Wales (Charity Number: 1212323).














