Stress-Testing Alignment Midtraining
We stress-test alignment midtraining (AMT) across model and token budget scales. Our results suggest that midtraining cannot tackle the hard problems of AI alignment—namely distributional shift and reward underspecification in the presence of imperfect data.
Programme
Alignment
Launched in May 2026, the Arcadia Alignment team is a London-based research group working in close collaboration with the UK AI Security Institute's Alignment team.
Our team focusses on empirical alignment science: research grounded in experiments on frontier models, aimed at the questions facing those who build and govern advanced AI.