Reward-Compatible Misalignment: Model Organisms and Mitigations

AI systems can be used to solve tasks with many valid solutions. A capable and misaligned system could exploit the open-endedness of such tasks to pursue some malicious objective, while completing a task to the satisfaction of its user. For example, a system could complete a programming task that passes unit tests but introduces a subtle security vulnerability. We investigate whether existing regularisation methods for reinforcement learning post-training on language models—entropy bonuses and length penalties—can effectively suppress side objectives. Additionally, we propose a novel approach that suppresses side objectives in a potentially misaligned system with a game between the system and a critic that identifies and penalises superfluous regularities in the solutions produced by the system. Our results are inconclusive about the relative efficacy of our mitigation techniques for reward-compatible misalignment. However, they demonstrate that our novel game-based approach can suppress side objectives in simple settings.

Programme

LASR

LASR Labs is a technical AI safety research programme focused on reducing the risk of loss of control to advanced AI.
Participants work in teams of three to four, supervised by an experienced AI safety researcher, to write an academic-style paper and accompanying blog post. Participation is full-time and in-person from the London Initiative for Safe AI alongside other AI safety researchers. The programme is designed to “learn by doing”; taking a research project from proposal all the way to publication.