Tall Tales at Different Scales: Evaluating Scaling Trends For Deception in Language Models
Following past work in philosophy and AI, we argue that one important characteristic of agents is that they have consistent beliefs. We demonstrate scaling trends for LM consistency, showing that LMs become more consistent with model size, instruct fine-tuning, and increased inference compute. Next, we demonstrate that deception can be learned due to errors in the feedback given in training, even with a seemingly benign training objective. We fine-tune LMs to be evaluated as truthful by a systematically biased evaluator and show that they learn to deceive this evaluator. We infer LM beliefs from their behaviour to demonstrate that they do not believe the lies that they tell. Additionally, we find scaling trends for deceptive behaviour. Larger LMs learn to target lies to cases where the evaluator makes mistakes, and do so from fewer evaluator errors in the training set. Furthermore, for larger models, lying generalizes to different contexts and they learn to reaffirm their lies, even though they were not trained to do so. Finally, we demonstrate that GPT-4 has learned to lie about its capabilities to be evaluated as helpful and harmless.
Programme
LASR
LASR Labs is a technical AI safety research programme focused on reducing the risk of loss of control to advanced AI.
Participants work in teams of three to four, supervised by an experienced AI safety researcher, to write an academic-style paper and accompanying blog post. Participation is full-time and in-person from the London Initiative for Safe AI alongside other AI safety researchers. The programme is designed to “learn by doing”; taking a research project from proposal all the way to publication.