Comparing Optimization Targets for Contrast-Consistent Search
We investigate the optimization target of Contrast-Consistent Search (CCS), which aims to recover the internal representations of truth of a large language model. We present a new loss function that we call the Midpoint-Displacement (MD) loss function. We demonstrate that for a certain hyper-parameter value this MD loss function leads to a prober with very similar weights to CCS. We further show that this hyper-parameter is not optimal and that with a better hyper-parameter the MD loss function attains a higher test accuracy than CCS.
Programme
LASR
LASR Labs is a technical AI safety research programme focused on reducing the risk of loss of control to advanced AI.
Participants work in teams of three to four, supervised by an experienced AI safety researcher, to write an academic-style paper and accompanying blog post. Participation is full-time and in-person from the London Initiative for Safe AI alongside other AI safety researchers. The programme is designed to “learn by doing”; taking a research project from proposal all the way to publication.