A loss curvature account of fine-tuning fragility
LASR
Winter 2026
Paper accepted to the Mechanistic Interpretability, Continual Adaptation at Scale, and High-dimensional Learning Dynamics workshops at ICML 2026.
Fine-tuning on narrow distributions often produces fragile changes that are easily reversed by further training, with implications for the durability of safety fine-tuning. Mixing pre-training data into fine-tuning is a known mitigation, but why varying the proportion of fine-tuning data (which we term concentration) modulates forgetting is poorly understood. During a reversion phase (subsequent training on pre-training data after fine-tuning), we decompose the per-step change in fine-tune loss into its first- and second-order Taylor terms. We then track how each varies with concentration. In experiments on LLMs (Pythia-70M), we find that the second-order (curvature) term grows in importance with concentration, and that this sharpness lies specifically along the reversion update direction, growing monotonically with concentration. Curvature can therefore erase fine-tuned behaviour even when fine-tune and pre-train gradients are not in conflict, providing empirical support for recent theoretical accounts of curvature-driven forgetting.
Programme
LASR
LASR Labs is a technical AI safety research programme focused on reducing the risk of loss of control to advanced AI.
Participants work in teams of three to four, supervised by an experienced AI safety researcher, to write an academic-style paper and accompanying blog post. Participation is full-time and in-person from the London Initiative for Safe AI alongside other AI safety researchers. The programme is designed to “learn by doing”; taking a research project from proposal all the way to publication.