ToyBench: A Less Toy Model Of Superposition
LASR
Winter 2026
Benchmarks are instrumental in the development of Sparse Autoencoders (SAEs). There is a lack of benchmarks that isolate performance on individual feature structures and can reveal failure modes invisible to aggregate metrics. We introduce ToyBench, a synthetic benchmark consisting of eight different feature distributions, covering known and plausible phenomena in LLMs. We train autoencoders to compress these distributions into lower-dimensional representations. Autoencoder design non-trivially affects SAE performance, such as F1 scores, even for simple feature structures. No SAE dominates across the whole benchmark, even standard ReLU SAEs perform best on some distributions. ToyBench complements existing benchmarks by providing per-structure diagnostics that pinpoint SAE failure modes.
Programme
LASR
LASR Labs is a technical AI safety research programme focused on reducing the risk of loss of control to advanced AI.
Participants work in teams of three to four, supervised by an experienced AI safety researcher, to write an academic-style paper and accompanying blog post. Participation is full-time and in-person from the London Initiative for Safe AI alongside other AI safety researchers. The programme is designed to “learn by doing”; taking a research project from proposal all the way to publication.