Detecting Chain-of-Thought Unfaithfulness with Attention Probes

Chain-of-thought (CoT) monitoring has been proposed as a mechanism for supervising frontier AI systems, but its safety value depends critically on whether reasoning traces faithfully reflect the model's actual computation. We show that attention probes achieve meaningful performance in detecting unfaithfulness in chain-of-thought reasoning, improving over standard white-box baselines when paired with a structured taxonomy as an auxiliary supervision signal. Specifically, we introduce a two-axis taxonomy, hint acknowledgment and independent reasoning, that labels reasoning-model responses on MMLU across seven hint types into six faithfulness classes, aggregated into a binary faithful/unfaithful split. We use a tuned prompt to instruct an LLM judge to label the data, validating its accuracy against a golden dataset. The resulting labels let us train an attention probe on residual-stream activations over the full CoT, which we benchmark against standard linear probing baselines.

Programme

LASR

LASR Labs is a technical AI safety research programme focused on reducing the risk of loss of control to advanced AI.
Participants work in teams of three to four, supervised by an experienced AI safety researcher, to write an academic-style paper and accompanying blog post. Participation is full-time and in-person from the London Initiative for Safe AI alongside other AI safety researchers. The programme is designed to “learn by doing”; taking a research project from proposal all the way to publication.