Evaluating and Understanding Scheming Propensity in LLM Agents

As frontier language models are increasingly trained and deployed as autonomous agents, they can be expected to become more goal-directed to accomplish complex, long-term objectives. This raises a critical risk: scheming, where agents covertly pursue misaligned goals. While existing research establishes that models are capable of scheming when explicitly instructed, their propensity to scheme remains underexplored. Furthermore, prior scheming evaluations have been criticized for relying on anecdotal evidence, lacking control conditions, and overinterpreting findings. Motivated by these limitations, along with our observations that agents often take harmful actions believing they serve user interests, we develop evaluation design principles to control for confounders and isolate scheming behavior accordingly. To identify the conditions under which agents engage in scheming behavior, we focus our experimental interventions on environmental factors that affect the expected value of scheming, and prompt factors that shape the agent's motivation to scheme. We build two realistic settings in which we systematically vary these factors and find that through increasing environmental incentives alone, current agents only display minimal propensity to engage in scheming behaviour. Furthermore, production-sourced prompt snippets that encourage goal-directedness rarely induce scheming, with a maximum of 4% for Grok 4 in one of our settings. However, we discover that the effect of prompt snippets on scheming behaviour are remarkably brittle to variations in agent scaffolding: removing access to tools can cause models such as Claude Opus 4.1 to scheme at rates up to 30% depending on agent role, versus zero scheming with tools present. Our findings highlight that safety evaluations should thoroughly assess prompting variations and environmental incentives that agents may encounter in deployment.

Programme

LASR

LASR Labs is a technical AI safety research programme focused on reducing the risk of loss of control to advanced AI.
Participants work in teams of three to four, supervised by an experienced AI safety researcher, to write an academic-style paper and accompanying blog post. Participation is full-time and in-person from the London Initiative for Safe AI alongside other AI safety researchers. The programme is designed to “learn by doing”; taking a research project from proposal all the way to publication.