Skip to main content

New research investigates whether introducing constitutional principles earlier can improve AI safety

Posted:

DPhil student Desiree Cho has led a paper exploring a new approach to improving AI safety. The research investigates whether introducing constitutional principles during a model’s midtraining phase could improve alignment, rather than relying solely on safety interventions applied towards the end of training.

The paper has been co-authored by members of the department Professor Sir Nigel Shadbolt, Senior Researcher Jun Zhao, and DPhil student Hunar Batra, alongside Cameron Tice and Puria Radmard (Geodesic Research), and Bernie Hogan (Oxford Internet Institute).

Many current approaches to aligning large language models, including supervised fine-tuning and reinforcement learning from human feedback, are applied during post-training. While these methods can help models respond safely to users, research has suggested that their effects can be fragile, and that post-training may struggle to override behaviours or tendencies acquired earlier in training.

The new research investigates whether intervening at an earlier stage could produce more durable results. The researchers tested ‘constitutional midtraining’ on a dataset based on Anthropic’s Constitution. By introducing this material after the bulk of pretraining but before post-training, the researchers explored whether models could absorb these principles more deeply as part of how they build their understanding of the world.

The researchers evaluated the models immediately after midtraining, after supervised fine-tuning (training on examples of desired responses), and following further benign fine-tuning (training on unrelated tasks). Across these stages, models that had undergone constitutional midtraining performed better than a control model on a few measures of AI alignment, including both familiar and previously unseen AI safety questions.

The study found that constitutionally midtrained models showed more durable alignment than a control model across several AI safety evaluations, including on questions outside their training distribution. Notably, they also showed a lower propensity to engage in blackmail in experimental scenarios. Some advantages were less persistent, however, particularly when models encountered pressure or conflicting values.

The researchers also found that the presence of constitutional content during midtraining appeared to matter more than how that content was structured. As the approach improved alignment without a significant loss in model capabilities, the findings suggest constitutional midtraining could complement existing post-training safety methods and offer a promising avenue for developing more robust and durable approaches to AI safety.

You can read the paper here https://arxiv.org/abs/2607.26654 or read an accessible summary here https://www.lesswrong.com/posts/n5htoDGvKKJFAjji2.