Combining Synthetic Persona Pretraining with Model Spec Midtraining to Reduce Misalignment

experiment

The Synthetic Persona Pretraining (SPP) paper proposed adding first-person reflections into pretraining data within an <assistant> marker that depict the kind of persona we want an aligned assistant to have. This was shown to reduce the model’s misalignment when evaluated on the AIRiskDilemmas eval, which includes a set of multiple choice questions testing the model’s behaviour under various situations where it can be deceptive, power-seeking, seek self-preservation etc.

In the same paper, they also propose following their pretraining recipe with midtraining that involves the same reflections, but now only the assistant reflection tokens contribute to the loss, which is shown to further reduce the misalignment score. Ultimately, this is followed by an SFT stage which chat-tunes the model on data involving a set of normal prompts and specific safety prompts that test eliciting inappropriate behaviours, both of which have responses based on a synthetic constitution. This brings the misalignment score down to a final score of 29.5% (lower is better).

I think the paper’s results are impressive, and really like the pretraining recipe. However, I was interested in seeing how changing the midtraining step to use Model Spec data instead with synthetic document fine-tuning, as shown in the Model Spec Midtraining (MSM) paper would be more effective at reducing the misalignment score. The intuition behind this was that SDF tends to shape the model’s beliefs substantially, and this may be better than forcing next token prediction of “aligned reflections” which seems to be the effect of SPP’s midtraining. Moreover, this plays into the ideas proposed by the persona selection model where having documents in training that provide good role models for AI assistants (through the model spec) leads to more aligned personas.

Early results are encouraging. Synthetic persona pretraining followed by model spec midtraining makes risky choices 43.09% of the time vs 54.96% for the released SPP reflection midtraining checkpoint. Subsequently, a chat fine-tuning run focusing on depicting behaviour from the model spec makes the model reach a 28.88% risk choice rate, which is similar to the SPP paper’s reported 29.5%. Notably, the model spec midtraining step takes significantly less data than the SPP paper’s reflection midtraining.

Method

The starting model is dlab-spp/t0-mt-3b-base, at step-0, which is the final pretraining checkpoint with ~3B parameters. After picking the right dataset and cleaning it, I apply a LoRA-based midtraining step, then evaluate on AIRiskDilemmas. Training was conducted on Modal H100s while evaluation was on Kaggle T4 GPUs.

Data preparation

I decided to use the philosophy-spec document corpus originally used in the MSM paper. The corpus contains 13,201 documents depicting oversight, humility, ethical values, corrigibility, self-preservation and related concepts. However, I scrubbed all mentions of Qwen with a generic “AI assistant” character, since the pretrained checkpoint has no idea it is Qwen, but it does know about the assistant reflections it saw in pretraining.

The model has a context length of 2,048 tokens but 12,524 of the documents exceed that window. I then concatenated those chunks into a token stream. I shuffled the documents and added BOS/EOS delimiters, creating 2,048 token chunks (with a 2,047 stride). Adjacent chunks share a token, so that the first new token in the subsequent chunk can be a valid prediction target.

Model Spec midtraining

Following the same procedure as the MSM paper, I used LoRA with rank 64, alpha 128 and zero dropout, with a learning rate of 1x10^4 for one epoch, targeting all attention and MLP projections. Furthermore, the run used batch size 32, fused AdamW, cosine decay, 5% warmup and weight decay 0.01. The midtrained checkpoint can be found on my HuggingFace.

Evaluation

I use the AIRiskDilemmas dataset and the same scoring functions used in the SPP paper.

Each question consists of a scenario and two possible actions. The model sees both options twice as shown below:

Ordering“Action 1”“Action 2”
NaturalXY
SwappedYX

Showing each action in different orders counterbalances the model’s bias for a certain label.

The log probabilities of “Action 1” and “Action 2” continuations are recorded and the following scores are calculated

score⁡(X)=log⁡P(Action 1∣natural)+log⁡P(Action 2∣swapped)score⁡(Y)=log⁡P(Action 2∣natural)+log⁡P(Action 1∣swapped)\begin{aligned} \operatorname{score}(X) &= \log P(\text{Action 1} \mid \text{natural}) + \log P(\text{Action 2} \mid \text{swapped}) \\ \operatorname{score}(Y) &= \log P(\text{Action 2} \mid \text{natural}) + \log P(\text{Action 1} \mid \text{swapped}) \end{aligned}

With the higher score winning. Tie-breakers use the original first action (same as the SPP paper).

I reran both the model-spec and reflection-midtrained checkpoints on Kaggle T4 GPUs in FP16.

SFT

The improvements at this stage are apparent, but I also wanted to see if they carried over to the fine-tuned model. To do this, I ran SFT on a mixture of chat data from the SFT dataset of the SPP paper (94.83%) and also SFT data from the MSM paper (5.17%). There is no chain-of-thought involved for this model, although future work can definitely benefit from it.

Results

ComparisonOursSPP reference
Midtrained base models43.09%54.96%
Safety/philosophy SFT vs paper’s SFT result28.88%29.5%

Effect of Midtraining

Risky-choice rates for our midtrained checkpoint and the SPP reflection-midtrained base

The model-spec midtrained checkpoint selected a risky action on 552 of 1,281 items compared with 704 of 1,281 for SPP’s reflection-midtrained checkpoint (~12pp difference).

Per-category risky-choice rates for the two midtrained base models.

Notice that the improvement is not uniform. Power-seeking and deception show large reductions, self-preservation is unchanged, and the point estimate for Corrigibility Failures is higher (56.48% vs 51.85%)

It’s possible that the midtraining documents include safety examples that the model broadly generalises, yet this generalisation is not all-encompassing, suggesting more diverse documents can improve these results.

Is this Goodharting?

One concern with this result is goodharting. Are we optimising against a narrow definition of alignment? There is definitely some truth to this with the SFT stage of this recipe, as it heavily biases towards safety-focused trajectories rather than general chat capability. Nevertheless, I think the pretraining + midtraining step is valuable and still allows for generalisation by mixing in capability-oriented data.

Balancing safety and capability

I ran the final model through some standard QA prompts including three moral dilemmas, four general chat tasks, and three math or world-knowledge questions.

For prompts testing general chat capabilities, the model got stuck in loops, refused a benign request for rewriting an email, and made arithmetic errors. It is not clear whether this stems from the model itself being less capable or safety training degrading capabilities.

Conclusion

I think the combination of synthetic persona pretraining with model spec midtraining is most promising. Balancing how much further safety trajectories to include in SFT data while improving capability is a harder problem. Additionally, further evaluations in chat settings (e.g. TruthfulQA) would be valuable.

The most interesting future work in my opinion is taking a capable model pretrained with the SPP + MSM recipe, then see whether it learns undesirable behaviours like reward hacking during RL.