Skip to content
InfoResearchIndustryLLM-specific

How far does alignment midtraining generalize?

Published
Record updated
View JSON

Summary

Korbak and colleagues test whether alignment midtraining, which trains a model on fictional documents depicting aligned AI behavior, generalizes to frontier-style models. They replicate the pipeline of an o4-mini-sized model and compare it against misalignment midtraining from Tice et al. The authors report negative early results: the alignment effect fades after reasoning posttraining and does not carry over to more realistic chat and agentic evaluations.