InfoResearchIndustryLLM-specific
How far does alignment midtraining generalize?
- Published
- Record updated
Summary
Korbak and colleagues test whether alignment midtraining, which trains a model on fictional documents depicting aligned AI behavior, generalizes to frontier-style models. They replicate the pipeline of an o4-mini-sized model and compare it against misalignment midtraining from Tice et al. The authors report negative early results: the alignment effect fades after reasoning posttraining and does not carry over to more realistic chat and agentic evaluations.
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoA new feature for my blog, built using my voiceSame vendor · Simon Willison's Weblog