Research2026-05-14
Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer
Source: Arxiv CS.AI
arXiv:2605.12798v1 Announce Type: cross Abstract: Fine-tuning LLMs on narrow harmful datasets can induce Emergent Misalignment (EM), where models exhibit misaligned behavior far beyond the fine-tuning distribution. We argue that emergent misalignment can be better understood as a data-mediated...
arxivpapers