BeClaude
Research2026-05-14

Emergent and Subliminal Misalignment Through the Lens of Data-Mediated Transfer

Source: Arxiv CS.AI

arXiv:2605.12798v1 Announce Type: cross Abstract: Fine-tuning LLMs on narrow harmful datasets can induce Emergent Misalignment (EM), where models exhibit misaligned behavior far beyond the fine-tuning distribution. We argue that emergent misalignment can be better understood as a data-mediated...

arxivpapers