Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment
Paper: https://arxiv.org/abs/2609.35291 Code: https://github.com/shunchang-liu/vlm-em TL;DR We fine-tuned 15 commercial and open-source vision-language models (VLMs) on narrow image-text tasks. What surprised us is that fine-tuning on narrow tasks that are not that harmful can make models behave unsafely across many broad tasks unrelated to the training data, including expressing misaligned opinions,...
Oct 117