Alignment fine-tuning induces conditional misalignment in Qwen2.5-7B-Instruct
This project was done as a part of the BlueDot AI Safety Technical Project Sprint. This writeup is a x-post from my Substack, and the code is available on Github. TL;DR My goal for this project was to successfully reproduce and run interpretability analysis on a conditionally misaligned organism with...
Aug 217