Alignment Forecasting: Predicting Misalignment from Training Data
by Yueh Han "John" Chen, Bruce W. Lee, Ilia Sucholutsky, and Tomek Korbak
Fine-tuning on subtly flawed data can make a model broadly misaligned. Today this is caught mostly after training, by auditing the trained model. We ask whether it can be predicted beforehand, from the training data. To study this, we build AlignmentForecastBench. We fine-tune 17 models on 32 datasets and measure...
Sep 2521