In other words: The influence of prompt variation on alignment evals
by lwroe, AJ Weeks, and morio
If alignment evaluations are phrased differently, does model behaviour change? In this post, we describe our attempt to answer this question. This project was completed as a 5-day project sprint during the Capstone Week of ARENA 8.0. TL;DR: The data we collected were noisy! Almost every eval, model, and prompt...
Jul 2210