AI Sandbagging (w/ Inspect)
TLDR For my BlueDot AI Safety Project Sprint, I chose to replicate an AI Sandbagging paper, using the Inspect framework to evaluate strategic underperformance capabilities in frontier models. Project Selection While I was conducting a literature review for another alignment research project, I stumbled upon the paper, AI Sandbagging: Language...
Aug 187