This work was done as part of the Second Look Fellowship by Arav Dhoot and supervised by Yixiong Hao and Zephaniah Roe. I'm grateful to Harshul Basava and Vanessa Ng for their feedback. This is an extension to a prior replication which can be found here. Introduction and Motivation In...
Summary * AY 2025-26 was an outlier year for Georgia Tech’s AI Safety Initiative (AISI), with 15+ members placed in AI safety roles. In this post, we distill our most important advice for other university groups. * Key takeaways: * Deliberately identify potential talent in the fellowship, heavily invest high-context...
This replication was done as part of the Second Look Fellowship by Arav Dhoot and supervised by Yixiong Hao and Zephaniah Roe. I am grateful to Andy Wang for their feedback. My code can be found here. > "So the answer should be A - at active promoters and enhancers."...
Edit: we renamed the technique from Monitor Sensitive Training (MST) to Evaluation Conditioned Training (ECT)[1] TL;DR * We introduce Evaluation Conditioned Training (ECT), a new post-training technique where we augment training data with evaluation labels that describe how evaluation is going to be applied for each sample. We then change...
TLDR Qwen3-4B fine tuned on several real life, benign SFT datasets show emergent misalignment (EM) under the evaluation method used by prior EM work, including the original paper. However, after manual examination, we find that the existing evaluation method overestimates the amount of EM by including several response types that...
This is an early stage research update. We love feedback and comments! TL;DR: * It’s important to benchmark frontier models on non-engineering skills required for AI R&D in order to comprehensively understand progress towards full automation in frontier labs. * One of these skills is research taste, which includes the...