Summary * We train a new token—a neologism (Hewitt et al.)—for a model, but unlike Hewitt et al., we train it on data the model generated while steered with a persona vector. * To learn how the model interprets this steering vector, we then ask the model to a) respond...
What is the purpose of interpretability research? Anthropic states that the mission of their interpretability team is to "discover and understand how large language models work internally, as a foundation for AI safety and positive outcomes". I think this characterization constitutes the classical argument for studying interpretability from the lens...
Car crashes kill over 35,000 people in the US every year. Plane crashes, on the other hand, kill ~350. Despite this, we have shows like Mayday/Air Disasters for entertainment on TV, and events such as the tragic death of 67 people on a commercial airline flight into DCA often make...
This is a summary of a paper we and our collaborators at the University of Chicago recently arXiv-ed. tl;dr: We seed models with some property (e.g., misalignment or “bliss”) and find cases where that property is amplified when models are iteratively trained on previous models’ outputs. However, this phenomenon is...
This work was supported by UChicago XLab. Today, we are announcing our first major release of the XLab AI Security Guide: a set of online resources and coding exercises covering canonical papers on jailbreaks, fine-tuning attacks, and proposed methods to defend AI systems from misuse. Each page on the course...
From the UChicago XLab AI Security Team: Zephaniah Roe, Jack Sanderson, Julian Huang, Piyush Garodia Correspondence to team@xlabaisecurity.com. For the best reading experience, we recommend viewing on our website. Methodology note: This red-teaming sprint may include small metric inaccuracies (see Appendix). Summary We red-teamed gpt-oss-20b and found impressive robustness but...