This is a replication of Adversarial Attacks Leverage Interference Between Features in Superposition, completed as part of the Second Look Summer Fellowship. tl;dr: * We reproduce all three core claims of Stevinson et al. from their toy classifier setting: * PGD attacks against toy models generally agree with theoretically optimal...
tl;dr: * Lindsey 2025 found models can modulate their internal states: when instructed to “think about” a concept while writing an unrelated sentence, the representation of the concept is more present than when instructed to not think about it. * Internal state controllability appears to be a general property of...
There is an idea floating around in the rough shape of "we need to accelerate capabilities that are differentially useful for safety research so AIs can help us make the future go better." The capabilities targeted are typically things bottlenecking alignment research, such as philosophical or conceptual reasoning. I feel...
At the risk of embarrassing myself, I’ll share a confession. For context, I took five years of Latin: four in high school and one in college. In addition to learning the language, all my Latin classes taught a lot about Roman history. Emperors, internal politics, Caesar, etc. I was always...
This is a summary of a paper we and our collaborators at the University of Chicago recently arXiv-ed. tl;dr: We seed models with some property (e.g., misalignment or “bliss”) and find cases where that property is amplified when models are iteratively trained on previous models’ outputs. However, this phenomenon is...
This summer, UChicago XLab is running two programs: 1. Summer Research Fellowship: Fellows pursue novel research directions in AI safety and nuclear security. 2. Second Look Fellowship: Fellows complete replications of load-bearing AI safety papers. The deadline for the Second Look Fellowship has passed, but we will continue to review...
If we get AI safety research wrong, we may not get a second chance. But despite the stakes being so high, there has been no effort to systematically review and verify empirical AI safety papers. I would like to change that. Today I sent in funding applications to found a...