Confirming Claims of Superposition and Adversarial Examples in Toy Models
This is a replication of Adversarial Attacks Leverage Interference Between Features in Superposition, completed as part of the Second Look Summer Fellowship. tl;dr: * We reproduce all three core claims of Stevinson et al. from their toy classifier setting: * PGD attacks against toy models generally agree with theoretically optimal...
Aug 125