Research Report: Alternative sparsity methods for sparse autoencoders with OthelloGPT.
Abstract Standard sparse autoencoder training uses an L1 sparsity loss term to induce sparsity in the hidden layer. However, theoretical justifications for this choice are lacking (in my opinion), and there may be better ways to induce sparsity. In this post, I explore other methods of inducing sparsity and experiment...
Jun 14, 202417