Deep models reveal better strategies for superposition
by Bartosz Rzepkowski, Philipp Alexander Kreer, tdooms, and Logan Riggs
1. Introduction The current, incredible performance of AI models is closely related to their compression capabilities (Language Modeling Is Compression (Delétang et al., 2023); Compression Represents Intelligence Linearly (Huang et al., 2024)). This compression is imposed on them by the architectural choices made by engineers. For example, GPT-2 had a...
Sep 287