x
Matryoshka NLAs: training activation verbalizers to frontload reconstruction-relevant information — LessWrong