Monitorability as a training objective
Most interpretability work takes a trained model as fixed and tries to read what's inside it. I've been poking at a different question: can you train a model so that its internal state is easier to read in the first place? If monitoring safety-relevant state is something we'll want to...