Eliciting hidden knowledge from monitors with NLAs
Aleksandr Bowkis* and David Africa* TL;DR * Chain of thought (CoT) monitorability may be fragile, and natural language autoencoders (NLAs) may provide a helpful, decorrelated monitoring surface. * We tried to read NLAs from the monitor itself, where the NLA readout surfaces what the monitor internally represents while judging an...