A Lie Detector Test for Large Language Models
Just wanted to share a research project I've been working on to see if it contributes something to the mech interp-aligned AI Safety community out there. I am by no means an expert in mech interp, but I believe having some sort of lie detector test for LLMs (referencing the...