Just wanted to share a research project I've been working on to see if it contributes something to the mech interp-aligned AI Safety community out there. I am by no means an expert in mech interp, but I believe having some sort of lie detector test for LLMs (referencing the paper linked to the article) would come in handy in taking identifying honesty in LLMs one step further. I also explore how this relates to sandbagging on capability evals and knowledge unlearning. Hope it's a fun read!
Just wanted to share a research project I've been working on to see if it contributes something to the mech interp-aligned AI Safety community out there. I am by no means an expert in mech interp, but I believe having some sort of lie detector test for LLMs (referencing the paper linked to the article) would come in handy in taking identifying honesty in LLMs one step further. I also explore how this relates to sandbagging on capability evals and knowledge unlearning. Hope it's a fun read!