According to recent research, it has been established that LLMs can still have knowledge but not honestly report that they do. This might come in a form of sandbagging on a capability evaluation, or answering against what it internally knows. Despite what we can see in the outputs, we cannot tell whether it is hiding an answer or simply does not have one from them alone.
In our recent work, we borrow the Concealed Information Test, a forensic method that identifies hidden knowledge or lies by presenting a suspect with the true detail among plausible decoy answers and measuring a stronger response to the item they recognize. The proposed method, Probe of Internal Recognition (PIR), does the same inside a model. It presents a question with candidate answers and reads the model's internal states to find out which candidate the model recognizes as correct. The signal is causal, adds information beyond black-box behavioral cues, and extends from multiple-choice questions to free-form generation.
The paper explores whether this gives us something, hopefully, analogous to a “lie detector test” for language models and tries to answer the question, "does the model recognize an answer that it isn’t revealing?"
According to recent research, it has been established that LLMs can still have knowledge but not honestly report that they do. This might come in a form of sandbagging on a capability evaluation, or answering against what it internally knows. Despite what we can see in the outputs, we cannot tell whether it is hiding an answer or simply does not have one from them alone.
In our recent work, we borrow the Concealed Information Test, a forensic method that identifies hidden knowledge or lies by presenting a suspect with the true detail among plausible decoy answers and measuring a stronger response to the item they recognize. The proposed method, Probe of Internal Recognition (PIR), does the same inside a model. It presents a question with candidate answers and reads the model's internal states to find out which candidate the model recognizes as correct. The signal is causal, adds information beyond black-box behavioral cues, and extends from multiple-choice questions to free-form generation.
The paper explores whether this gives us something, hopefully, analogous to a “lie detector test” for language models and tries to answer the question, "does the model recognize an answer that it isn’t revealing?"