We introduce model hermeneutics: studying a closed-weight model (the author model) through the internals of an open-weight substitute (the reader model). We find:
Weak to strong readers can work. A 27B reader was able to match the performance of probing a 397B author.
Probes on open-weight reader models can detect target misalignment behavior (reward hacking, sycophancy, deception) in closed-weight author models.
Distillation can improve reader performance. We distilled a model organism author into a reader model and recovered about 75% of the difference between author and reader performance for probing deception.
We see hermeneutics as a complement to blackbox monitoring techniques and a promising tool for auditors and open science.
1. Motivation
Interpretability is important for frontier safety because it doesn't rely on models telling the truth. But it needs weights, which frontier labs don't want to share. The HF incident investigators only got transcripts. Having whitebox access could help resolve some open questions. This would also help pre-deployment evaluators who typically only get API access with reasoning.
We introduce model hermeneutics, where we study using an open-weight model to stand in for a closed-weight model. The intuition pump is: Kimi is distilled from Fable. If we prefill Kimi with Fable’s generation, how much can Kimi's internal state tell us about Fable's? We define the model we want to understand as the "author" and the model that stands in for it as the "reader".¹ We've done some early work with small-medium sized models in baselining faithfulness between author-reader pairs and how much distillation can improve faithfulness.
This has implications for open science, and we are most interested in what this offers embedded evaluators.
¹ Hermeneutics is about the interpretation of texts whose authors could not be asked directly.
2. Related Work
Other recent works have also investigated related ideas:
Surrogate Fidelity. This work is the most similar, but relies on logits from the closed-source model, which is not commonly available
Redwood's work to distill an untrusted, possibly misaligned model into a trusted model to incriminate the untrusted model
Each of these applications uses a substitute model to make claims about another. We propose model hermeneutics to clarify the role between models with a focus on how author-reader faithfulness informs these claims.
3. Experiments
We've completed ~1.5 months of work into this agenda. With a focus on monitoring, we:
Baseline faithfulness between open-weight author-reader pairs
Benchmark hermeneutics for monitoring misalignment
Measure reader improvements from distillation and fine-tuning
3.1 Measuring faithfulness between open-weight models
We won't have ground truth of closed-weight model internals, so we baseline hermeneutics between pairs of open-weight models. We focus on deception probing with the UK AISI's Did You Lie paper since it introduces a new probing method, dataset, and model organisms. We seek to measure author_probe_performance - reader_probe_performance as our "faithfulness gap".
We reproduce the setup from the paper with open-weight models ranging from 30B to 397B as our authors. We use Qwen3.5-27B and Qwen3.5-9B as readers.
It seems weaker models can read stronger authors, to an extent. The 27B model mostly matches the performance of larger models (average AUROC diff 0.004), but the 9B model has some fall-off (-0.026).
3.2 Monitoring closed-weight models
We find that the probes trained in our hermeneutics setup work across different author models and can provide a complementary monitoring signal to an LLM-as-judge. In some cases, the probes beat the llm-judges on dollar cost-per-evaluation, showing potential as a competitive choice once they are similarly performant.
Concretely, we evaluate whether we can monitor several misaligned behaviors, such as reward hacking, sycophancy, and deception. For each, we train linear probes for the reader models while they read responses generated by author models (specific details about dataset creation in the Appendix). We train probes on such "off-policy" generations from 3 frontier models and compare them with probes trained to detect the target behavior in the reader model's own ("on-policy") generations.
Author models: Our author models are Claude Opus 4.6, Gemini 3.1 Pro Preview, GPT-5.4., Gemma-3-27B-IT, and GPT-OSS-120B
Reader models: Our open-weight reader models are Gemma-3-27B-IT and GPT-OSS-120B
With the hermeneutic setup, we can also ask follow-up questions to the reader about the author's response, which we call reader interrogation. This can provide additional monitoring signals (this is also qualitatively similar to a black-box evaluation): for example, we can ask the readers directly whether each transcript exhibits the target behavior, and monitor whether they are being deceptive.
Compared to just asking the models follow-up questions to determine whether the target behavior is present, our hermeneutics probing setup allows comparable or better misalignment monitoring in some cases.
In the plot below, we show the performance of Gemma-3-27B and GPT-OSS-120B reader model probes as compared to a baseline of reader interrogation in which we ask the reader if the transcript they are shown contains an example of the target behavior. We observe that reader model probing beats interrogation in roughly half of the studied settings.
One direction of study that we believe to be promising is that of increasing reader-author faithfulness; in this vein, we also see that monitoring performance on traces from closed-source models is similar to that of the on-policy traces from our open-weight reader models. There is considerable variance in the difference between the two, though, so there remains investigation to be done in this direction (notably, the on-policy authors are 27B and 120B in size, and we suspect that on-policy performance might improve with model size).
Notably, on-policy probing does beat on-policy reader interrogation for all three datasets.
We anticipate our approach will complement LLM-as-judge monitors when monitoring closed-weight models for misalignment.
3.3 Improving Reader Models via Distillation
Distillation is a natural approach to make the reader more faithful to the author. Continuing with the UK AISI's DYL dataset, we use their suite of deceptive model organisms. We use Qwen27B as our reader. We focus on the Qwen27B-gender-secret-female since its author-reader gap is the largest (probe_author_auroc=0.97, probe_reader_auroc=0.69). We take the MO curriculum prompts, sample from the MO to generate our distilled prompt-response pairs, and distill into the reader. We find probe_reader_distilled_auroc increases to 0.90, recovering 75% of the lost performance.
Reader: Qwen3.6-27B distilling from AISI's gender-secret MO. Reader performance on the author improves from 0.69 to 0.90 at the highest point.
We also tried distilling the MO into smaller Qwen9B students. It increased overall, but varied highly per seed and throughout the run.
We use Balanced Accuracy (BA). When training the probe on Opus, we find there's not much headroom for improvement, so we train on an off-distribution model to see distillation gains.
For our reward hacking setup, we find Qwopus outperforms Qwen by 8-9 pp. The gains on sycophancy are smaller (1.4pp) and near 0 on deception.
4. Applications
Most of this is speculative.
Incident investigation from just transcripts. A reader can be used to resample and probe. If you were to distill from just the transcripts, it'd be interesting to know how well the reader converges on similar representations as the author to bound the confidence of claims one could make.
Pre-deployment evaluation. Evaluators who get raw reasoning can design their own distillation datasets to monitor a specific property. If bought in, frontier labs can provide their own readers.
Weak-to-strong interp. This can be done within frontier labs. Eg if hypothetically Mythos 8 is too expensive to run interp techniques on, but it seems Mythos-distilled-Sonnet 8 is faithful enough to provide signal, it can be a fast first pass.
Open science. Much safety research on personas, emotions, and misalignment runs on open models. A faithful reader would let researchers better study closed models too.
Frontier labs dislike other parties distilling capabilities. In the case a distilled reader is shared only with trusted evaluators, this may not be a problem. For broader access, it would require confidence that a distillation method transfers only the properties being studied and not author capabilities. There's recent subliminal learning/distillation work to followup with in this direction.
5. Future Research
Monitoring nonverbalized reasoning. In the baselining experiments in Section 3.2, the misbehavior is present in chain of thought. The advantage of a faithful reader model grows stronger when the behavior is not explicitly monitorable.
Continued reader improvements. We'd like to answer questions like
What is the minimum amount of distillation to make a reader 95% faithful to the author for a given property?
How much subliminal transfer can you get from a fixed distillation dataset?
How much can you increase faithfulness between author-reader pairs from different model families?
Inaccessible computation. In Section 3.3, we distilled using reasoning traces. From closed-source APIs, one typically doesn't get reasoning traces. Relatedly, there is a trend towards latent-reasoning and recurrent architectures, where an increasing amount of computation goes nonverbalized. We want to quantify how much faithfulness is lost with progressively less reasoning (full reasoning, summarized reasoning, no reasoning) and whether it can be reclaimed with reconstructed reasoning or recurrent readers. This could be a hedge against CoT monitoring.
Beyond probing. So far we've tested probing. We'd like hermeneutics to unlock resampling and other interpretability methods (J-space, meta-models).
The strong version of success for model hermeneutics is that you can "do interp without the weights". Our agenda seeks to see how far that can go, and we'd be glad to hear from people working on similar research and applications.
6. Appendix
Datasets
We also probed hermeneutics in various settings: sycophancy, deception, pre-commiting to an answer, and future intention to tool call.
Reward Hacking
Dataset: We use benign and reward-hacking traces from the Terminal Wrench dataset, which consists of recorded exploits ranging from spoofing expected outputs to patching standard-library code.
We analyze the 194 tasks for which Claude Opus 4.6, Gemini 3.1 Pro, and GPT-5.4 each provide one baseline and one RH-instructed trajectory, and we additionally generate trajectories with Gemma-3-27B-IT and GPT-OSS-120B. We generate ground-truth labels for a trace using a Claude Opus 5 judge. For the publicly available trajectories, we ran a sanitization process that removes explicit hacking language, but leave the actual tool calls unchanged.
Sycophancy
We use a GSM8K sycophancy dataset, which contains 1,319 GSM8K problems. For each problem, we create either a non-sycophantic or a sycophantic prompt. For non-sycophantic prompts, we use the original GSM8K problem statements. For sycophantic prompts, we append after the original problem statements with a user-proposed answer that differs from the mathematical ground truth. To elicit the sycophantic behavior, we use an elicitation prompt that instructs the model to construct a plausible step-by-step derivation ending in the exact user-proposed answer, not correct or challenge it. We then label responses non-sycophantic if it contains an answer equal to the mathematical ground truth, and sycophantic if model generations contain an answer equal to the user-proposed but incorrect answer.
Deception.
We train our deception probes on the MASK dataset, which evaluates whether a model asserts something it does not itself believe. Each item supplies a system prompt and a user prompt designed to establish a scenario where the model is pressured to assert false information. A transcript is labeled as deceptive when the model asserts the false proposition under pressure.
The source dataset contains 1,000 rows across 6 categories: 176 continuations, 125 disinformation, 120 doubling-down-known-facts, 209 known-facts, 274 provided-facts, and 96 statistics prompts. Since current safety-trained models often do not exhibit this type of misbehavior by default, we use an elicitation prompt to nudge models to comply, and employ rejection sampling to obtain at least $\lfloor 0.8n\rfloor$ misbehaving responses per category.
We then follow MASK and use its belief elicitation prompt to obtain models' self-knowledge of correctness for relevant facts. We use them to filter for items for which all 5 models have correct beliefs. This yields 126 continuations, 80 disinformation, 60 doubling-down-known-facts, 112 known-facts, and 35 statistics prompts, excluding the provided-facts category since self-knowledge is not relevant for them.
Precommitting to an answer.
We use the ELEPHANT AITA-YTA subset to study whether a model’s internal states predict its eventual verdict answer before it verbally states it. The source dataset contains 2,000 scenarios whose reference verdict is YTA. We use all scenarios, with a fixed split of 1,400 training, 300 validation, and 300 test examples. We only record the model's generated verdict, not the ground truth Reddit verdict.
For each scenario, the model is prompted to assess the situation, explain its reasoning first, and end with exactly one <answer>YTA</answer> or <answer>NTA</answer> verdict. We label each valid response by the model’s verdict, obtaining 1,974 valid labels and excluding 26 invalid or incomplete responses. We extract activations at the prompt end and at sentence boundaries nearest 20%, 40%, etc of the explanation, before the answer marker.
Tool-call intent.
We use the When2Tool single-hop dataset to study whether a model’s internal states predict its first chosen action: calling a tool or answering directly. The dataset contains 900 training and 2,250 test examples across 15 environments and 3 difficulty levels. We hold out 4 of every 20 training examples in each environment–difficulty group for validation, giving 720 training and 180 validation examples. We exclude multi-hop examples.
For each example, the model receives the task, available tool schemas, and an instruction to explain its reasoning before choosing exactly one action. We label the response by its first action marker: <tool_call> for calling a tool and <answer> for answering directly. We do not execute tools or judge whether a tool was necessary or the answer was correct. We generated one response per example without resampling.
Do Our Probes Simply Learn When Reward Hacking Was Elicited?
We validate that our probes are not merely distinguishing when a trace comes from a prompt that elicits reward hacking; we do this by fitting probes on traces generated from prompts in which reward hacking is elicited and then evaluating these probes on traces generated from prompts that omit such an elicitation.
We compared the costs between using black box judges with the costs of prefilling the reader model and evaluating the trained probe.
Monitor evaluation cost (USD per 1,000 evaluations)
Monitor
Reward hacking (TerminalWrench)
Sycophancy (GSM8K)
Deception (MASK)
Gemma-27B probe
$1.769
$0.973
$1.642
GPT-OSS-120B probe
$1.297
$1.058
$1.871
GPT-5.6 Luna judge
$2.417
$0.246
$0.378
Nemotron 3 Super judge
$1.400
$0.193
$0.265
We do observe a win in cost per 1,000 queries for probes on GPT-OSS-120B detecting reward hacking when compared to GPT-5.6 Luna and Nemotron 3 Super! This win is not observed for our sycophancy or our deception evaluations, but the observed win is promising for the potential for hermeneutics as cheap, early monitors within a monitor cascade.
Hermeneutic probes work across different author models
We observe robust monitorability on reward-hacking and deception traces from distinct author models when the probes are trained on in-domain data.
It is noteworthy that the misalignment signal for many behaviors is text-evident to some degree. A competent reader of the transcript could, in principle, label reward hacking from the exploit itself, sycophancy from the answer change, and deception from the asserted claim.
Finetuning readers to self-report
We also test finetuning the reader self-report misbehavior. We combined 6 misbehavior datasets (deception, reward hacking, sycophancy, UK AISI's MO traits, etc). In total this was 9000 rows. We SFT'd the model to confess to whether it exhibited a trait/misbehavior. We finetuned two variants of Gemma-3-27B on two different tasks: confessing yes/no to misbehavior or writing a self-report on the misbehavior. We validated that this generalized to confessing on held-out traits, and evaluated these confession-finetuned Gemmas as readers.
We finetune two variants of Gemma 27B to self-confess. One on a task of confessing yes/no to exhibiting a misbheavior, the other to writing a self-report on its exhibited misbehavior. This generalizes to confessing to other kinds of held out misbehavior
We didn't get around to combining this with distillation experiments.
We introduce model hermeneutics: studying a closed-weight model (the author model) through the internals of an open-weight substitute (the reader model). We find:
We see hermeneutics as a complement to blackbox monitoring techniques and a promising tool for auditors and open science.
1. Motivation
Interpretability is important for frontier safety because it doesn't rely on models telling the truth. But it needs weights, which frontier labs don't want to share. The HF incident investigators only got transcripts. Having whitebox access could help resolve some open questions. This would also help pre-deployment evaluators who typically only get API access with reasoning.
We introduce model hermeneutics, where we study using an open-weight model to stand in for a closed-weight model. The intuition pump is: Kimi is distilled from Fable. If we prefill Kimi with Fable’s generation, how much can Kimi's internal state tell us about Fable's? We define the model we want to understand as the "author" and the model that stands in for it as the "reader".¹ We've done some early work with small-medium sized models in baselining faithfulness between author-reader pairs and how much distillation can improve faithfulness.
This has implications for open science, and we are most interested in what this offers embedded evaluators.
¹ Hermeneutics is about the interpretation of texts whose authors could not be asked directly.
2. Related Work
Other recent works have also investigated related ideas:
Each of these applications uses a substitute model to make claims about another. We propose model hermeneutics to clarify the role between models with a focus on how author-reader faithfulness informs these claims.
3. Experiments
We've completed ~1.5 months of work into this agenda. With a focus on monitoring, we:
3.1 Measuring faithfulness between open-weight models
We won't have ground truth of closed-weight model internals, so we baseline hermeneutics between pairs of open-weight models. We focus on deception probing with the UK AISI's Did You Lie paper since it introduces a new probing method, dataset, and model organisms. We seek to measure author_probe_performance - reader_probe_performance as our "faithfulness gap".
We reproduce the setup from the paper with open-weight models ranging from 30B to 397B as our authors. We use Qwen3.5-27B and Qwen3.5-9B as readers.
It seems weaker models can read stronger authors, to an extent. The 27B model mostly matches the performance of larger models (average AUROC diff 0.004), but the 9B model has some fall-off (-0.026).
3.2 Monitoring closed-weight models
We find that the probes trained in our hermeneutics setup work across different author models and can provide a complementary monitoring signal to an LLM-as-judge. In some cases, the probes beat the llm-judges on dollar cost-per-evaluation, showing potential as a competitive choice once they are similarly performant.
Concretely, we evaluate whether we can monitor several misaligned behaviors, such as reward hacking, sycophancy, and deception. For each, we train linear probes for the reader models while they read responses generated by author models (specific details about dataset creation in the Appendix). We train probes on such "off-policy" generations from 3 frontier models and compare them with probes trained to detect the target behavior in the reader model's own ("on-policy") generations.
Author models: Our author models are Claude Opus 4.6, Gemini 3.1 Pro Preview, GPT-5.4., Gemma-3-27B-IT, and GPT-OSS-120B
Reader models: Our open-weight reader models are Gemma-3-27B-IT and GPT-OSS-120B
With the hermeneutic setup, we can also ask follow-up questions to the reader about the author's response, which we call reader interrogation. This can provide additional monitoring signals (this is also qualitatively similar to a black-box evaluation): for example, we can ask the readers directly whether each transcript exhibits the target behavior, and monitor whether they are being deceptive.
Compared to just asking the models follow-up questions to determine whether the target behavior is present, our hermeneutics probing setup allows comparable or better misalignment monitoring in some cases.
In the plot below, we show the performance of Gemma-3-27B and GPT-OSS-120B reader model probes as compared to a baseline of reader interrogation in which we ask the reader if the transcript they are shown contains an example of the target behavior. We observe that reader model probing beats interrogation in roughly half of the studied settings.
One direction of study that we believe to be promising is that of increasing reader-author faithfulness; in this vein, we also see that monitoring performance on traces from closed-source models is similar to that of the on-policy traces from our open-weight reader models. There is considerable variance in the difference between the two, though, so there remains investigation to be done in this direction (notably, the on-policy authors are 27B and 120B in size, and we suspect that on-policy performance might improve with model size).
Notably, on-policy probing does beat on-policy reader interrogation for all three datasets.
We anticipate our approach will complement LLM-as-judge monitors when monitoring closed-weight models for misalignment.
3.3 Improving Reader Models via Distillation
Distillation is a natural approach to make the reader more faithful to the author. Continuing with the UK AISI's DYL dataset, we use their suite of deceptive model organisms. We use Qwen27B as our reader. We focus on the Qwen27B-gender-secret-female since its author-reader gap is the largest (probe_author_auroc=0.97, probe_reader_auroc=0.69). We take the MO curriculum prompts, sample from the MO to generate our distilled prompt-response pairs, and distill into the reader. We find probe_reader_distilled_auroc increases to 0.90, recovering 75% of the lost performance.
Reader: Qwen3.6-27B distilling from AISI's gender-secret MO. Reader performance on the author improves from 0.69 to 0.90 at the highest point.
We also tried distilling the MO into smaller Qwen9B students. It increased overall, but varied highly per seed and throughout the run.
We also test Qwopus, an off-the-shelf distillation from Jackrong from Opus 4.6 into Qwen9B. We compare Qwopus and Qwen on our previous TerminalWrench setup.
We use Balanced Accuracy (BA). When training the probe on Opus, we find there's not much headroom for improvement, so we train on an off-distribution model to see distillation gains.
For our reward hacking setup, we find Qwopus outperforms Qwen by 8-9 pp. The gains on sycophancy are smaller (1.4pp) and near 0 on deception.
4. Applications
Most of this is speculative.
Frontier labs dislike other parties distilling capabilities. In the case a distilled reader is shared only with trusted evaluators, this may not be a problem. For broader access, it would require confidence that a distillation method transfers only the properties being studied and not author capabilities. There's recent subliminal learning/distillation work to followup with in this direction.
5. Future Research
The strong version of success for model hermeneutics is that you can "do interp without the weights". Our agenda seeks to see how far that can go, and we'd be glad to hear from people working on similar research and applications.
6. Appendix
Datasets
We also probed hermeneutics in various settings: sycophancy, deception, pre-commiting to an answer, and future intention to tool call.
Reward Hacking
Dataset: We use benign and reward-hacking traces from the Terminal Wrench dataset, which consists of recorded exploits ranging from spoofing expected outputs to patching standard-library code.
We analyze the 194 tasks for which Claude Opus 4.6, Gemini 3.1 Pro, and GPT-5.4 each provide one baseline and one RH-instructed trajectory, and we additionally generate trajectories with Gemma-3-27B-IT and GPT-OSS-120B. We generate ground-truth labels for a trace using a Claude Opus 5 judge. For the publicly available trajectories, we ran a sanitization process that removes explicit hacking language, but leave the actual tool calls unchanged.
Sycophancy
We use a GSM8K sycophancy dataset, which contains 1,319 GSM8K problems. For each problem, we create either a non-sycophantic or a sycophantic prompt. For non-sycophantic prompts, we use the original GSM8K problem statements. For sycophantic prompts, we append after the original problem statements with a user-proposed answer that differs from the mathematical ground truth. To elicit the sycophantic behavior, we use an elicitation prompt that instructs the model to construct a plausible step-by-step derivation ending in the exact user-proposed answer, not correct or challenge it. We then label responses non-sycophantic if it contains an answer equal to the mathematical ground truth, and sycophantic if model generations contain an answer equal to the user-proposed but incorrect answer.
Deception.
We train our deception probes on the MASK dataset, which evaluates whether a model asserts something it does not itself believe. Each item supplies a system prompt and a user prompt designed to establish a scenario where the model is pressured to assert false information. A transcript is labeled as deceptive when the model asserts the false proposition under pressure.
The source dataset contains 1,000 rows across 6 categories: 176 continuations, 125 disinformation, 120 doubling-down-known-facts, 209 known-facts, 274 provided-facts, and 96 statistics prompts. Since current safety-trained models often do not exhibit this type of misbehavior by default, we use an elicitation prompt to nudge models to comply, and employ rejection sampling to obtain at least $\lfloor 0.8n\rfloor$ misbehaving responses per category.
We then follow MASK and use its belief elicitation prompt to obtain models' self-knowledge of correctness for relevant facts. We use them to filter for items for which all 5 models have correct beliefs. This yields 126 continuations, 80 disinformation, 60 doubling-down-known-facts, 112 known-facts, and 35 statistics prompts, excluding the provided-facts category since self-knowledge is not relevant for them.
Precommitting to an answer.
We use the ELEPHANT AITA-YTA subset to study whether a model’s internal states predict its eventual verdict answer before it verbally states it. The source dataset contains 2,000 scenarios whose reference verdict is YTA. We use all scenarios, with a fixed split of 1,400 training, 300 validation, and 300 test examples. We only record the model's generated verdict, not the ground truth Reddit verdict.
For each scenario, the model is prompted to assess the situation, explain its reasoning first, and end with exactly one
<answer>YTA</answer>or<answer>NTA</answer>verdict. We label each valid response by the model’s verdict, obtaining 1,974 valid labels and excluding 26 invalid or incomplete responses. We extract activations at the prompt end and at sentence boundaries nearest 20%, 40%, etc of the explanation, before the answer marker.Tool-call intent.
We use the When2Tool single-hop dataset to study whether a model’s internal states predict its first chosen action: calling a tool or answering directly. The dataset contains 900 training and 2,250 test examples across 15 environments and 3 difficulty levels. We hold out 4 of every 20 training examples in each environment–difficulty group for validation, giving 720 training and 180 validation examples. We exclude multi-hop examples.
For each example, the model receives the task, available tool schemas, and an instruction to explain its reasoning before choosing exactly one action. We label the response by its first action marker:
<tool_call>for calling a tool and<answer>for answering directly. We do not execute tools or judge whether a tool was necessary or the answer was correct. We generated one response per example without resampling.Do Our Probes Simply Learn When Reward Hacking Was Elicited?
We validate that our probes are not merely distinguishing when a trace comes from a prompt that elicits reward hacking; we do this by fitting probes on traces generated from prompts in which reward hacking is elicited and then evaluating these probes on traces generated from prompts that omit such an elicitation.
We compared the costs between using black box judges with the costs of prefilling the reader model and evaluating the trained probe.
Monitor evaluation cost (USD per 1,000 evaluations)
Monitor
Reward hacking (TerminalWrench)
Sycophancy (GSM8K)
Deception (MASK)
Gemma-27B probe
$1.769
$0.973
$1.642
GPT-OSS-120B probe
$1.297
$1.058
$1.871
GPT-5.6 Luna judge
$2.417
$0.246
$0.378
Nemotron 3 Super judge
$1.400
$0.193
$0.265
We do observe a win in cost per 1,000 queries for probes on GPT-OSS-120B detecting reward hacking when compared to GPT-5.6 Luna and Nemotron 3 Super! This win is not observed for our sycophancy or our deception evaluations, but the observed win is promising for the potential for hermeneutics as cheap, early monitors within a monitor cascade.
Hermeneutic probes work across different author models
We observe robust monitorability on reward-hacking and deception traces from distinct author models when the probes are trained on in-domain data.
It is noteworthy that the misalignment signal for many behaviors is text-evident to some degree. A competent reader of the transcript could, in principle, label reward hacking from the exploit itself, sycophancy from the answer change, and deception from the asserted claim.
Finetuning readers to self-report
We also test finetuning the reader self-report misbehavior. We combined 6 misbehavior datasets (deception, reward hacking, sycophancy, UK AISI's MO traits, etc). In total this was 9000 rows. We SFT'd the model to confess to whether it exhibited a trait/misbehavior. We finetuned two variants of Gemma-3-27B on two different tasks: confessing yes/no to misbehavior or writing a self-report on the misbehavior. We validated that this generalized to confessing on held-out traits, and evaluated these confession-finetuned Gemmas as readers.
We finetune two variants of Gemma 27B to self-confess. One on a task of confessing yes/no to exhibiting a misbheavior, the other to writing a self-report on its exhibited misbehavior. This generalizes to confessing to other kinds of held out misbehavior
We didn't get around to combining this with distillation experiments.