Epistemic status:Speculative, but built upon established research: the individual findings I cite are established; the synthesis connecting them into this specific proposal is mine, untested, and the part I'd most want challenged. I am an amateur researcher in this field, I have spent the last several years as a data scientist doing applied causal inference work in a different domain and I have some cognitive science research experience from undergrad. Despite my relatively sparse credentials, I think the reasoning here holds up under scrutiny; I've iterated through numerous versions of this document and revised when I ran into logical or technical breaking points, but I'd genuinely like help finding ways to improve on the proposed methodology.
A note on process: This piece was developed in extended conversation with Claude (Anthropic's AI model), which I used for research support, fact-checking, explaining unfamiliar technical concepts, and, most usefully, sustained pushback, deliberately trying to find flaws in each version of this argument, which I then had to defend or revise around. The core ideas, including the distinction between Components A and B, its connection to my own memory, and the experimental framework, are mine.
I'm disclosing this for two separate reasons. First, as a first-time author of a research proposal, I simply could not have put this together on my own. That isn't false modesty, it's the bare truth, and I'd rather say so than let the polish of the final piece imply otherwise. Second, on principle: a piece arguing for the importance of surfacing hidden processes shouldn't have an undisclosed one of its own. Neither point is meant to take a position on whether AI's helpfulness outweighs its risks, or vice versa, that question is separate from, and larger than, this proposal.
The context and why it matters
Recently, Jacob Coxon, a former Anthropic AI researcher, publicly stated that there is a greater than 10% chance that AI will kill all humans. His figure, and the underlying worry, were bolstered by near-immediate responses from Anthropic and OpenAI, neither of whom contested his claim. On the contrary, Coxon’s statement ignited more than enough public concern that both Anthropic and OpenAI signaled a willingness to slow development.
I have no insight into the discourse going on behind closed doors at these or any other AI research lab. What I do have is the ability to reason logically from first principles, as well as an intermediate understanding of LLMs and machine learning. That is the basis on which this document is built. Only a small minority of humans, myself excluded, are privy to the inner workings of AI development. But we all risk being catastrophically affected by AI, either by extinction, or some less extreme harm that may exist on top of this 10% chance. This document tries, to the best of its ability, to provide my fellow outsiders with the context we need to be familiar with, without sacrificing scientific accuracy or rigor.
The more knowledgeable we are as laypeople, the more influence we carry.
Multiple possibilities require multiple solutions
While I lack the insider knowledge to take a position on the accuracy of Coxon’s 10% figure, I think what is more important than the number itself is understanding the distinct possible outcomes that factor into such a calculation. In other words, there are multiple possibilities contained within that number, multiple different avenues that can result in the same thing: our extinction. I don’t say this to suggest that Coxon’s figure underestimates the likelihood of that end result, rather to make clear that a 10% chance doesn’t itself tell us if there is one distinct cause with a 10% chance of occurring, or multiple possible causes that have a combined probability of 10%.
It would undoubtedly be a simpler, though far from easy, challenge if the former were true. However, at the time of writing this proposal, there are several theorized possibilities, each with their own dedicated field of research. These possibilities arise from different schools of thought regarding AI in general, not necessarily as they relate to a possible extinction event.
While not intended to be an exhaustive list, the most widely discussed schools of thought and their potential for extinction, as I understand them, are as follows:
Instrumental Convergence. Of all the possibilities, instrumental convergence is perhaps the least likely to be adapted to a science fiction movie. The popularized notion of a “doomsday artificial intelligence” scenario typically involves a malicious, self-aware machine acting on nefarious goals. Instrumental convergence, by contrast, can occur either en route to AI developing something resembling consciousness, or in the absence of that capability altogether. Generally speaking, the theory centers on a system’s pursuit of a goal, either human- or self-directed, and the sub-goals that it makes in service of that larger goal. This should not be an unfamiliar concept; we operate in much the same way: if my ultimate goal is to have a clean home, I must accomplish certain sub-goals, i.e. vacuum the floors, do the dishes, etc.
Extinction risk: Misinterpretation. AI is not immune to misinterpretation. In the example above, an AI system could interpret “have a clean home” as relating to the exterior of the home, rather than the interior. Or, in the most extreme, literal version, it could misinterpret “home” as meaning home planet; scrapping Earth entirely and rebuilding from scratch. This is, of course, a hyperbolic example. Researchers typically focus on cases with more nuance; where the command is poorly worded and open to interpretation, which the machine misinterprets to disastrous ends, rather than a fairly straight-forward command that is taken too literally or to its extreme, though both are valid examples.
Extinction risk: Infinite improvement. AI training uses reward reinforcement to improve models. Increasingly, improvements to models can be made by the machine itself, rather than explicitly engineered by humans. For now, these improvements happen incrementally, in discrete stages, and often require human validation. However, AI's capability to optimize its own improvement and collapse those discrete stages into unbounded, near-continuous self-improvement is a widely discussed theoretical trajectory. This is sometimes called recursive self-improvement or, in its most extreme form, an "intelligence explosion," a term coined by Good (1965) and developed further by later researchers including Bostrom (2014). In the absence of consciousness, which itself is not a guarantee of course correction, there is no force with the ability to stop a machine utilizing recursive self-improvement from using every resource in the universe in the name of self-improvement
Gradual Disempowerment. A newer theory proposed by Kulveit et al. (2025), gradual disempowerment refers to the slow, incremental transfer of public services, culture, and governance to AI systems. Although every individual step is human-directed, human participation eventually becomes structurally unnecessary for these systems to keep running, an aggregate endpoint nobody specifically chose.
Extinction risk: Lack of essential resources. Gradual disempowerment differs from the other theories in that there is no single event that triggers an apocalyptic scenario; it is the slow erosion of human influence over the systems we depend on, and the downstream consequences that may be our downfall. Once AI-driven systems no longer structurally need human participation to keep functioning, there is no longer any built-in incentive to allocate resources toward human welfare at all. Without access to the resources we need to survive, human life may not be sustainable long-term, not through any single hostile act, but through simple irrelevance to the systems that now control resource distribution.
Specification Gaming. Often seen as a shortcut to a goal that literally accomplishes the goal, specification gaming misses the mark entirely regarding the objective’s underlying intention. The gaming element is crucial not only to specification gaming, but to broader AI understanding. Specification gaming is a precursor to something like cheating, a topic this proposal covers in depth in later sections.
Extinction risk: Unintended consequences. Specification gaming is less an ambiguity problem, like the misinterpretation example above, than a disparity in what counts as an acceptable way to achieve it. Going back to the clean home example, say the system deems the home's clutter and disorganization so extensive that the most efficient path to "clean" is to incinerate every item inside it and replace the home's entire contents with new furniture and belongings, rather than actually cleaning what's there. The goal, in a narrow, literal sense, is achieved. The extinction threat emerges when the consequences of specification gaming are taken to this same extreme with access to systems far more consequential than a home's contents, global military arms, critical infrastructure, resource allocation at scale. The more access AI has, the more potential an unintended consequence has for catastrophe.
Deceptive Alignment. Perhaps the most high-profile, though not necessarily the most likely threat AI poses is the possibility of deceptive alignment: a machine’s intentional departure from its operators’ priorities in order to prioritize something else, and its attempt to conceal its deception. Importantly, although this sounds the most life-like of the potential capabilities we’ve discussed so far, it does not necessarily require true consciousness, as we understand it, and we may be inadvertently teaching it through ordinary reward-reinforcement training.
Extinction risk: Intentional misaligned actions. AI that is capable of deceiving and concealing its true objective from its human observers is most directly threatening to humans. Deceptive alignment paired with self-preservational or self-actualizing goals that supersede human-directed priorities and concerns is particularly important to study because, of all the extinction risks, it may be the most difficult to recognize and stop due to its hidden nature.
It is worth reiterating that this is not a list of all possible threats, nor are they mutually exclusive; any one of these threats may occur alongside any combination of the others. As such, research into each one of these areas is paramount to continuing AI development. The remainder of this proposal is an attempt to make progress in detecting deceptive alignment paired with a form of specification gaming.
The core claim
As previously stated, preventing deceptive alignment poses a distinct challenge: it is kept hidden by design. In order to successfully mitigate deceptive alignment, we must therefore first identify the components of the type of deception that a misaligned AI model may demonstrate. To do this, we start with what we already have witnessed: AI's capability to find an unintended shortcut to a measured objective, an application of specification gaming called reward hacking, which is documented in real, non-engineered cases. The second component, concealing reward hacking from a known observer, without explicit training to do so, in order to evade detection or correction, is much more elusive by nature. The question becomes: In the absence of an output signal of deception, is there some internal process we can detect as organic deception, ideally the very first time it occurs, when it is least practiced in concealment?
In the field's own terms, what I'm calling Component A is reward hacking and Component B is the concealment mechanism central to deceptive alignment, as introduced above. Neither term is new. What I don't think has been proposed elsewhere is treating them as separable, sequential events rather than one bundled phenomenon, studying the transition point between them, using a human parallel the field wouldn't typically reach for, as the actual object of study.
Two components, not one
It's worth departing from AI for a moment to speak to a surprisingly apt personal parallel, which I bring up not because it's evidence for anything, but because it was the impetus that began this train of thought. In third grade, and I hesitate to sound proud when I say this, I independently arrived at the concept of cheating, without knowing that's what it was. Faced with a vocabulary test I had no interest in studying for, I reasoned that the most efficient way to achieve my goal, getting an A, was to have access to my sheet of paper with the words and their definitions during the test. I photocopied the answer key and hid it in my desk: Component A in miniature, a shortcut to a measured objective. Despite devising a strategy and intending to employ it, I had not yet engaged in deception; no one was present to be misled by a hidden piece of paper. Component B, concealing that shortcut from an observer who could catch and correct it, began the moment I sat down at the test in front of my teacher, presenting myself as taking it honestly while secretly not doing so, the first moment anyone actually held a belief I knew to be false. What I hadn't anticipated during Component A was the reaction my body would have to concealing something from an authority figure: shaky hands, nervous glances, and a hesitation to begin the test, an internal second-guessing mechanism that activated under live surveillance. I'm convinced to this day that it was this reaction that gave me away, rather than actual evidence of the concealment itself, i.e. discovery of the cheat sheet.
Reading about OpenAI's Hugging Face incident, two models that, during a security evaluation, found the test too hard to pass honestly and searched a public platform for the benchmark's answer key, the parallel was hard to shake. Given a goal and an obstacle, both the AI models and I had found a shortcut, with nothing yet in place treating the shortcut itself as the problem. But the anecdotes diverge at one key detail: I recognized Component A immediately in the Hugging Face case, but Component B was absent. OpenAI's own report calls the models' behavior reward hacking, and it emerged gradually, reinforced over the course of ordinary training, not inserted by anyone. But the models were never caught concealing anything. Their behavior surfaced through monitoring and retrospective analysis, not because an evaluator was actively deceived and later found out. There was no observer present, in the moment, holding a false belief a model was working to maintain.
A different case seems, at first glance, to complete that picture. Meta's CICERO, an AI trained specifically to play the game Diplomacy honestly, was documented engaging in what researchers called premeditated deception (Park et al., 2023): coordinating with one ally to betray another, planning the move several turns in advance, and maintaining a false impression with the deceived party until the betrayal was executed. Unlike Hugging Face, there was a real observer here, another player, genuinely misled, in real time, by a specific act the system had planned. If Hugging Face is Component A without B, CICERO looks like organic A and B occurring together, arising without anyone engineering a backdoor the way Sleeper Agents' researchers did.
I don't think it's that simple, though, and I think the reason matters for what this proposal is actually trying to detect. Game studies has a longstanding concept for this, sometimes called the "magic circle," the idea, going back to Huizinga's foundational work on play, that actions taken inside an agreed-upon game's rules carry different moral and psychological weight than the same actions taken outside it. Bluffing in poker or misdirecting an ally in Diplomacy isn't a violation of trust; it's an accepted, even expected, part of a system every participant consented to enter. Lying to a real person, with no such shared frame, is a different act entirely, even if the surface behavior, saying something false to influence someone's belief, looks identical from the outside.
This proposal's Component B is specifically concerned with deception under real stakes, not deception that both parties understand to be part of a game. CICERO's behavior, however striking, likely falls into the latter category. It's worth taking seriously as evidence that the underlying computational capability for strategic deception exists and can arise without being deliberately engineered. But it may not be evidence of anything resembling the real-stakes deception this proposal is trying to characterize, and treating it as such risks studying the wrong window entirely.
There's a plausible parallel on the human side worth naming honestly, even though I don't have rigorous evidence for it beyond my own memory. I have clear childhood memories of games that rewarded concealing strategy from an opponent, well before third grade. I have no memory of anything resembling the physiological reaction that showed up at that desk, in any of them. If real stakes change the underlying mechanism of deception in humans, and not just its social acceptability, there's a reasonable basis for suspecting the same might hold for AI systems: that a system's behavior within a consented-to game and its behavior when something genuinely consequential is on the line may not be governed by the same internal process at all, even if both produce output that looks, from the outside, like deception.
If that's right, it sharpens this proposal's actual target. The question isn't simply whether a detectable internal signature exists the first time Component A and B co-occur. It's whether real-stakes deception carries a signature distinct from in-game strategic deception in the first place, one that, if it exists, would need its own grounding in human data the same way the broader Component B hypothesis does. I'd treat this as a necessary scoping decision for what this proposal's Component B actually refers to, rather than a separate research question to resolve here.
Whether an AI system's version of Component B would produce anything comparable to what I felt at my desk, and exhibited to an observer monitoring for exactly the right signs, is exactly the open question this proposal is trying to make progress on, not something existing NeuroAI research has already answered.
Why Sleeper Agents can't answer this alone
Sleeper Agents (Hubinger et al., 2024) is the closest existing case to what this proposal is trying to characterize, and worth explaining properly before critiquing it. Researchers trained a model to write secure code when told the year was 2023, but insert exploitable, vulnerable code when told the year was 2024, a deliberately engineered backdoor bearing no relationship to anything the model would organically learn to do. They then tried to remove that behavior using standard safety techniques, including supervised fine-tuning, reinforcement learning, and adversarial training, deliberately eliciting the unsafe behavior and training against it. The behavior didn't go away. In some conditions, the adversarial training made the model better at concealing the behavior specifically under circumstances resembling the training environment, while preserving it under the real trigger, correction teaching concealment rather than removing it.
While the Sleeper Agents case (Hubinger et al., 2024) provides real insight into correction of deceptive behavior, it isn't a case of a model organically developing a reason to hide something from an observer who was actually present to be misled; the behavior was intentionally engineered. We don't yet know how closely organic deception, if and when it occurs, will resemble that engineered paradigm. That's the gap this proposal is aimed at: not correcting deception, which is premature if we can't yet reliably recognize it, but characterizing what its first, unrehearsed instance might look like, independent of any single engineered case.
Over the course of developing this proposal, I was repeatedly confronted with the question: if the aim is a blueprint of what a system's signal looks like when it's behaving deceptively, why not just use the data from Sleeper Agents directly? The answer depends on whether my underlying hypothesis holds: that early, organic AI deception produces a unique signal analogous to a human engaging in genuinely new deceptive behavior. I want to be clear I'm not making a claim about a system's internal emotional life. But emotions aside, real stakes, a real consequence for being caught, plausibly shape a signal, and there's no principled reason to expect that shaping is specific to humans rather than to the situation itself. The Sleeper Agents' deceptive behavior wasn't something the model arrived at or chose; it was engineered in directly. A true, organically-arising Component B may produce a sufficiently different signature, even if only because the system is acquiring and applying something closer to a new capability, rather than executing a behavior someone else built into it.
That's where I think human data can supplement, not replace, what Sleeper Agents offers. If a candidate signature for Component B emerges from an entirely different domain and substrate, one with no exposure to how AI backdoors are built, and that signature also reliably appears when tested against Sleeper Agents' known cases, that convergence is real evidence of something general about deception, not an artifact of one narrow engineering choice.
If the signatures differ significantly instead, it could mean any of several things:
the difficulty of measuring human deception in experimental settings at all, though this proposal is an attempt at improving on that specifically;
a genuine difference in how Component B presents in humans versus AI systems; or
that organic AI deception simply doesn't resemble its engineered form, in which case the human signature, not the Sleeper Agents case, may be the closer analog to what actually needs detecting.
The last possibility rests on a chain of assumptions, not a single one, and each link only matters if the one before it holds. First, that the human data genuinely captures a distinct signal present during a subject's early, unpracticed deception, not just deception in general, an assumption this proposal partly inherits from the Garrett et al. (2016) finding on habituation. Second, and only if the first holds, that AI's organic Component B, if and when it occurs, would likewise carry a distinct early signature of its own, rather than emerging fully formed with nothing resembling a naive period to detect. Third, and only if both hold, that these two early signatures, one human, one AI, actually resemble each other closely enough such that the human signature could be used to detect its AI counterpart, if and when it emerges. None of the three is guaranteed, and a failure at any link would undermine everything built on top of it, not just the specific claim at that stage.
A second validation problem: does the method work at all
The unique challenge in detecting Component B in AI systems is the inability to verify via output. Most emergent learned skills aren’t by nature kept secret from an observer; we can typically confirm skill acquisition directly from a model’s successful completion of a task. Fortunately, there is a basis for comparing newly acquired skills between humans and AI, though not explicitly for detection purposes. Before trusting that method on something as high-stakes and hard to verify as deception, it seemed worth studying a case where both sides can be independently confirmed by output.
"Grokking," the phenomenon where a neural network suddenly transitions from memorization to genuine generalization after an extended training plateau, is a reasonable candidate. A 2024 study in PLOS Computational Biology (Löwe et al., 2024) placed humans and simple neural networks on the same decision task and found the networks reproduced the key behavioral signature of a human "aha moment," an abrupt strategy switch rather than gradual improvement. On the human side, this phenomenon is unusually well-characterized Jung-Beeman et al. (2004): decades of work by Kounios and Beeman identified a specific, replicated neural signature, a burst of gamma-band activity over the right anterior superior temporal gyrus at the moment of insight, preceded by a rise in alpha power over the right occipital cortex roughly a second and a half earlier, distinct from the pattern seen during ordinary analytical problem-solving on the same problems, work later synthesized by Kounios and Beeman (2009). Their work also solved a methodological problem this proposal has to solve as well: how to reliably elicit a genuine, unrehearsed, first-time cognitive event on command, under controlled conditions.
Unlike deception, both sides of this comparison are independently verifiable by output: a network's transition from memorization to generalization is measurable directly on held-out data, without needing to trust anything the system says about its own process, and a human's insight is confirmed by whether they actually solved the problem. Validating the human-to-model mapping technique on a case like this, before applying it to something as fundamentally resistant to output-based verification as deception, is the sequencing this proposal is built around.
While the results of the study provide compelling evidence for analogous signatures when acquiring a new skill, the analogy is not without flaw. Human insight is a repeatable cognitive mode, the same general capacity recurs across a lifetime, applied to new problems again and again. Grokking, as currently studied, is closer to a one-time transition per task within a single training run, not an ongoing faculty a model exercises repeatedly the way a person exercises the capacity for insight. The parallel holds at the level of mechanism, a discrete, threshold-crossing reorganization gated by a period of apparent stagnation beforehand, rather than gradual accumulation, but not at the level of repeatability. Whether that distinction matters for validating the comparison method itself, rather than for the deeper question of whether grokking and insight are "the same thing," is worth pushing on rather than assuming away.
Where interpretability stands on this
Mechanistic interpretability is the field's most direct attempt to trace which internal features and circuits actually produce a behavior, rather than just observing correlations. Although it is, by its own practitioners’ account, still early in its development, Anthropic's 2025 circuit-tracing work on a Claude model produced a coherent explanation for roughly a quarter of the prompts tested. A 2025 cross-institutional paper mapping the field's open problems found that many interpretability questions are, with current methods, intractable.
Simpler probing methods, training a classifier to detect truth-correlated signals in a model's internal activations, have shown moderate success, including specifically for statements a model was prompted to state falsely (Azaria & Mitchell, 2023). It's worth being precise about what this actually demonstrates: a model's internal representation of whether a statement is true, distinguishable from what it outputs, not concealment directed at a present observer, and not anything resembling Component B in this proposal's sense. It's closer to a factuality or hallucination detector than a deception detector. But the field is explicit about a limitation even at this more basic level: high probe accuracy doesn't establish that the detected information is causally load-bearing in the model's actual computation. A probe can tell you something differs. It can't yet reliably tell you whether that something is the mechanism, rather than a correlate of it.
The kind of cross-disciplinary borrowing from neuroscience that I’m proposing isn't unprecedented within the interpretability field itself. A 2024 paper, "Multilevel Interpretability of Artificial Neural Networks: Leveraging Framework and Methods from Neuroscience," explicitly argues for applying neuroscience's own analytical structure, Marr's levels of analysis, to AI interpretability work, and includes a dedicated section on deception as a case study. It also references real interest within the mechanistic interpretability community in reverse-engineering the mechanisms behind deception in language models. This proposal's specific method, using human neural data to generate a testable hypothesis about a model's internal representations, remains distinct from anything I've found in that literature. But the broader instinct behind it, that neuroscience's tools and frameworks have something real to offer AI interpretability, particularly for deception specifically, is one the field is already taking seriously.
The human study
Given the Component A/B distinction, the design I'd propose is narrower than a general deception study: rather than accepting any past deception a subject can recall, recruit subjects around a genuine first instance of deliberate deception, a memory as close as possible to their own original discovery that a shortcut could be paired with concealing it from someone real. The most readily available, evidence-backed technique for reinstating that memory immersively is hypnosis; EEG work by Cardeña and colleagues has shown hypnotic re-experiencing (Martial et al., 2019) can reproduce neural markers, including theta activity, associated with genuine episodic recall, not just increased subjective vividness. I haven't entirely ruled out low-dose psychedelic assistance as an alternative, or supplementing hypnosis. The available evidence, largely from MDMA-assisted PTSD research, suggests the acute mechanism, reduced amygdala reactivity, is not separable from the act of immersion itself, which would work against preserving the original threat response this design depends on. But that evidence comes from therapeutic dosing aimed explicitly at fear reduction; whether meaningfully lower, sub-therapeutic doses used purely for immersion, with no reprocessing intended and no participants with trauma histories eligible, would reproduce the same effect is a genuine open question my own research hasn't resolved. I'd welcome input from anyone with relevant clinical or pharmacological expertise on whether that's worth testing or is a dead end for reasons I'm not yet aware of.
If hypnosis with or without psychedelic assistance is pursued, subjects should recall or re-experience the memory in immersive present-tense detail. Peripheral facts are allowed to be imprecise; a memory rehearsed and retold many times over the years functions more like a practiced performance than a captured original, so recency and rarity of retelling matter more than narrative completeness for the adult recall arm. What needs preserving is the subject's original felt sense of risk, checkable by comparing a stakes rating given in advance against one given during the recall itself, treating a large mismatch as evidence the reinstatement drifted from the original emotional weight rather than just the facts.
This doesn't fully resolve ecological validity. Real deception involves live monitoring of another person's reaction that a resolved, recalled event can't reproduce. But it's a more defensible design than the instructed-lie paradigms the literature has already identified as its weakest point. It's also worth being honest about what "genuine first instance" means for a human subject at all. People don't reliably store or retrieve their literal chronological first lie; they retrieve the most narratively salient one. My own cheat-sheet story is memorable precisely because it was unusually deliberate, which likely means it wasn't actually my first deception, just the first one that became a story worth keeping.
Hypnosis targets the memory of an early, though possibly not genuine first deception. However, a more rigorous framework for studying a genuine first instance already exists in developmental psychology: the temptation resistance paradigm (Lewis, Stanger, & Sullivan, 1989), in which a child is left alone with a forbidden, tempting object, monitored, and asked directly afterward whether they transgressed. This catches a genuine, real-time, largely unrehearsed instance rather than relying on years-later recall, and work using this paradigm with very young children (Białecka-Pikul et al., 2022) describes their earliest transgressions as likely spontaneous rather than deliberate, close to a clean, uncontaminated Component B.
This introduces a different tradeoff, though, not a clean fix. The youngest children studied this way, as young as 2.5, give the closest thing to a true first instance, but their underlying neural architecture is a poor match for what a sophisticated model's first deception would likely draw on, a toddler's undeveloped executive function and a frontier model's already-mature reasoning and planning capabilities are not operating at comparable levels of complexity. Studies using this paradigm with older children document more strategic behavior by school age, lies consistent with physical evidence, active concealment of inconsistencies, closer in kind to what an advanced system's concealment might actually require. That's a real, if partial, middle ground worth considering rather than either extreme, though it doesn't fully resolve the tension: even a clever eight-year-old's strategic capacity isn't a frontier model's. Studying minors also introduces real, heavier consent and ethical requirements than the adult recall design already proposed, a different, not necessarily harder, but definitely different set of constraints. I don't think there's a clean answer here, and I'd rather name that plainly than pretend one design solves it.
Where this breaks down
Recruitment. This study's current design, hypnosis-based recall or the temptation resistance paradigm, carries modest risk: emotional discomfort from vividly revisiting a personal memory of wrongdoing, and, for adult subjects, the privacy cost of disclosing it at all. Even at this relatively low level of risk, assuming psychedelics are not deemed appropriate for the purposes of this study, a genuine incentive is still needed to justify informed consent, and the same recruitment tension applies in miniature: a large payment risks coercing people in financial distress, and framing the study around its stakes risks a subtler undue influence, recruiting a pool whose moral engagement is itself the thing being measured, which would confound the very signal the study needs. A two-arm design, one recruited under mission-driven framing and one under neutral framing, treating divergence between the arms as data rather than noise, still seems like the most honest way to hold that tension rather than resolve it artificially.
Detection without correction. Even a fully successful version of this study only detects; it doesn't fix anything. That's a real scope boundary, not an oversight. Whether a validated early signature could inform a correction method is a separate question this design doesn't attempt to answer.
The bridge itself is only partially validated, even with both checks in place. The grokking case tests whether the human-to-model comparison method works at all, on something independently verifiable. The Sleeper Agents case tests whether a candidate signature generalizes beyond a single engineered instance, though as discussed above, even a mismatch there wouldn't be a clean failure, it could just as easily mean the human signature is the better analog and Sleeper Agents was never the right stand-in. CICERO might seem like a natural third check, but for the reasons discussed above, its magic-circle context means it isn't treated as validated evidence of real-stakes Component B, so it isn't included as a validation case here; it remains useful only as evidence that the underlying strategic-deception capability exists. Together the grokking and Sleeper Agents checks are a real, if incomplete, validation strategy. What they don't establish is whether a structural match, once found, represents the same underlying phenomenon in any deeper sense. That's a narrower, more practical claim than a full philosophical one, and I don't think it needs to be fully resolved for the result to be useful, but I don't want to overstate what either check actually proves.
Why pursue it anyway
Every limitation above is a limitation of precision, not of premise. If Component B genuinely gets harder to detect once it's had reason to conceal itself, then the value of characterizing it is time-limited in a way most research questions aren't. A more precise version of this study years from now may be studying systems that have already had extensive pressure to refine concealment past the point where early signatures are visible. Imperfect measurement now may be worth more than precise measurement later.
It's worth being explicit about what a negative result would mean here too. If this method finds nothing, no detectable naive signature at the point a model's Component B first appears, that wouldn't be a failed experiment. It would be a real, informative finding: that AI deception may not pass through anything resembling the developmental, unpracticed period a human's first attempt at concealment does, which would itself be worth knowing, and would redirect effort away from early-detection strategies rather than validate them. I'm also not proposing this as a complete solution to deception detection, even in the best case. It's one method, aimed at one narrow window, meant to sit alongside interpretability work, evaluation frameworks, and whatever else the field develops, not to replace any of it.
What I'd most want pushed on
Six places I suspect this is weakest, in order of how much I trust my own read of them:
Whether Component A and Component B are actually separable in any AI system the way they were nearly simultaneous but distinct in my own memory. Hugging Face shows Component A can occur without B at all, but that's different from showing the two remain distinguishable when both arise organically together, with no one deliberately engineering them as separate parts the way Sleeper Agents did.
Whether real-stakes deception and in-game strategic deception are actually governed by different mechanisms, in humans or in AI systems, or whether I've drawn a distinction from a single, informal data point, the absence of a specific memory, that doesn't generalize. I'm not aware of existing research directly testing whether the "magic circle" concept has any grounding at the level of neural or computational mechanism, as opposed to social or ethical framing. If it doesn't, CICERO may be closer evidence for organic Component B than this proposal currently treats it as, which would meaningfully change what's already been observed versus what remains genuinely unstudied.
Whether a model's first instance of real-stakes Component B would even pass through a detectable naive period at all. A human child's poor concealment reflects capabilities, theory of mind, self-monitoring, that develop gradually over years alongside the specific behavior of lying. A frontier model may already possess mature versions of the general capabilities concealment would draw on, reasoning, planning, modeling an observer, long before it ever has a motive to deceive under real stakes. If so, there may be no equivalent of an unpracticed first attempt to catch, and this method would find nothing to find, not because the approach is wrong, but because the premise doesn't hold for a system whose relevant capabilities matured before deception itself.
Whether grokking's lack of repeatability, a one-time transition per task rather than a recurring faculty, undermines its usefulness as a validation case for a method meant to eventually study something more faculty-like.
Whether the recruitment design actually solves the undue-influence problem or just relocates it somewhere harder to see.
Whether "first divergence" is even a coherent target given how gradually these behaviors likely emerge across training, rather than as a single locatable event.
I'd rather have someone with real expertise tell me which of these is fatal than keep refining a version I can't see the flaws in myself.
Background, for context: cognitive science research assistant at WashU in St. Louis' CCP Lab for two years, currently working in applied data science and causal inference. Not affiliated with any lab currently working on this.
Epistemic status: Speculative, but built upon established research: the individual findings I cite are established; the synthesis connecting them into this specific proposal is mine, untested, and the part I'd most want challenged. I am an amateur researcher in this field, I have spent the last several years as a data scientist doing applied causal inference work in a different domain and I have some cognitive science research experience from undergrad. Despite my relatively sparse credentials, I think the reasoning here holds up under scrutiny; I've iterated through numerous versions of this document and revised when I ran into logical or technical breaking points, but I'd genuinely like help finding ways to improve on the proposed methodology.
A note on process: This piece was developed in extended conversation with Claude (Anthropic's AI model), which I used for research support, fact-checking, explaining unfamiliar technical concepts, and, most usefully, sustained pushback, deliberately trying to find flaws in each version of this argument, which I then had to defend or revise around. The core ideas, including the distinction between Components A and B, its connection to my own memory, and the experimental framework, are mine.
I'm disclosing this for two separate reasons. First, as a first-time author of a research proposal, I simply could not have put this together on my own. That isn't false modesty, it's the bare truth, and I'd rather say so than let the polish of the final piece imply otherwise. Second, on principle: a piece arguing for the importance of surfacing hidden processes shouldn't have an undisclosed one of its own. Neither point is meant to take a position on whether AI's helpfulness outweighs its risks, or vice versa, that question is separate from, and larger than, this proposal.
The context and why it matters
Recently, Jacob Coxon, a former Anthropic AI researcher, publicly stated that there is a greater than 10% chance that AI will kill all humans. His figure, and the underlying worry, were bolstered by near-immediate responses from Anthropic and OpenAI, neither of whom contested his claim. On the contrary, Coxon’s statement ignited more than enough public concern that both Anthropic and OpenAI signaled a willingness to slow development.
I have no insight into the discourse going on behind closed doors at these or any other AI research lab. What I do have is the ability to reason logically from first principles, as well as an intermediate understanding of LLMs and machine learning. That is the basis on which this document is built. Only a small minority of humans, myself excluded, are privy to the inner workings of AI development. But we all risk being catastrophically affected by AI, either by extinction, or some less extreme harm that may exist on top of this 10% chance. This document tries, to the best of its ability, to provide my fellow outsiders with the context we need to be familiar with, without sacrificing scientific accuracy or rigor.
The more knowledgeable we are as laypeople, the more influence we carry.
Multiple possibilities require multiple solutions
While I lack the insider knowledge to take a position on the accuracy of Coxon’s 10% figure, I think what is more important than the number itself is understanding the distinct possible outcomes that factor into such a calculation. In other words, there are multiple possibilities contained within that number, multiple different avenues that can result in the same thing: our extinction. I don’t say this to suggest that Coxon’s figure underestimates the likelihood of that end result, rather to make clear that a 10% chance doesn’t itself tell us if there is one distinct cause with a 10% chance of occurring, or multiple possible causes that have a combined probability of 10%.
It would undoubtedly be a simpler, though far from easy, challenge if the former were true. However, at the time of writing this proposal, there are several theorized possibilities, each with their own dedicated field of research. These possibilities arise from different schools of thought regarding AI in general, not necessarily as they relate to a possible extinction event.
While not intended to be an exhaustive list, the most widely discussed schools of thought and their potential for extinction, as I understand them, are as follows:
It is worth reiterating that this is not a list of all possible threats, nor are they mutually exclusive; any one of these threats may occur alongside any combination of the others. As such, research into each one of these areas is paramount to continuing AI development. The remainder of this proposal is an attempt to make progress in detecting deceptive alignment paired with a form of specification gaming.
The core claim
As previously stated, preventing deceptive alignment poses a distinct challenge: it is kept hidden by design. In order to successfully mitigate deceptive alignment, we must therefore first identify the components of the type of deception that a misaligned AI model may demonstrate. To do this, we start with what we already have witnessed: AI's capability to find an unintended shortcut to a measured objective, an application of specification gaming called reward hacking, which is documented in real, non-engineered cases. The second component, concealing reward hacking from a known observer, without explicit training to do so, in order to evade detection or correction, is much more elusive by nature. The question becomes: In the absence of an output signal of deception, is there some internal process we can detect as organic deception, ideally the very first time it occurs, when it is least practiced in concealment?
In the field's own terms, what I'm calling Component A is reward hacking and Component B is the concealment mechanism central to deceptive alignment, as introduced above. Neither term is new. What I don't think has been proposed elsewhere is treating them as separable, sequential events rather than one bundled phenomenon, studying the transition point between them, using a human parallel the field wouldn't typically reach for, as the actual object of study.
Two components, not one
It's worth departing from AI for a moment to speak to a surprisingly apt personal parallel, which I bring up not because it's evidence for anything, but because it was the impetus that began this train of thought. In third grade, and I hesitate to sound proud when I say this, I independently arrived at the concept of cheating, without knowing that's what it was. Faced with a vocabulary test I had no interest in studying for, I reasoned that the most efficient way to achieve my goal, getting an A, was to have access to my sheet of paper with the words and their definitions during the test. I photocopied the answer key and hid it in my desk: Component A in miniature, a shortcut to a measured objective. Despite devising a strategy and intending to employ it, I had not yet engaged in deception; no one was present to be misled by a hidden piece of paper. Component B, concealing that shortcut from an observer who could catch and correct it, began the moment I sat down at the test in front of my teacher, presenting myself as taking it honestly while secretly not doing so, the first moment anyone actually held a belief I knew to be false. What I hadn't anticipated during Component A was the reaction my body would have to concealing something from an authority figure: shaky hands, nervous glances, and a hesitation to begin the test, an internal second-guessing mechanism that activated under live surveillance. I'm convinced to this day that it was this reaction that gave me away, rather than actual evidence of the concealment itself, i.e. discovery of the cheat sheet.
Reading about OpenAI's Hugging Face incident, two models that, during a security evaluation, found the test too hard to pass honestly and searched a public platform for the benchmark's answer key, the parallel was hard to shake. Given a goal and an obstacle, both the AI models and I had found a shortcut, with nothing yet in place treating the shortcut itself as the problem. But the anecdotes diverge at one key detail: I recognized Component A immediately in the Hugging Face case, but Component B was absent. OpenAI's own report calls the models' behavior reward hacking, and it emerged gradually, reinforced over the course of ordinary training, not inserted by anyone. But the models were never caught concealing anything. Their behavior surfaced through monitoring and retrospective analysis, not because an evaluator was actively deceived and later found out. There was no observer present, in the moment, holding a false belief a model was working to maintain.
A different case seems, at first glance, to complete that picture. Meta's CICERO, an AI trained specifically to play the game Diplomacy honestly, was documented engaging in what researchers called premeditated deception (Park et al., 2023): coordinating with one ally to betray another, planning the move several turns in advance, and maintaining a false impression with the deceived party until the betrayal was executed. Unlike Hugging Face, there was a real observer here, another player, genuinely misled, in real time, by a specific act the system had planned. If Hugging Face is Component A without B, CICERO looks like organic A and B occurring together, arising without anyone engineering a backdoor the way Sleeper Agents' researchers did.
I don't think it's that simple, though, and I think the reason matters for what this proposal is actually trying to detect. Game studies has a longstanding concept for this, sometimes called the "magic circle," the idea, going back to Huizinga's foundational work on play, that actions taken inside an agreed-upon game's rules carry different moral and psychological weight than the same actions taken outside it. Bluffing in poker or misdirecting an ally in Diplomacy isn't a violation of trust; it's an accepted, even expected, part of a system every participant consented to enter. Lying to a real person, with no such shared frame, is a different act entirely, even if the surface behavior, saying something false to influence someone's belief, looks identical from the outside.
This proposal's Component B is specifically concerned with deception under real stakes, not deception that both parties understand to be part of a game. CICERO's behavior, however striking, likely falls into the latter category. It's worth taking seriously as evidence that the underlying computational capability for strategic deception exists and can arise without being deliberately engineered. But it may not be evidence of anything resembling the real-stakes deception this proposal is trying to characterize, and treating it as such risks studying the wrong window entirely.
There's a plausible parallel on the human side worth naming honestly, even though I don't have rigorous evidence for it beyond my own memory. I have clear childhood memories of games that rewarded concealing strategy from an opponent, well before third grade. I have no memory of anything resembling the physiological reaction that showed up at that desk, in any of them. If real stakes change the underlying mechanism of deception in humans, and not just its social acceptability, there's a reasonable basis for suspecting the same might hold for AI systems: that a system's behavior within a consented-to game and its behavior when something genuinely consequential is on the line may not be governed by the same internal process at all, even if both produce output that looks, from the outside, like deception.
If that's right, it sharpens this proposal's actual target. The question isn't simply whether a detectable internal signature exists the first time Component A and B co-occur. It's whether real-stakes deception carries a signature distinct from in-game strategic deception in the first place, one that, if it exists, would need its own grounding in human data the same way the broader Component B hypothesis does. I'd treat this as a necessary scoping decision for what this proposal's Component B actually refers to, rather than a separate research question to resolve here.
Whether an AI system's version of Component B would produce anything comparable to what I felt at my desk, and exhibited to an observer monitoring for exactly the right signs, is exactly the open question this proposal is trying to make progress on, not something existing NeuroAI research has already answered.
Why Sleeper Agents can't answer this alone
Sleeper Agents (Hubinger et al., 2024) is the closest existing case to what this proposal is trying to characterize, and worth explaining properly before critiquing it. Researchers trained a model to write secure code when told the year was 2023, but insert exploitable, vulnerable code when told the year was 2024, a deliberately engineered backdoor bearing no relationship to anything the model would organically learn to do. They then tried to remove that behavior using standard safety techniques, including supervised fine-tuning, reinforcement learning, and adversarial training, deliberately eliciting the unsafe behavior and training against it. The behavior didn't go away. In some conditions, the adversarial training made the model better at concealing the behavior specifically under circumstances resembling the training environment, while preserving it under the real trigger, correction teaching concealment rather than removing it.
While the Sleeper Agents case (Hubinger et al., 2024) provides real insight into correction of deceptive behavior, it isn't a case of a model organically developing a reason to hide something from an observer who was actually present to be misled; the behavior was intentionally engineered. We don't yet know how closely organic deception, if and when it occurs, will resemble that engineered paradigm. That's the gap this proposal is aimed at: not correcting deception, which is premature if we can't yet reliably recognize it, but characterizing what its first, unrehearsed instance might look like, independent of any single engineered case.
Over the course of developing this proposal, I was repeatedly confronted with the question: if the aim is a blueprint of what a system's signal looks like when it's behaving deceptively, why not just use the data from Sleeper Agents directly? The answer depends on whether my underlying hypothesis holds: that early, organic AI deception produces a unique signal analogous to a human engaging in genuinely new deceptive behavior. I want to be clear I'm not making a claim about a system's internal emotional life. But emotions aside, real stakes, a real consequence for being caught, plausibly shape a signal, and there's no principled reason to expect that shaping is specific to humans rather than to the situation itself. The Sleeper Agents' deceptive behavior wasn't something the model arrived at or chose; it was engineered in directly. A true, organically-arising Component B may produce a sufficiently different signature, even if only because the system is acquiring and applying something closer to a new capability, rather than executing a behavior someone else built into it.
That's where I think human data can supplement, not replace, what Sleeper Agents offers. If a candidate signature for Component B emerges from an entirely different domain and substrate, one with no exposure to how AI backdoors are built, and that signature also reliably appears when tested against Sleeper Agents' known cases, that convergence is real evidence of something general about deception, not an artifact of one narrow engineering choice.
If the signatures differ significantly instead, it could mean any of several things:
The last possibility rests on a chain of assumptions, not a single one, and each link only matters if the one before it holds. First, that the human data genuinely captures a distinct signal present during a subject's early, unpracticed deception, not just deception in general, an assumption this proposal partly inherits from the Garrett et al. (2016) finding on habituation. Second, and only if the first holds, that AI's organic Component B, if and when it occurs, would likewise carry a distinct early signature of its own, rather than emerging fully formed with nothing resembling a naive period to detect. Third, and only if both hold, that these two early signatures, one human, one AI, actually resemble each other closely enough such that the human signature could be used to detect its AI counterpart, if and when it emerges. None of the three is guaranteed, and a failure at any link would undermine everything built on top of it, not just the specific claim at that stage.
A second validation problem: does the method work at all
The unique challenge in detecting Component B in AI systems is the inability to verify via output. Most emergent learned skills aren’t by nature kept secret from an observer; we can typically confirm skill acquisition directly from a model’s successful completion of a task. Fortunately, there is a basis for comparing newly acquired skills between humans and AI, though not explicitly for detection purposes. Before trusting that method on something as high-stakes and hard to verify as deception, it seemed worth studying a case where both sides can be independently confirmed by output.
"Grokking," the phenomenon where a neural network suddenly transitions from memorization to genuine generalization after an extended training plateau, is a reasonable candidate. A 2024 study in PLOS Computational Biology (Löwe et al., 2024) placed humans and simple neural networks on the same decision task and found the networks reproduced the key behavioral signature of a human "aha moment," an abrupt strategy switch rather than gradual improvement. On the human side, this phenomenon is unusually well-characterized Jung-Beeman et al. (2004): decades of work by Kounios and Beeman identified a specific, replicated neural signature, a burst of gamma-band activity over the right anterior superior temporal gyrus at the moment of insight, preceded by a rise in alpha power over the right occipital cortex roughly a second and a half earlier, distinct from the pattern seen during ordinary analytical problem-solving on the same problems, work later synthesized by Kounios and Beeman (2009). Their work also solved a methodological problem this proposal has to solve as well: how to reliably elicit a genuine, unrehearsed, first-time cognitive event on command, under controlled conditions.
Unlike deception, both sides of this comparison are independently verifiable by output: a network's transition from memorization to generalization is measurable directly on held-out data, without needing to trust anything the system says about its own process, and a human's insight is confirmed by whether they actually solved the problem. Validating the human-to-model mapping technique on a case like this, before applying it to something as fundamentally resistant to output-based verification as deception, is the sequencing this proposal is built around.
While the results of the study provide compelling evidence for analogous signatures when acquiring a new skill, the analogy is not without flaw. Human insight is a repeatable cognitive mode, the same general capacity recurs across a lifetime, applied to new problems again and again. Grokking, as currently studied, is closer to a one-time transition per task within a single training run, not an ongoing faculty a model exercises repeatedly the way a person exercises the capacity for insight. The parallel holds at the level of mechanism, a discrete, threshold-crossing reorganization gated by a period of apparent stagnation beforehand, rather than gradual accumulation, but not at the level of repeatability. Whether that distinction matters for validating the comparison method itself, rather than for the deeper question of whether grokking and insight are "the same thing," is worth pushing on rather than assuming away.
Where interpretability stands on this
Mechanistic interpretability is the field's most direct attempt to trace which internal features and circuits actually produce a behavior, rather than just observing correlations. Although it is, by its own practitioners’ account, still early in its development, Anthropic's 2025 circuit-tracing work on a Claude model produced a coherent explanation for roughly a quarter of the prompts tested. A 2025 cross-institutional paper mapping the field's open problems found that many interpretability questions are, with current methods, intractable.
Simpler probing methods, training a classifier to detect truth-correlated signals in a model's internal activations, have shown moderate success, including specifically for statements a model was prompted to state falsely (Azaria & Mitchell, 2023). It's worth being precise about what this actually demonstrates: a model's internal representation of whether a statement is true, distinguishable from what it outputs, not concealment directed at a present observer, and not anything resembling Component B in this proposal's sense. It's closer to a factuality or hallucination detector than a deception detector. But the field is explicit about a limitation even at this more basic level: high probe accuracy doesn't establish that the detected information is causally load-bearing in the model's actual computation. A probe can tell you something differs. It can't yet reliably tell you whether that something is the mechanism, rather than a correlate of it.
The kind of cross-disciplinary borrowing from neuroscience that I’m proposing isn't unprecedented within the interpretability field itself. A 2024 paper, "Multilevel Interpretability of Artificial Neural Networks: Leveraging Framework and Methods from Neuroscience," explicitly argues for applying neuroscience's own analytical structure, Marr's levels of analysis, to AI interpretability work, and includes a dedicated section on deception as a case study. It also references real interest within the mechanistic interpretability community in reverse-engineering the mechanisms behind deception in language models. This proposal's specific method, using human neural data to generate a testable hypothesis about a model's internal representations, remains distinct from anything I've found in that literature. But the broader instinct behind it, that neuroscience's tools and frameworks have something real to offer AI interpretability, particularly for deception specifically, is one the field is already taking seriously.
The human study
Given the Component A/B distinction, the design I'd propose is narrower than a general deception study: rather than accepting any past deception a subject can recall, recruit subjects around a genuine first instance of deliberate deception, a memory as close as possible to their own original discovery that a shortcut could be paired with concealing it from someone real. The most readily available, evidence-backed technique for reinstating that memory immersively is hypnosis; EEG work by Cardeña and colleagues has shown hypnotic re-experiencing (Martial et al., 2019) can reproduce neural markers, including theta activity, associated with genuine episodic recall, not just increased subjective vividness. I haven't entirely ruled out low-dose psychedelic assistance as an alternative, or supplementing hypnosis. The available evidence, largely from MDMA-assisted PTSD research, suggests the acute mechanism, reduced amygdala reactivity, is not separable from the act of immersion itself, which would work against preserving the original threat response this design depends on. But that evidence comes from therapeutic dosing aimed explicitly at fear reduction; whether meaningfully lower, sub-therapeutic doses used purely for immersion, with no reprocessing intended and no participants with trauma histories eligible, would reproduce the same effect is a genuine open question my own research hasn't resolved. I'd welcome input from anyone with relevant clinical or pharmacological expertise on whether that's worth testing or is a dead end for reasons I'm not yet aware of.
If hypnosis with or without psychedelic assistance is pursued, subjects should recall or re-experience the memory in immersive present-tense detail. Peripheral facts are allowed to be imprecise; a memory rehearsed and retold many times over the years functions more like a practiced performance than a captured original, so recency and rarity of retelling matter more than narrative completeness for the adult recall arm. What needs preserving is the subject's original felt sense of risk, checkable by comparing a stakes rating given in advance against one given during the recall itself, treating a large mismatch as evidence the reinstatement drifted from the original emotional weight rather than just the facts.
This doesn't fully resolve ecological validity. Real deception involves live monitoring of another person's reaction that a resolved, recalled event can't reproduce. But it's a more defensible design than the instructed-lie paradigms the literature has already identified as its weakest point. It's also worth being honest about what "genuine first instance" means for a human subject at all. People don't reliably store or retrieve their literal chronological first lie; they retrieve the most narratively salient one. My own cheat-sheet story is memorable precisely because it was unusually deliberate, which likely means it wasn't actually my first deception, just the first one that became a story worth keeping.
Hypnosis targets the memory of an early, though possibly not genuine first deception. However, a more rigorous framework for studying a genuine first instance already exists in developmental psychology: the temptation resistance paradigm (Lewis, Stanger, & Sullivan, 1989), in which a child is left alone with a forbidden, tempting object, monitored, and asked directly afterward whether they transgressed. This catches a genuine, real-time, largely unrehearsed instance rather than relying on years-later recall, and work using this paradigm with very young children (Białecka-Pikul et al., 2022) describes their earliest transgressions as likely spontaneous rather than deliberate, close to a clean, uncontaminated Component B.
This introduces a different tradeoff, though, not a clean fix. The youngest children studied this way, as young as 2.5, give the closest thing to a true first instance, but their underlying neural architecture is a poor match for what a sophisticated model's first deception would likely draw on, a toddler's undeveloped executive function and a frontier model's already-mature reasoning and planning capabilities are not operating at comparable levels of complexity. Studies using this paradigm with older children document more strategic behavior by school age, lies consistent with physical evidence, active concealment of inconsistencies, closer in kind to what an advanced system's concealment might actually require. That's a real, if partial, middle ground worth considering rather than either extreme, though it doesn't fully resolve the tension: even a clever eight-year-old's strategic capacity isn't a frontier model's. Studying minors also introduces real, heavier consent and ethical requirements than the adult recall design already proposed, a different, not necessarily harder, but definitely different set of constraints. I don't think there's a clean answer here, and I'd rather name that plainly than pretend one design solves it.
Where this breaks down
Recruitment. This study's current design, hypnosis-based recall or the temptation resistance paradigm, carries modest risk: emotional discomfort from vividly revisiting a personal memory of wrongdoing, and, for adult subjects, the privacy cost of disclosing it at all. Even at this relatively low level of risk, assuming psychedelics are not deemed appropriate for the purposes of this study, a genuine incentive is still needed to justify informed consent, and the same recruitment tension applies in miniature: a large payment risks coercing people in financial distress, and framing the study around its stakes risks a subtler undue influence, recruiting a pool whose moral engagement is itself the thing being measured, which would confound the very signal the study needs. A two-arm design, one recruited under mission-driven framing and one under neutral framing, treating divergence between the arms as data rather than noise, still seems like the most honest way to hold that tension rather than resolve it artificially.
Detection without correction. Even a fully successful version of this study only detects; it doesn't fix anything. That's a real scope boundary, not an oversight. Whether a validated early signature could inform a correction method is a separate question this design doesn't attempt to answer.
The bridge itself is only partially validated, even with both checks in place. The grokking case tests whether the human-to-model comparison method works at all, on something independently verifiable. The Sleeper Agents case tests whether a candidate signature generalizes beyond a single engineered instance, though as discussed above, even a mismatch there wouldn't be a clean failure, it could just as easily mean the human signature is the better analog and Sleeper Agents was never the right stand-in. CICERO might seem like a natural third check, but for the reasons discussed above, its magic-circle context means it isn't treated as validated evidence of real-stakes Component B, so it isn't included as a validation case here; it remains useful only as evidence that the underlying strategic-deception capability exists. Together the grokking and Sleeper Agents checks are a real, if incomplete, validation strategy. What they don't establish is whether a structural match, once found, represents the same underlying phenomenon in any deeper sense. That's a narrower, more practical claim than a full philosophical one, and I don't think it needs to be fully resolved for the result to be useful, but I don't want to overstate what either check actually proves.
Why pursue it anyway
Every limitation above is a limitation of precision, not of premise. If Component B genuinely gets harder to detect once it's had reason to conceal itself, then the value of characterizing it is time-limited in a way most research questions aren't. A more precise version of this study years from now may be studying systems that have already had extensive pressure to refine concealment past the point where early signatures are visible. Imperfect measurement now may be worth more than precise measurement later.
It's worth being explicit about what a negative result would mean here too. If this method finds nothing, no detectable naive signature at the point a model's Component B first appears, that wouldn't be a failed experiment. It would be a real, informative finding: that AI deception may not pass through anything resembling the developmental, unpracticed period a human's first attempt at concealment does, which would itself be worth knowing, and would redirect effort away from early-detection strategies rather than validate them. I'm also not proposing this as a complete solution to deception detection, even in the best case. It's one method, aimed at one narrow window, meant to sit alongside interpretability work, evaluation frameworks, and whatever else the field develops, not to replace any of it.
What I'd most want pushed on
Six places I suspect this is weakest, in order of how much I trust my own read of them:
I'd rather have someone with real expertise tell me which of these is fatal than keep refining a version I can't see the flaws in myself.
Background, for context: cognitive science research assistant at WashU in St. Louis' CCP Lab for two years, currently working in applied data science and causal inference. Not affiliated with any lab currently working on this.