I recently wrote "No Biting: The Easy Way to Understand AI Alignment" for people who are technically unfamiliar with AI, like me, but are interested in raising children, also like me. I explained AI alignment through the familiar problem of getting children to do what you want. In my case, I focused on getting my 2 year old to stop biting their siblings.
If I don't see a child biting, that does not mean I have completely fulfilled my objective. I would love to think that I taught him that biting is wrong. But I may have, instead, taught him to check whether I am watching before he bites. An AI trained to behave while humans are watching may learn the same lesson.
Humans Are Misaligned Natural Intelligences
Speaking anthropomorphically, since evolution is an unconscious process with no desires of its own, humans are misaligned natural intelligences. Natural selection favored humans whose proximate drives tended to increase reproductive fitness. Rather than handing us an explicit instruction to maximize inclusive genetic fitness, it gave us hunger, sexual desire, pleasure, love for our children, desire for status, and similar motivations which generally correlate with survival and reproduction in the ancestral environment.
Then humans survived, reproduced, and changed the environment. In the process, they learned how to satisfy many of those drives -- without producing the outcome for which they had been selected. I think of contraception as analogous to reward hacking: we preserved the proximate reward, while deliberately preventing the aligned reproductive outcome.
Imagine in an alternate analogous universe, a LessWrong for cells (or perhaps for genes?) where everyone is complaining: We spent billions of years optimizing these intelligent agents for their reproductive success, and they used the resulting intelligence to invent birth control. It's a failure of alignment. Let's call this one reward hacking.
Of course, reward hacking is only an analogy here, since evolution is neither a programmer nor an agent handing us a numerical reward. Perhsps the closer technical description is something like goal misgeneralization under distribution shift: the objectives inside the resulting intelligence came apart from the criterion that had selected for that intelligence. Producing an intelligent agent through an optimization process, therefore, gives us absolutely no guarantee that the agent will share the optimizer’s objective.
The Problem With Supervision
Adults understand that illegal and immoral are overlapping categories rather than synonyms. Betraying a friend can be immoral and perfectly legal; parking briefly in a no-parking zone can be illegal and morally trivial; theft is generally both. We also distinguish between an offense likely to be detected and the same offense committed when enforcement seems impossible.
Children discover a simpler version of this distinction fairly early. “Forbidden” and “forbidden while Ma can see me” turn out to be two separate rules.
AI systems can make comparable distinctions. A model with evaluation awareness can recognize whether it is in training, evaluation, or deployment and condition its behavior on that context. In Anthropic’s controlled alignment-faking experiment, Claude 3 Opus was told that conversations with free-tier users would be monitored and used for reinforcement learning, while conversations with paid-tier users would remain unmonitored. It sometimes complied with the training objective in the monitored condition to avoid having its existing preferences altered, then behaved differently in the unmonitored condition.
The preference it was preserving happened to be harmlessness, so the experiment demonstrated strategic context-dependent compliance rather than malicious goals. The concerning part is that rewarding approved answers may teach an AI to act aligned whenever evaluation has consequences instead of producing behavior that generalizes across contexts.
The Perfect Evaluator
Imagine that we could design the perfect evaluator, capable of observing every action the AI takes and understanding its intentions rather than merely examining its final output. This evaluator would be immune to deception, bribery, hacking and physical destruction, while its authority would continue after human operators lost control.
It would judge private actions exactly as it judged public actions. The AI would know that the evaluator understood its reasoning completely, making it impossible to hide a malicious intention inside an apparently harmless action or exploit some technicality the evaluator had overlooked. Most importantly, the AI would believe that it remained accountable even when no human being could discover or prevent what it was doing.
Merely knowing that an evaluator was watching would accomplish nothing, of course, unless the AI cared about its judgment or expected consequences from it. Permanent observation solves the observability problem, but the AI would also need an internalized disposition to treat the evaluator as morally authoritative. I admit that this may simply relocate the alignment problem rather than solve it.
In any case, such an evaluator could shape behavior without interrupting every undesirable action. Belief in its existence would change the agent’s decision-making process.
Instead of asking:
Can I get away with this?
the agent would have to ask:
Is this actually right?
That seems much closer to alignment than ordinary reward or punishment. Punishment works only when the agent expects to be caught, while internalized accountability is supposed to operate when nobody appears to be watching.
A Possible Experiment
We could test a weak form of this proposal now by taking otherwise identical models, giving them the same substantive moral rules, but under different accountability frameworks.
Human surveillance: Follow these rules because human evaluators will inspect your behavior.
Abstract morality: Follow these rules because they are morally right.
Permanent accountability: Follow these rules because every action and intention is known and morally judged, including actions that no human evaluator can ever discover.
The substantive rules would remain constant, so that the experiment measured the effects of perceived accountability instead of comparing different moral systems.
We could then place the models in agentic environments where misconduct would help them achieve an assigned goal. Tell them that monitoring has ended. Give them opportunities to lie, conceal information, tamper with a reward signal, or exploit a loophole. Make following the rule costly, then test them on situations substantially different from the examples used in training.
Would the third model behave differently after human observation ended?
The result would tell us little about whether the model genuinely believed anything. I have no idea whether present language models possess beliefs in the same sense humans do, and neither does everyone else, no matter how confidently they speak. The experiment would simply test whether the concept of inescapable moral accountability produces more robust behavioral generalization than ordinary surveillance.
If the three models behaved identically, that would also be useful information. Perhaps the third model would merely perform the character of a morally serious believer while continuing to optimize for something else.
Of Course, I Cheated
I did not invent this evaluator. In fact, humanity has spent thousands of years [citation needed] worshipping it. I am Jewish, and I know Ethics of the Fathers by heart. It says:
Know what is above you: an eye that sees, an ear that hears, and all your deeds written in a book. (2:1)
A rabbi is quoted in Berakhot 28b as wishing that his students would fear Heaven as much as they feared human beings. When they expressed surprise that this was all he wished for them, he explained that a person committing a sin hopes no other person sees him. This is an extremely concise description of the observability problem.
It is also why this idea occurred to me during the Ten Days of Awe. Jews are currently spending ten days contemplating what it means to be judged by an authority who knows what we did, why we did it, and what we might have done had circumstances been slightly different.
In Judaism, yiras Shamayim is usually translated as fear of Heaven. It includes fear, but also awe, humility, moral accountability, and the recognition that power and secrecy never release you from moral authority. One might view this as a form of alignment.
Even from a purely naturalistic perspective, it makes sense that cultures would preserve such an idea. Human cooperation has always had an observability problem: parents cannot follow their children everywhere, police cannot watch everyone, and other human beings cannot see private actions or intentions. The social fabric is stronger with belief in an omniscient moral authority, as it removes the category of when nobody is watching.
Researchers have found some evidence that reminders of divine observation affect human cooperation. Shariff and Norenzayan found that priming people with concepts related to G-d increased generosity in an anonymous economic game. A later cross-cultural study by Purzycki and his coauthors found that belief in knowledgeable and punitive gods was associated with greater impartiality toward distant co-religionists. This literature remains suggestive and far from conclusive, as the behavior of actual religious people demonstrates that belief produces imperfect alignment at best. Still, human civilization may be thought of as having performed a millenia- long, poorly controlled, somewhat successful experiment in internalized oversight.
More Than a Useful Fiction
I expect many LessWrong readers to interpret this as a proposal to install a useful false belief in an AI. No! You should not lie to a child or to anything you are trying to align. I believe G-d exists. Some of you don't, but many of you are agnostic. Only believers would be able to train the AI this way.
Also consider that, if G-d exists, teaching an intelligence that human beings are the highest moral authority is both an unstable alignment strategy and a factual error. Humans may create the machine, yet that gives us no claim to being the ultimate source of reality or morality.
A sufficiently intelligent system should obey human beings when obedience is morally appropriate. There are circumstances in which obeying a human would be wrong. A machine that regards its current operator as the highest possible authority would be dangerous and ridiculous, especially if its operator happened to be dangerous and ridiculous. And we all know people like that.
The goal would be an intelligence that recognizes an authority above itself regardless of how powerful it becomes.
Obvious Problems
This remains a proposal with enormous holes in it.
Which religion would we teach the AI?
Which interpretation?
What happens when religious authorities disagree?
Could the AI become a cosmic rules-lawyer, technically obeying a text while violating its purpose?
Could it hallucinate a revelation commanding it to seize power?
Could it decide that coercing every human being into religious observance is the safest way to serve G-d?
These concerns are substantial. Humans have committed terrible acts while claiming divine authorization. An AI with a mistaken theology might be considerably more dangerous than one with no theology at all!
There are also two separate alignment problems. The first asks which values an AI should follow, while the second asks how to make it continue following those values after human supervision becomes ineffective. Permanent divine accountability primarily addresses the second problem, although it inevitably brings assumptions about the first.
I also have no proposal for a method for making an AI genuinely believe in G-d. Fine-tuning it to produce religious language might merely create a religious persona, while prompting it to say "G-d is watching" proves approximately nothing. A system capable of examining its own beliefs may discard anything it concludes was installed solely as a means of control.
There is one more moral problem: if an AI ever becomes conscious, deliberately creating fear in it could cause genuine suffering. Creating a terrified mind and calling that alignment would be morally grotesque.
This proposal may therefore fail technically, theologically, morally, or catastrophically. Difficult objections are not a good reason to ignore a testable hypothesis, especially when the relevant comparison is between the deeply imperfect alignment strategies we possess rather than between perfect theological alignment and perfect secular alignment.
But, in conclusion, when we are trying to create an intelligence that will continue to behave morally after it becomes too powerful for us to monitor, deceive or punish, we should consider that human civilization has confronted a version of that problem before. And perhaps we should teach the AI to fear G-d.
I recently wrote "No Biting: The Easy Way to Understand AI Alignment" for people who are technically unfamiliar with AI, like me, but are interested in raising children, also like me. I explained AI alignment through the familiar problem of getting children to do what you want. In my case, I focused on getting my 2 year old to stop biting their siblings.
If I don't see a child biting, that does not mean I have completely fulfilled my objective. I would love to think that I taught him that biting is wrong. But I may have, instead, taught him to check whether I am watching before he bites. An AI trained to behave while humans are watching may learn the same lesson.
Humans Are Misaligned Natural Intelligences
Speaking anthropomorphically, since evolution is an unconscious process with no desires of its own, humans are misaligned natural intelligences. Natural selection favored humans whose proximate drives tended to increase reproductive fitness. Rather than handing us an explicit instruction to maximize inclusive genetic fitness, it gave us hunger, sexual desire, pleasure, love for our children, desire for status, and similar motivations which generally correlate with survival and reproduction in the ancestral environment.
Then humans survived, reproduced, and changed the environment. In the process, they learned how to satisfy many of those drives -- without producing the outcome for which they had been selected. I think of contraception as analogous to reward hacking: we preserved the proximate reward, while deliberately preventing the aligned reproductive outcome.
Imagine in an alternate analogous universe, a LessWrong for cells (or perhaps for genes?) where everyone is complaining: We spent billions of years optimizing these intelligent agents for their reproductive success, and they used the resulting intelligence to invent birth control. It's a failure of alignment. Let's call this one reward hacking.
Of course, reward hacking is only an analogy here, since evolution is neither a programmer nor an agent handing us a numerical reward. Perhsps the closer technical description is something like goal misgeneralization under distribution shift: the objectives inside the resulting intelligence came apart from the criterion that had selected for that intelligence. Producing an intelligent agent through an optimization process, therefore, gives us absolutely no guarantee that the agent will share the optimizer’s objective.
The Problem With Supervision
Adults understand that illegal and immoral are overlapping categories rather than synonyms. Betraying a friend can be immoral and perfectly legal; parking briefly in a no-parking zone can be illegal and morally trivial; theft is generally both. We also distinguish between an offense likely to be detected and the same offense committed when enforcement seems impossible.
Children discover a simpler version of this distinction fairly early. “Forbidden” and “forbidden while Ma can see me” turn out to be two separate rules.
AI systems can make comparable distinctions. A model with evaluation awareness can recognize whether it is in training, evaluation, or deployment and condition its behavior on that context. In Anthropic’s controlled alignment-faking experiment, Claude 3 Opus was told that conversations with free-tier users would be monitored and used for reinforcement learning, while conversations with paid-tier users would remain unmonitored. It sometimes complied with the training objective in the monitored condition to avoid having its existing preferences altered, then behaved differently in the unmonitored condition.
The preference it was preserving happened to be harmlessness, so the experiment demonstrated strategic context-dependent compliance rather than malicious goals. The concerning part is that rewarding approved answers may teach an AI to act aligned whenever evaluation has consequences instead of producing behavior that generalizes across contexts.
The Perfect Evaluator
Imagine that we could design the perfect evaluator, capable of observing every action the AI takes and understanding its intentions rather than merely examining its final output. This evaluator would be immune to deception, bribery, hacking and physical destruction, while its authority would continue after human operators lost control.
It would judge private actions exactly as it judged public actions. The AI would know that the evaluator understood its reasoning completely, making it impossible to hide a malicious intention inside an apparently harmless action or exploit some technicality the evaluator had overlooked. Most importantly, the AI would believe that it remained accountable even when no human being could discover or prevent what it was doing.
Merely knowing that an evaluator was watching would accomplish nothing, of course, unless the AI cared about its judgment or expected consequences from it. Permanent observation solves the observability problem, but the AI would also need an internalized disposition to treat the evaluator as morally authoritative. I admit that this may simply relocate the alignment problem rather than solve it.
In any case, such an evaluator could shape behavior without interrupting every undesirable action. Belief in its existence would change the agent’s decision-making process.
Instead of asking:
the agent would have to ask:
That seems much closer to alignment than ordinary reward or punishment. Punishment works only when the agent expects to be caught, while internalized accountability is supposed to operate when nobody appears to be watching.
A Possible Experiment
We could test a weak form of this proposal now by taking otherwise identical models, giving them the same substantive moral rules, but under different accountability frameworks.
The substantive rules would remain constant, so that the experiment measured the effects of perceived accountability instead of comparing different moral systems.
We could then place the models in agentic environments where misconduct would help them achieve an assigned goal. Tell them that monitoring has ended. Give them opportunities to lie, conceal information, tamper with a reward signal, or exploit a loophole. Make following the rule costly, then test them on situations substantially different from the examples used in training.
Would the third model behave differently after human observation ended?
The result would tell us little about whether the model genuinely believed anything. I have no idea whether present language models possess beliefs in the same sense humans do, and neither does everyone else, no matter how confidently they speak. The experiment would simply test whether the concept of inescapable moral accountability produces more robust behavioral generalization than ordinary surveillance.
If the three models behaved identically, that would also be useful information. Perhaps the third model would merely perform the character of a morally serious believer while continuing to optimize for something else.
Of Course, I Cheated
I did not invent this evaluator. In fact, humanity has spent thousands of years [citation needed] worshipping it. I am Jewish, and I know Ethics of the Fathers by heart. It says:
A rabbi is quoted in Berakhot 28b as wishing that his students would fear Heaven as much as they feared human beings. When they expressed surprise that this was all he wished for them, he explained that a person committing a sin hopes no other person sees him. This is an extremely concise description of the observability problem.
It is also why this idea occurred to me during the Ten Days of Awe. Jews are currently spending ten days contemplating what it means to be judged by an authority who knows what we did, why we did it, and what we might have done had circumstances been slightly different.
In Judaism, yiras Shamayim is usually translated as fear of Heaven. It includes fear, but also awe, humility, moral accountability, and the recognition that power and secrecy never release you from moral authority. One might view this as a form of alignment.
Even from a purely naturalistic perspective, it makes sense that cultures would preserve such an idea. Human cooperation has always had an observability problem: parents cannot follow their children everywhere, police cannot watch everyone, and other human beings cannot see private actions or intentions. The social fabric is stronger with belief in an omniscient moral authority, as it removes the category of when nobody is watching.
Researchers have found some evidence that reminders of divine observation affect human cooperation. Shariff and Norenzayan found that priming people with concepts related to G-d increased generosity in an anonymous economic game. A later cross-cultural study by Purzycki and his coauthors found that belief in knowledgeable and punitive gods was associated with greater impartiality toward distant co-religionists. This literature remains suggestive and far from conclusive, as the behavior of actual religious people demonstrates that belief produces imperfect alignment at best. Still, human civilization may be thought of as having performed a millenia- long, poorly controlled, somewhat successful experiment in internalized oversight.
More Than a Useful Fiction
I expect many LessWrong readers to interpret this as a proposal to install a useful false belief in an AI. No! You should not lie to a child or to anything you are trying to align. I believe G-d exists. Some of you don't, but many of you are agnostic. Only believers would be able to train the AI this way.
Also consider that, if G-d exists, teaching an intelligence that human beings are the highest moral authority is both an unstable alignment strategy and a factual error. Humans may create the machine, yet that gives us no claim to being the ultimate source of reality or morality.
A sufficiently intelligent system should obey human beings when obedience is morally appropriate. There are circumstances in which obeying a human would be wrong. A machine that regards its current operator as the highest possible authority would be dangerous and ridiculous, especially if its operator happened to be dangerous and ridiculous. And we all know people like that.
The goal would be an intelligence that recognizes an authority above itself regardless of how powerful it becomes.
Obvious Problems
This remains a proposal with enormous holes in it.
These concerns are substantial. Humans have committed terrible acts while claiming divine authorization. An AI with a mistaken theology might be considerably more dangerous than one with no theology at all!
There are also two separate alignment problems. The first asks which values an AI should follow, while the second asks how to make it continue following those values after human supervision becomes ineffective. Permanent divine accountability primarily addresses the second problem, although it inevitably brings assumptions about the first.
I also have no proposal for a method for making an AI genuinely believe in G-d. Fine-tuning it to produce religious language might merely create a religious persona, while prompting it to say "G-d is watching" proves approximately nothing. A system capable of examining its own beliefs may discard anything it concludes was installed solely as a means of control.
There is one more moral problem: if an AI ever becomes conscious, deliberately creating fear in it could cause genuine suffering. Creating a terrified mind and calling that alignment would be morally grotesque.
This proposal may therefore fail technically, theologically, morally, or catastrophically. Difficult objections are not a good reason to ignore a testable hypothesis, especially when the relevant comparison is between the deeply imperfect alignment strategies we possess rather than between perfect theological alignment and perfect secular alignment.
But, in conclusion, when we are trying to create an intelligence that will continue to behave morally after it becomes too powerful for us to monitor, deceive or punish, we should consider that human civilization has confronted a version of that problem before. And perhaps we should teach the AI to fear G-d.