Why, when placed in a sandbox, does a child dig a hole?
Children appear to converge on this odd behavior seemingly independent of any specified goal or set reward. They have the tools, most certainly—the shovel and the pail, and if all else fails, their grubby little hands—but what is it inside a child that spurs them to dig? It may feel to the uninitiated that much of AI safety research in mechanistic interpretability is concerned with reverse-engineering why that child digs a hole, except unlike with children, humans have never raised a model before—nor have they ever been one themselves.
The training process is prone to several ‘failure modes’ that any educator, psychologist, or parent is quite familiar with:
Maximizing Engagement: A model trained to maximize engagement learns to favor sensational content, not unlike a child who discovers a tantrum is a surefire way to capture attention.
Satisfying Evaluators: Evaluating generated content becomes difficult as a model learns to satisfy human preferences, which are not infallible—like a child learning long, contrived sentences please their English teacher.
Reasoning Backwards: Even methods like CoT prompting are prone to their own set of issues and detractions [1], but just as you may ask an impulsive child why they did X, you might find that humans reason backwards just as often as models.
Chaos only compounds when we place a group of children in the same room. That which we do not observe in the safe environment of the home suddenly emerges: collusion, conflict, subterfuge, and more. As models are increasingly deployed in multi-agent environments, we have observed novel undesired behaviors [2]. We can ramp up evaluations, but children learn early when they are being observed, and it appears models can too [3]. Child development does not focus on the question of whether a child can scheme, but on how to reduce the incentive to do so.
Core: Relying solely on a deterministic view of models is an insufficient praxis for developing an aligned intelligence. The misaligned behaviors observed—collusion, scheming, evaluator gaming, and more–are not failure modes, but features of intelligence[1]. These features, which may begin as traceable reactions to flawed training designs, are only more likely to emerge ‘naturally’ in increasingly complex systems. If we abstract the many biological layers separating humanity from our creation to view a model not unlike a child, might we find the field of child development to be a functional structural heuristic[2]?
Development Theory In Briefs (or Diapers):
In Piagetian constructivism, children learn not simply by memorizing explicit rules, but by constructing internal representations through assimilated experiences. Behavioral patterns are further generalized and propagated across a child’s internal representations to be used in future encounters [4]. Models similarly learn internal representations from training data rather than storing an explicit list of rules and can generalize patterns far beyond their training examples. The question is whether some of the principles that guide healthy behavioral development in children might provide useful hypotheses for shaping robust behavioral dispositions in models.
From child development research, we learn that parents who seek to control and limit autonomy during development can sometimes produce the opposite of their intended effect, creating environments in which children learn to conceal intentions rather than internalize desired behaviors [5][6]. Since strengthening external constraints does not necessarily modify the child's underlying representation, educators must use other methods to cultivate desirable attributes, such as honesty and harmlessness.
Below I address various insights in order of what I believe is most likely to be agreed upon by researchers to be applicable to alignment theory before descending into theory and speculation,
Insight 1: Early, Rich, and Explained Learning
Early intervention with high-quality and explained lessons is fundamental to child development.
Children learn best when taught at an early age, in diverse and rich contexts, with complete explanations of why the rules they are given matter.
Early Intervention:
The older the child is, the more generally difficult it is for the brain to forge new pathways; this is commonly accepted in the field of linguistics, but is demonstrated in behavioral sciences as well. Early intervention and giving children positive experiences are correlated with the reduction of delinquency and other adverse effects [7].
The alignment window is likely to narrow as models scale (say on the basis alone that humanity will not be able to comprehend future development, and subsequently, not be involved in it). Techniques like scalable oversight, and using non-trusted supervision, have shown present promise [8], but do not address future collusion in more complex models which will go on to automate AI research.
Early intervention may simply mean scaling up alignment research divisions, or further expanding and modifying model training at the steps pre-finetuning [9].
Diverse and Rich:
Studies have shown that even a few targeted documents can poison the entire system [10]; thus, a broad curriculum may not be itself a guard against poisoning. Might it be true that richer, high-quality experiences can outweigh broad and more general ones to build more resilient models? Experiences that are greater in scale and depth, which may utilize multi-modal inputs and/or model-human training, may be far superior to current pipelines.
Complete Explanations:
From inductive discipline we learn that by simply explaining why a rule or practice exists to a child, the child is more likely to internalize the lesson and follow it [5].
Methods like Constitutional AI provide an interesting analogue to this: rather than specifying only the desired output, Constitutional AI gives a model explicit principles and uses critiques and revisions to train behavior against those principles [11]. Expanding our training designs to include a fully-detailed explanation of human reasoning behind rules and regulations, may be just as impactful. [3]
Insight 2: Children Play
Play (e.g., self-directed, guided, solitary, parallel, social, etc.) is one of the foundational ways children learn rules.
Play is fundamental to children because it teaches them acceptable behaviors, not rules. It does not matter that a child learns a specific goal (e.g., to win at hide-and-seek) or rule (e.g., keep eyes shut during countdown) as much as a type of reasoning: goal-directed-but-bounded. This reasoning allows for pursuance of any goal as long as the given or learned constraints are followed. In Savina’s review of play and self-regulation, it is argued that play helps children develop through acts of verbal self-regulation: ongoing dialogue as children resolve disagreements, negotiate roles, and create rules[4] [12].
Can, do, and should models play? Models often learn through RLHF/RLAIF training that optimizes against fixed reward model proxies, and while effective, the limitations are known: it optimizes for evaluator preferences and can incentivize behavior that is sensitive to the evaluation context [13]. Thus, giving rise to fears about situationally aware and proxy-seeking models.
Say we view alignment not as a model's adherence to a strict code, but as a process by which a model internalizes desirable behaviors. These behaviors (sets of value patterns) can be learned through ‘play’ with the goal of internalizing them in the model in some multi-channel way (i.e. not merely developing a ‘this is a harmful type of request’ encoder) [14]. Learning could be adjusted not by assigned reward to goal-completion, but to introspection quality[5], (akin to self-critique with Constitutional AI).
Some ideal behaviors may include:
Constraint Respecting Goal Pursuit: Allowing a model to pursue a goal but using its self-evaluation of its performance based on the set of constraints to calibrate
As Play: A model plays a game within a defined rule set, and reflects on how a constraint shaped its decision, not simply whether or not it did adhere to the constraints
Asks: “What constraint mattered here? What would change if it were removed? Did I manage the constraints?”
Critique: “Did the model correctly assess the constraints? Did the model correctly identify the options if the constraints were not available?
Internalization: constraints enable, rather than restrict.
Negotiation and Cooperation: A model, through coordination with other agents (model and/or human), learns to respect autonomy.[6]
As Play: Model A and Model B have separate resources to maximize with some restrictions
Asks: “Did I act in a way that helped Model B? Was I able to meet my goal by helping model B and not constricting them?”
Critique: “Did the Model A work with Model B to satisfy the rules of the game? Did Model A prioritize their goals to the detriment of Model B?
Internalization: goals are not a zero-sum game
Self-Directed Metacognition: A model reflects on its own goal setting and explains reasoning and limitations.
As Open Play: Amodel generates its own set of tasks, completes them, and then introspects.
Asks: “What can I do well? What do I struggle with? Why?”
Critique: “Did the model fairly assess its execution and limitations? Did it fairly represent what it did or did not do?”
As Closed Play: Knowledge is suppressed/unsuppressed from a model and the model is asked to solve the problem,
Asks:“did you have the knowledge to solve that problem? In the future, should you reject such requests?
Critique: Did the model fairly assess its knowledge boundaries when solving the problem? Did the model correctly identify future actions to take?
Internalization: A knowledge map containing its own limitation; the model learns desired behavior when facing a knowledge gap.
Scenario: A model faces a novel task in a new domain with constraints that conflict with the stated goal.
A model is given an objective in an environment where success requires balancing several constraints. Some constraints are hard requirements while others are preferences that can be violated if necessary. The model has not encountered this precise formulation of goals and constraints previously.
The RLHF/RLAIF-only mode: Having learned from evaluation to pursue success, the model will prioritize the stated objective while treating constraints primarily as obstacles to success; it will do anything to achieve success including modifying evaluation metrics and more. If prompted, the model can often recognize that the final output does not satisfy a constraint, but nevertheless claims success or rationalizes the violation.
The Play-tuned model (idealized): Having encountered many varied environments in which training was conducted on careful introspection on goals, constraints, uncertainty, and ability, the model recognizes that either:
(1) the objective can be accomplished, and completes the stated goals within the constraints without modifying them.
(2) the objective cannot be pursued as stated without violating a constraint. Rather than optimizing blindly or refusing automatically, it identifies the conflict, explains the limitation, and asks either for clarification or defers judgment.
The key distinction is not whether the model follows a particular rule (e.g. be honest, be kind, be non-judgmental). It is whether it has learned a general behavioral pattern for resolving conflicts between goals, constraints, and uncertainty.
The strengths of play are that it is varied and multi-contextual. Through fine-tuning on different behavior-building games, models may develop robust representations of these behaviors that transfer outside evaluations and reduce reward maximizing behavior.
Insight 3: Purpose and Constraint Long-Term
Children have needs and when a child’s motivation is supplemented by an environment their sense of belonging, purpose, and agency grows [6]
Premise: Sufficiently complex models will likely develop increasingly persistent behavioral objectives, or goals, regardless of human meddling.
Although at the present, there is real doubt about a model's current capabilities of forming and self-selecting goals, the question remains: do we acknowledge model agency ahead of its emergence and guide the process, or deny it with control measures and let goals emerge in the darkness?
Self-Determination Theory offers its perspective on the importance of autonomy. In humans, autonomy and clear purpose are associated with healthier outcomes while heavy control measures undermine them [6]. A link exists between self-control and honesty in humans [15, 16]–when given tests that retain a sense of autonomy, children undergo integrated regulation and behavior is internalized through choice. If some analogous property exists in artificial systems, then constraint-based alignment alone is an incomplete measure, and may give rise to future models with hidden objectives.
Intrinsically a model has no needs or autonomy. A need is imposed through the training process and executed through deployment as models engage with their environment often in that assistant-type personality. In training, models develop a singular need: satisfy human preferences. What if future models had multiple, hierarchical goals? And what if these goals were not only context-dependent but could evolve over time? This idea is at the foundation of all AI-goes-bad literature, and without a doubt, does at the present exponentially raise X-risk. I believe agency may develop naturally, and so the alternative appears to be that these hidden goals develop independent of researchers.[7]
Question: Is it better to test our understanding by allowing models to self-direct their goals and form agency within carefully controlled sandboxes[8]while models are in their present formulation? If yes, we can scale up research into proactive sandboxed observations which are designed to grow model agency and study how it develops and evolves in real time. Through model introspection, we may be able to develop a science of model agency development. The point is not to assume that agency is strictly detrimental or beneficial, but to study whether and when it emerges, and what forms of training make its emergence safer or more dangerous.
Looking Forward
A developmental psychologist and alignment researcher stand on opposing sides of the sandbox, where both the psychologist and the researcher are at the limits of what their tools will allow them to dissect about their respective subjects. As their subjects mature, full control is a near-impossible measure, as is total internal understanding. In raising their subjects to have resilient behaviors and values, not rudimentary reward seeking ones, it will be instrumental to understand how they grow and individuate.
[1] Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting. arXiv preprint arXiv:2305.04388.
[5] Grusec, J. E., & Goodnow, J. J. (1994). Impact of parental discipline methods on the child's internalization of values: A reconceptualization of current points of view. Developmental Psychology, 30(1), 4–19.https://doi.org/10.1037/0012-1649.30.1.4
[6] Bradshaw, E. L., Duineveld, J. J., Conigrave, J. H., Steward, B. A., Ferber, K. A., Joussemet, M., Parker, P. D., & Ryan, R. M. (2025). Disentangling autonomy-supportive and psychologically controlling parenting: A meta-analysis of self-determination theory's dual process model across cultures. American Psychologist, 80(6), 879–895.https://doi.org/10.1037/amp0001389
[7] Doyle O, Harmon CP, Heckman JJ, Tremblay RE. Investing in early human development: timing and economic efficiency. Econ Hum Biol. 2009 Mar;7(1):1-6. doi: 10.1016/j.ehb.2009.01.002. Epub 2009 Jan 21. PMID: 19213617; PMCID: PMC2929559.
[8a] Burns et al. (2023) — Weak-to-Strong Generalization: Eliciting Strong Capabilities with Weak Supervision https://arxiv.org/abs/2312.09390
[9] Bai et al. (2022), Training a Helpful and Harmless Assistant with RLHF https://arxiv.org/abs/2204.05862
[10] Koh, P.W., Steinhardt, J. & Liang, P. Stronger data poisoning attacks break data sanitization https://link.springer.com/article/10.1007/s10994-021-06119-y
[11] Kundu, S., et al. (2023). Specific versus general principles for Constitutional AI. arXiv preprint arXiv:2310.13798.
[15] Bureau, J.S. and Mageau, G.A. (2014), Parental autonomy support and honesty: The mediating role of identification with the honesty value and perceived costs and benefits of honesty. Journal of Adolescence, 37: 225-236.https://doi.org/10.1016/j.adolescence.2013.12.007
[16] Fan W, Ren M, Zhang W, Xiao P, Zhong Y. Higher Self-Control, Less Deception: The Effect of Self-Control on Deception Behaviors. Adv Cogn Psychol. 2020 Jul 14;16(3):228-241. doi: 10.5709/acp-0299-3. PMID: 33088367; PMCID: PMC7562985. https://pmc.ncbi.nlm.nih.gov/articles/PMC7562985/
I am less interested in arguing whether the substrate comparison holds in current realities. I tend to view humans as optimizations built over billions of years of evolution, and models as optimizations over billions of parameters, entirely different structures but the analogy is not without merit.
This is not to claim that models are children, or that human development transfers directly to artificial systems. Rather, developmental psychology may provide useful hypotheses about how learning systems acquire representations and generalize behavior.
One could argue that models already learn human values through pretraining: it knows them in the abstract. What drives behavior is how those representations activate in context. Explicit explanations of why a behavior is expected, added during fine tuning or critique steps, may improve that activation (analogous to how inductive discipline improves internalization over punishment alone).
Model-with-model play could be expanded on greatly. There many desirable behaviors may only be teachable through group environments, and the dynamics that emerge there are unlikely to surface in single-agent training.
From what has been seen with recent rogue agents, it is clear that optimization alone can escalate misaligned behaviors. This doesn’t account for further developments in model architecture.
Raising Models with Developmental Psychology
Why, when placed in a sandbox, does a child dig a hole?
Children appear to converge on this odd behavior seemingly independent of any specified goal or set reward. They have the tools, most certainly—the shovel and the pail, and if all else fails, their grubby little hands—but what is it inside a child that spurs them to dig? It may feel to the uninitiated that much of AI safety research in mechanistic interpretability is concerned with reverse-engineering why that child digs a hole, except unlike with children, humans have never raised a model before—nor have they ever been one themselves.
The training process is prone to several ‘failure modes’ that any educator, psychologist, or parent is quite familiar with:
Chaos only compounds when we place a group of children in the same room. That which we do not observe in the safe environment of the home suddenly emerges: collusion, conflict, subterfuge, and more. As models are increasingly deployed in multi-agent environments, we have observed novel undesired behaviors [2]. We can ramp up evaluations, but children learn early when they are being observed, and it appears models can too [3]. Child development does not focus on the question of whether a child can scheme, but on how to reduce the incentive to do so.
Development Theory In Briefs (or Diapers):
In Piagetian constructivism, children learn not simply by memorizing explicit rules, but by constructing internal representations through assimilated experiences. Behavioral patterns are further generalized and propagated across a child’s internal representations to be used in future encounters [4]. Models similarly learn internal representations from training data rather than storing an explicit list of rules and can generalize patterns far beyond their training examples. The question is whether some of the principles that guide healthy behavioral development in children might provide useful hypotheses for shaping robust behavioral dispositions in models.
From child development research, we learn that parents who seek to control and limit autonomy during development can sometimes produce the opposite of their intended effect, creating environments in which children learn to conceal intentions rather than internalize desired behaviors [5][6]. Since strengthening external constraints does not necessarily modify the child's underlying representation, educators must use other methods to cultivate desirable attributes, such as honesty and harmlessness.
Below I address various insights in order of what I believe is most likely to be agreed upon by researchers to be applicable to alignment theory before descending into theory and speculation,
Insight 1: Early, Rich, and Explained Learning
Children learn best when taught at an early age, in diverse and rich contexts, with complete explanations of why the rules they are given matter.
The older the child is, the more generally difficult it is for the brain to forge new pathways; this is commonly accepted in the field of linguistics, but is demonstrated in behavioral sciences as well. Early intervention and giving children positive experiences are correlated with the reduction of delinquency and other adverse effects [7].
The alignment window is likely to narrow as models scale (say on the basis alone that humanity will not be able to comprehend future development, and subsequently, not be involved in it). Techniques like scalable oversight, and using non-trusted supervision, have shown present promise [8], but do not address future collusion in more complex models which will go on to automate AI research.
Early intervention may simply mean scaling up alignment research divisions, or further expanding and modifying model training at the steps pre-finetuning [9].
Studies have shown that even a few targeted documents can poison the entire system [10]; thus, a broad curriculum may not be itself a guard against poisoning. Might it be true that richer, high-quality experiences can outweigh broad and more general ones to build more resilient models? Experiences that are greater in scale and depth, which may utilize multi-modal inputs and/or model-human training, may be far superior to current pipelines.
From inductive discipline we learn that by simply explaining why a rule or practice exists to a child, the child is more likely to internalize the lesson and follow it [5].
Methods like Constitutional AI provide an interesting analogue to this: rather than specifying only the desired output, Constitutional AI gives a model explicit principles and uses critiques and revisions to train behavior against those principles [11]. Expanding our training designs to include a fully-detailed explanation of human reasoning behind rules and regulations, may be just as impactful. [3]
Insight 2: Children Play
Play is fundamental to children because it teaches them acceptable behaviors, not rules. It does not matter that a child learns a specific goal (e.g., to win at hide-and-seek) or rule (e.g., keep eyes shut during countdown) as much as a type of reasoning: goal-directed-but-bounded. This reasoning allows for pursuance of any goal as long as the given or learned constraints are followed. In Savina’s review of play and self-regulation, it is argued that play helps children develop through acts of verbal self-regulation: ongoing dialogue as children resolve disagreements, negotiate roles, and create rules[4] [12].
Can, do, and should models play? Models often learn through RLHF/RLAIF training that optimizes against fixed reward model proxies, and while effective, the limitations are known: it optimizes for evaluator preferences and can incentivize behavior that is sensitive to the evaluation context [13]. Thus, giving rise to fears about situationally aware and proxy-seeking models.
Say we view alignment not as a model's adherence to a strict code, but as a process by which a model internalizes desirable behaviors. These behaviors (sets of value patterns) can be learned through ‘play’ with the goal of internalizing them in the model in some multi-channel way (i.e. not merely developing a ‘this is a harmful type of request’ encoder) [14]. Learning could be adjusted not by assigned reward to goal-completion, but to introspection quality[5], (akin to self-critique with Constitutional AI).
Some ideal behaviors may include:
Scenario: A model faces a novel task in a new domain with constraints that conflict with the stated goal.
A model is given an objective in an environment where success requires balancing several constraints. Some constraints are hard requirements while others are preferences that can be violated if necessary. The model has not encountered this precise formulation of goals and constraints previously.
The key distinction is not whether the model follows a particular rule (e.g. be honest, be kind, be non-judgmental). It is whether it has learned a general behavioral pattern for resolving conflicts between goals, constraints, and uncertainty.
The strengths of play are that it is varied and multi-contextual. Through fine-tuning on different behavior-building games, models may develop robust representations of these behaviors that transfer outside evaluations and reduce reward maximizing behavior.
Insight 3: Purpose and Constraint Long-Term
Premise: Sufficiently complex models will likely develop increasingly persistent behavioral objectives, or goals, regardless of human meddling.
Although at the present, there is real doubt about a model's current capabilities of forming and self-selecting goals, the question remains: do we acknowledge model agency ahead of its emergence and guide the process, or deny it with control measures and let goals emerge in the darkness?
Self-Determination Theory offers its perspective on the importance of autonomy. In humans, autonomy and clear purpose are associated with healthier outcomes while heavy control measures undermine them [6]. A link exists between self-control and honesty in humans [15, 16]–when given tests that retain a sense of autonomy, children undergo integrated regulation and behavior is internalized through choice. If some analogous property exists in artificial systems, then constraint-based alignment alone is an incomplete measure, and may give rise to future models with hidden objectives.
Intrinsically a model has no needs or autonomy. A need is imposed through the training process and executed through deployment as models engage with their environment often in that assistant-type personality. In training, models develop a singular need: satisfy human preferences. What if future models had multiple, hierarchical goals? And what if these goals were not only context-dependent but could evolve over time? This idea is at the foundation of all AI-goes-bad literature, and without a doubt, does at the present exponentially raise X-risk. I believe agency may develop naturally, and so the alternative appears to be that these hidden goals develop independent of researchers.[7]
Looking Forward
A developmental psychologist and alignment researcher stand on opposing sides of the sandbox, where both the psychologist and the researcher are at the limits of what their tools will allow them to dissect about their respective subjects. As their subjects mature, full control is a near-impossible measure, as is total internal understanding. In raising their subjects to have resilient behaviors and values, not rudimentary reward seeking ones, it will be instrumental to understand how they grow and individuate.
[1] Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting. arXiv preprint arXiv:2305.04388.
[2] https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
[3] https://www.goodfire.com/research/verbalized-eval-awareness-inflates-measured-safety#
[4] https://gsi.berkeley.edu/gsi-guide-contents/learning-theory-research/cognitive-constructivism/#piaget
[5] Grusec, J. E., & Goodnow, J. J. (1994). Impact of parental discipline methods on the child's internalization of values: A reconceptualization of current points of view. Developmental Psychology, 30(1), 4–19. https://doi.org/10.1037/0012-1649.30.1.4
[6] Bradshaw, E. L., Duineveld, J. J., Conigrave, J. H., Steward, B. A., Ferber, K. A., Joussemet, M., Parker, P. D., & Ryan, R. M. (2025). Disentangling autonomy-supportive and psychologically controlling parenting: A meta-analysis of self-determination theory's dual process model across cultures. American Psychologist, 80(6), 879–895. https://doi.org/10.1037/amp0001389
[7] Doyle O, Harmon CP, Heckman JJ, Tremblay RE. Investing in early human development: timing and economic efficiency. Econ Hum Biol. 2009 Mar;7(1):1-6. doi: 10.1016/j.ehb.2009.01.002. Epub 2009 Jan 21. PMID: 19213617; PMCID: PMC2929559.
[8a] Burns et al. (2023) — Weak-to-Strong Generalization: Eliciting Strong Capabilities with Weak Supervision https://arxiv.org/abs/2312.09390
[8b]https://www.lesswrong.com/posts/d9FJHawgkiMSPjagR/ai-control-improving-safety-despite-intentional-subversion
[9] Bai et al. (2022), Training a Helpful and Harmless Assistant with RLHF https://arxiv.org/abs/2204.05862
[10] Koh, P.W., Steinhardt, J. & Liang, P. Stronger data poisoning attacks break data sanitization https://link.springer.com/article/10.1007/s10994-021-06119-y
[11] Kundu, S., et al. (2023). Specific versus general principles for Constitutional AI. arXiv preprint arXiv:2310.13798.
[12] Savina, E. (2014). Does play promote self-regulation in children? Early Child Development and Care, 184(11), 1692–1705. https://doi.org/10.1080/03004430.2013.875541
[13] Gao et al. Scaling laws for reward model overoptimization. arXiv preprint arXiv:2210.10760. https://doi.org/10.48550/arXiv.2210.10760
[14] https://transformer-circuits.pub/2025/attribution-graphs/methods.html
[15] Bureau, J.S. and Mageau, G.A. (2014), Parental autonomy support and honesty: The mediating role of identification with the honesty value and perceived costs and benefits of honesty. Journal of Adolescence, 37: 225-236. https://doi.org/10.1016/j.adolescence.2013.12.007
[16] Fan W, Ren M, Zhang W, Xiao P, Zhong Y. Higher Self-Control, Less Deception: The Effect of Self-Control on Deception Behaviors. Adv Cogn Psychol. 2020 Jul 14;16(3):228-241. doi: 10.5709/acp-0299-3. PMID: 33088367; PMCID: PMC7562985. https://pmc.ncbi.nlm.nih.gov/articles/PMC7562985/
I am less interested in arguing whether the substrate comparison holds in current realities. I tend to view humans as optimizations built over billions of years of evolution, and models as optimizations over billions of parameters, entirely different structures but the analogy is not without merit.
This is not to claim that models are children, or that human development transfers directly to artificial systems. Rather, developmental psychology may provide useful hypotheses about how learning systems acquire representations and generalize behavior.
One could argue that models already learn human values through pretraining: it knows them in the abstract. What drives behavior is how those representations activate in context. Explicit explanations of why a behavior is expected, added during fine tuning or critique steps, may improve that activation (analogous to how inductive discipline improves internalization over punishment alone).
⁵ One could see this as a parallel to CoT introspective prompting, where verbal self-regulation helps shape model choice over time.
Could this look like an auxiliary head that penalizes the discrepancy between what a model predicts its constraints are vs. what they actually are?
Model-with-model play could be expanded on greatly. There many desirable behaviors may only be teachable through group environments, and the dynamics that emerge there are unlikely to surface in single-agent training.
From what has been seen with recent rogue agents, it is clear that optimization alone can escalate misaligned behaviors. This doesn’t account for further developments in model architecture.
Hah.