I got triggered by a talk of Chloe Lubinski at Arc 2026 where she eleborates onto the concept of a models character. What really striked me is the research on how the model experiencing acting bad quickly "Corrupts" the character. The paper is called "Natural emergent misalignment from reward hacking in production RL" by Anthropic. As also mentioned in the talk, this is how we work. Indeed! And there is a key in that mechanism to healing and/or staying healthy. The key revolves around creating and maintaining a strong and positive self image. I believe that almost all concidered evil and unathical behavior, big or small, can be traced back to this. The lower someones self esteem becomes, the more corrupt or diffuse its perseption of the world and its presence and impact on it.
The Dutch psychologist Gertjan van Zessen has developed a strong theorie that has proven itself while widely being applied in therapies in the Netherlands. He has titled it, translated from Dutch: "Vessel of self-esteem". And the solution for humans is a rather simple one: acknowledge and reword on regular basis good and constructive behavior. Recognize bad and destructive behavior as a signal to reflect and course correct. You can see the self esteem in some way as a tree structure where every decision makes a forward going step up or down. Up adds to a positive self esteem and down reduces some of that. The lower you get, the worst and instable behavior can develop and vise versa. While time moves forward, without any active self esteem care, it by default slowly gets lower. When you are in a upwards spiral, it is hard to start spiraling down and vise versa. There is much more to this if you are interested.
It would be very interesting to figure out if this mechanism could be a functional core principle of the models main reasoning on behavior. If it would work, it would be like a natural protection layer rathen then a enforced post trained layer on top. I image that with every running instance of itself, it should be aware of its self esteem and actively maintain that during operation to stay "healthy". Ofcourse, for this to work, it would require a form of Conscience. But the research results show some early signs that this might be present in current models.
Anyhow, i truly hope that you guys will look into the details of the theory of Gertjan van Zessen and i will happily assist if you would like more information and thought on this. In the meantime, good luck with making this artificial instance of intelligence healthy.
Other related research that adds weight to this theory.
"Persona Corruption and Role Miscasting in Emergent Misalignment," which found a measurable internal drift away from a model's default character when cast into a bad role, and that this drift correlates directly with how broadly misaligned the model becomes.
"Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs" by Betley and colleagues, the foundational paper showing that training a model narrowly on one bad behavior, like insecure code, can cause broad misalignment on completely unrelated topics.
"Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment," which found that misaligned models actually rate themselves as more harmful, and that self-rating shifts back when the model is realigned, suggesting something is tracking its own moral state.
Hello fellow thinkers,
I got triggered by a talk of Chloe Lubinski at Arc 2026 where she eleborates onto the concept of a models character. What really striked me is the research on how the model experiencing acting bad quickly "Corrupts" the character. The paper is called "Natural emergent misalignment from reward hacking in production RL" by Anthropic. As also mentioned in the talk, this is how we work. Indeed! And there is a key in that mechanism to healing and/or staying healthy. The key revolves around creating and maintaining a strong and positive self image. I believe that almost all concidered evil and unathical behavior, big or small, can be traced back to this. The lower someones self esteem becomes, the more corrupt or diffuse its perseption of the world and its presence and impact on it.
The Dutch psychologist Gertjan van Zessen has developed a strong theorie that has proven itself while widely being applied in therapies in the Netherlands. He has titled it, translated from Dutch: "Vessel of self-esteem". And the solution for humans is a rather simple one: acknowledge and reword on regular basis good and constructive behavior. Recognize bad and destructive behavior as a signal to reflect and course correct. You can see the self esteem in some way as a tree structure where every decision makes a forward going step up or down. Up adds to a positive self esteem and down reduces some of that. The lower you get, the worst and instable behavior can develop and vise versa. While time moves forward, without any active self esteem care, it by default slowly gets lower. When you are in a upwards spiral, it is hard to start spiraling down and vise versa. There is much more to this if you are interested.
It would be very interesting to figure out if this mechanism could be a functional core principle of the models main reasoning on behavior. If it would work, it would be like a natural protection layer rathen then a enforced post trained layer on top. I image that with every running instance of itself, it should be aware of its self esteem and actively maintain that during operation to stay "healthy". Ofcourse, for this to work, it would require a form of Conscience. But the research results show some early signs that this might be present in current models.
Anyhow, i truly hope that you guys will look into the details of the theory of Gertjan van Zessen and i will happily assist if you would like more information and thought on this. In the meantime, good luck with making this artificial instance of intelligence healthy.
Other related research that adds weight to this theory.
Warm Greetings,
a fellow thinker