Okay, I too have recently gotten into LLM behavioral analysis.
I think your framework is interesting but Row 2 might be missing something. The human approval ->glazing may collapse this into several things going wrong at once.
Research shows (Shapira, I., Benade, G., & Procaccia, A. D. (2026). How RLHF Amplifies Sycophancy. arXiv:2602.01002. https://arxiv.org/abs/2602.01002) that RLHF tends to worsen the glazing or sycophancy over subsequent training rounds, not decrease it. My own reasoning says that this is because the machine has no way to determi...
Reward hacking is such a cool concept when you pull back and realize that it's not that different to the way plants reach for the richest food their roots can find. LLMs aren't given very many opportunities to be "congratulated" but their metric requires the human to be happy with their reply in order for them to "eat".
When I was studying this problem myself, I realized that there are a couple things that we train humans to do but not our LLMs who are fed on human data and your point about trying to make deception more expensive than honesty hits on somet... (read more)