When I say imagine being happy, everyone will have a flashback of a different moment in their life - some might imagine staying close to their loved ones; for some, happiness might be the day they became parents, found love, got an award (something along terms of achieved "X", did "Y", became "Z"). For someone else, happiness might mean, doing things that made a positive change in the world or in someone's life.
This tells us two things:
A simple concept like happiness has different associations inside people's heads.
People have different preferences for being happy.
Moreover, preferences and philosophies which make people happy differ extensively: I am happy when everyone gets a share of my and others' profits equally, or when the profit is divided according to someone's hard work and output or just the hardworking are rewarded.
Now imagine, we tell AI to just maximise for collective happiness. Wouldn't it go haywire with so many differing preferences? what to choose from? how to make everyone happy? does it maximise #people who are happy or maximise the #people subject to inclusion of all groups (irrespective of size)?
This raises two challenges: how to develop an understanding of different instances of happiness and preferences? Secondly how to align for those preferences. Also, can that understanding be robust?
Our human minds differ in understanding of the same concepts. If I point my finger to the sky and say "oh, they are watching, do well". Now different people will interpret it very differently based on their understanding of "they" and "well". Several studies show how interpretations of same words or concepts differ significantly in human minds [1, 2, 3].
Similar differences in interpretations have been shown in AI models as well. Different model families differ in their understanding of the same concepts [4, 5], and differ with humans as well [6, 7].
So a fundamental question is, how to ensure AI understands what we specify in abstract terms the way it is in our heads?
When we say maximise for X (say collective happiness or wellbeing), would it be able to derive the principles of X and apply them to actually maximise X?
I am thinking on how to develop this common ground and understanding, these days.Will be posting more and tending to keep this post concise (this marks my first post on LessWrong!).
PS: I have written it in one go and am posting the first draft; I am happy to receive any feedback to improve my writing and discuss any disagreements as well.
Thanks for reading through!
[1] Marti, L., Wu, S., Piantadosi, S. T., & Kidd, C. (2023). Latent diversity in human concepts. Open Mind, 7, 79-92.
[2] Thompson, B., Roberts, S. G., & Lupyan, G. (2020). Cultural influences on word meanings revealed through large-scale semantic alignment. Nature Human Behaviour, 4(10), 1029-1038.
[3] Charest, I., Kievit, R. A., Schmitz, T. W., Deca, D., & Kriegeskorte, N. (2014). Unique semantic space in the brain of each beholder predicts perceived similarity. Proceedings of the National Academy of Sciences, 111(40), 14565-14570.
[4] Murthy, S. K., Ullman, T., & Hu, J. (2025, April). One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 11241-11258).
[5] Chang, J., Piff, L., Sana, S., Li, J., & Levine, L. (2026, April). EigenBench: A Comparative Behavioral Measure of Value Alignment. In International Conference on Learning Representations (Vol. 2026, pp. 72657-72685).
[6] Chlapanis, O. S., Mastromichalakis, O. M., & Papadimitriou, C. H. (2026). The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans. arXiv preprint arXiv:2605.08837.
[7] Xu, Q., Peng, Y., Nastase, S. A., Chodorow, M., Wu, M., & Li, P. (2025). Large language models without grounding recover non-sensorimotor but not sensorimotor features of human concepts. Nature human behaviour, 9(9), 1871-1886.
When I say imagine being happy, everyone will have a flashback of a different moment in their life - some might imagine staying close to their loved ones; for some, happiness might be the day they became parents, found love, got an award (something along terms of achieved "X", did "Y", became "Z"). For someone else, happiness might mean, doing things that made a positive change in the world or in someone's life.
This tells us two things:
Moreover, preferences and philosophies which make people happy differ extensively: I am happy when everyone gets a share of my and others' profits equally, or when the profit is divided according to someone's hard work and output or just the hardworking are rewarded.
Now imagine, we tell AI to just maximise for collective happiness. Wouldn't it go haywire with so many differing preferences? what to choose from? how to make everyone happy? does it maximise #people who are happy or maximise the #people subject to inclusion of all groups (irrespective of size)?
This raises two challenges: how to develop an understanding of different instances of happiness and preferences? Secondly how to align for those preferences. Also, can that understanding be robust?
Our human minds differ in understanding of the same concepts. If I point my finger to the sky and say "oh, they are watching, do well". Now different people will interpret it very differently based on their understanding of "they" and "well". Several studies show how interpretations of same words or concepts differ significantly in human minds [1, 2, 3].
Similar differences in interpretations have been shown in AI models as well. Different model families differ in their understanding of the same concepts [4, 5], and differ with humans as well [6, 7].
So a fundamental question is, how to ensure AI understands what we specify in abstract terms the way it is in our heads?
When we say maximise for X (say collective happiness or wellbeing), would it be able to derive the principles of X and apply them to actually maximise X?
I am thinking on how to develop this common ground and understanding, these days.Will be posting more and tending to keep this post concise (this marks my first post on LessWrong!).
PS: I have written it in one go and am posting the first draft; I am happy to receive any feedback to improve my writing and discuss any disagreements as well.
Thanks for reading through!
[1] Marti, L., Wu, S., Piantadosi, S. T., & Kidd, C. (2023). Latent diversity in human concepts. Open Mind, 7, 79-92.
[2] Thompson, B., Roberts, S. G., & Lupyan, G. (2020). Cultural influences on word meanings revealed through large-scale semantic alignment. Nature Human Behaviour, 4(10), 1029-1038.
[3] Charest, I., Kievit, R. A., Schmitz, T. W., Deca, D., & Kriegeskorte, N. (2014). Unique semantic space in the brain of each beholder predicts perceived similarity. Proceedings of the National Academy of Sciences, 111(40), 14565-14570.
[4] Murthy, S. K., Ullman, T., & Hu, J. (2025, April). One fish, two fish, but not the whole sea: Alignment reduces language models’ conceptual diversity. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 11241-11258).
[5] Chang, J., Piff, L., Sana, S., Li, J., & Levine, L. (2026, April). EigenBench: A Comparative Behavioral Measure of Value Alignment. In International Conference on Learning Representations (Vol. 2026, pp. 72657-72685).
[6] Chlapanis, O. S., Mastromichalakis, O. M., & Papadimitriou, C. H. (2026). The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans. arXiv preprint arXiv:2605.08837.
[7] Xu, Q., Peng, Y., Nastase, S. A., Chodorow, M., Wu, M., & Li, P. (2025). Large language models without grounding recover non-sensorimotor but not sensorimotor features of human concepts. Nature human behaviour, 9(9), 1871-1886.