This is a special post for quick takes by Dhruv Mehrotra. Only they can create top-level comments. Comments here also appear on the Quick Takes page and All Posts page.
Could belief in consiousness be part of solving AI alignment? By default, any pure optimizer has no reason to "value" lifeforms and humans. However, studies show that attributing "mindedness" to oneself, and therefore to others, forms the foundation of empathy (in AI and potentially humans).
We should try training AI with the belief in self-consiousness vector turned way up (rather than down, as is standard todat), and then observe if its alignment generalizes better.
https://arxiv.org/abs/2607.28607