Compassion Aligned Machine Learning (CaML) is conducting research and deploying benchmarks with the aim of implementing broad compassion for all sentient beings into AI systems. In so doing, we’re attempting to integrate ideas from across the broad field of alignment research, including those focussed on specific technical questions, and those considering more strategic questions around who or what AI should be aligned with. To do this, it's important for us to understand how the alignment community is conceptualizing and prioritizing different ideas that CaML cares about.
In this spirit, we recently conducted a survey on controversial questions in alignment, to encourage more discussion on topics important to CaML, and roughly gauge where both core researchers and wider community members stand. Questions ranged from the value of benchmarking as eval awareness increases, to the implications of AGI for animals, to thoughts about AI s-risks. We’ve shared some interesting findings from this survey below.
Our intent with this survey was not to conduct an exhaustive analysis of the community’s position on every issue. This was an initial probe primarily designed to generate discussion and feedback on issues important to CaML. That said, we’d love to see research more accurately tracking core areas of (dis)agreement, consensus and controversy within alignment research. If we know where there is uncertainty, we can better understand where our research should focus next.
What We Did
Our survey consisted of two stages.
In the first, we shared 14 polls with a panel of 14 researchers, where they privately responded to statements on an 11-point agree-disagree scale. They could also share the reasoning behind their answers. We received responses from the following 14 researchers:
Scott Alexander (Astral Codex Ten)
Cameron Berg (Reciprocal Research)
Oscar Horta (Animal Ethics)
Jeff Sebo (NYU CMEP)
Chris Pang (NVIDIA)
Tobias Baumann (CRS)
Michael Dickens (Independent)
Dawn Drescher (Impartial Priorities)
Tristan Katz (Rethink Priorities)
Max Taylor (ACE)
David Manheim (ALTER)
Cameron Holmes (Independent)
Javier Ferrando (Amazon)
Raphael Sarfati (Goodfire)
Please note that respondents took these polls in an individual capacity, and responses do not represent the views of their employers
In the second stage, we released the survey totalling 15 polls (with one additional poll to the panel survey) to the broader community. Respondents answered via an EA Forum post (where all the polls are still viewable), and could publicly share their reasoning for different responses. After feedback from the expert panel, we included clarifications alongside the polls. We received a total of 1444 votes over all 15 polls, and over 100 comments providing further reasoning.
Once the survey was completed, we analysed two primary variables across the 15 questions: the lean (whether respondents on average voted agree or disagree), and standard deviation (SD) to see the spread of responses on a given question. Lean is measured from -5 meaning strongly disagree to +5 meaning strongly agree (with 0 being neither agree or disagree).
What We Found
Collective lean per statement, sorted by community lean. Grey connectors join the community and panel means for the same statement; rows where the dots sit on opposite sides of the dashed line are statements on which the two samples lean in opposite directions. The right-hand column gives the standard deviation of responses for the community (C) and panel (P). *Asked to the community only.
Most People see the s-risk field as underperforming
In two polls, we asked explicitly about the value of S-risk work:
Poll six: At the margin, S-risk work in AI is more important than X-risk work.
Poll eight: We are making good progress in the AI S-risk space, and research is on track.
On both of these questions, the community leant towards disagreeing with the statement (lean of -1.2 on poll six, and -2.0 on poll eight). Only 3 of 73 respondents agreed that s-risk progress was on track and not one strongly agreed (lean of more than 1.5).
Poll six saw the biggest distance between the community and the expert panel, with the panel leaning agree (+1.2) that s-risk work was more valuable at the margin. The panel was somewhat split on this though: this poll had the highest standard deviation (3.2) of the set. This may have been somewhat representative of the different research interests of the panel, with some explicitly more focussed on s-risks than x-risks or vice versa.
Nevertheless, both the panel and the community disagreed that s-risk work was on track. Some text responses gave more detail as to why:
‘It’s notoriously difficult to say due to lack of clear feedback loops’
‘Depends a lot on how good progress is defined’
‘The space seems highly contested, deeply illegible, few viable research directions are immediately present’
Regardless of how respondents prioritise s-risks, there was at least agreement that, as it stands, what progress actually looks like in the s-risk space is difficult to identify, be that due to the field’s far-future focus, small size, or general uncertainty. This suggests a primary challenge for those currently focussed on s-risk research: identify and communicate how progress could be measured in concrete ways in the field, and consider what robustly good things can be done under uncertainty.
People want compassion promoted in AI systems
The final poll in the community survey asked ‘Insofar as alignment continues to promote overlapping sets of characteristics (e.g. helpful, harmless, honest, corrigible), should compassion be one of those characteristics?’ This question was not shared with the panel, but it had the highest rate of agreement on the EA Forum of any poll: 89% of respondents agreed (lean +2.4, SD 1.9). Few voters articulated their reasoning on this question, though some responders did add that care must be taken in how this value is conceived of and pursued:
[Agreeing] ‘I think so, but be careful what you wish for.’
[Disagreeing] ‘Compassion seems a bit ill defined and can be interpreted in vastly more or less paternalistic ways.’
This poll points directly at CaML’s core mission and suggests broad support among respondents for the kinds of values we’d like to instil in AI systems. What we need to show to labs and the alignment community now, is how broad compassion can be implemented alongside other important values in a way that is safe and robust.
Vote distributions for three featured statements, on the same -5 (strongly disagree) to +5 (strongly agree) scale. Blue bars are community votes; coral dots below the baseline are the 14 individual panel responses (the compassion statement was community-only). The shapes show the findings: no community votes above +1.5 on s-risk progress, two opposing camps on non-human neglect, and the majority leaning agree on compassion.
It’s controversial whether non-human suffering is neglected by AI Safety
‘Non-human suffering is neglected by the AI safety community’ has the highest standard deviation among the community polls (3.1). In agreement, one respondent said: ‘as I understand, many key players are explicitly anthropocentrists. Hell, some key players care more about money than humans, let alone non-humans.’ In disagreement, another respondent said work had already been done to make Claude care about non-human suffering, and that it wasn’t even clear if this was a valuable direction to pursue (due to potential backfire effects, and the generalisation of human alignment to non-humans). Some of the disagreement here may have been due to the respondents’ reference point: those voting agree were more likely to mean ‘non-human suffering is neglected relative to the moral stakes’; disagree voters saw it as ‘neglected relative to common sense views’’. This was splits was visible in some panel responses —
[Agreeing] ‘It seems extremely likely that most [...] suffering is experienced by animals other than humans. Non-human mammals alone outnumber humans by at least an order of magnitude. But it seems like a tiny minority of AI safety researchers are taking the risks of AI for non-human suffering seriously.’
[Disagreeing] ‘It's certainly not neglected relative to human views generally - it's relatively highly overrepresented, and it's unlikely to matter if we don't manage alignment in general.’
Given that CaML’s positioning is very strongly in favour of more non-human representation in AI Safety and development, the lack of consensus here represents a challenge for us and similar organisations: if we want to increase representation of non-human interests in AI safety, how do we demonstrate the importance of this work? This is something we hope to emphasise more in forthcoming research on scaling compassion, and our animal cruelty benchmark currently in development.
As above though, there was strong support for instilling compassion as a value within AI systems. CaML is specifically interested in broad compassion that is robust and generalisable across different classes of sentient beings.
The panel was more uncertain about deceptive AIs and mech interp
The community’s second strongest result overall was that 84% agreed that deceptive AIs will evade mechanistic interpretability tools to hide their features from humans and operators. The panel was more evenly split on this (lean of +0.4, SD 2.6), with some respondents saying it would depend on specifics:
‘Only if we apply strong training pressure or give very large affordances over training (i.e. successor training). With reasonable precautions I don't think mech interp is adversarially difficult’
‘This is hard to say and would depend a lot on specifics’
‘This will be a cat-and-mouse, evolutionary-arms-race-style dynamic’
Vote distribution for `Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools''. Blue bars are community votes; coral dots below the baseline are the 14 individual panel responses. The community mass sits largely on the agree side (84% above neutral), while the panel is less confident.
Caveats
As mentioned above, these polls should not be interpreted as an exhaustive analysis of where the alignment community stands. We wanted to share the results primarily to encourage further discussion and work on these questions, as well as get feedback on CaML's research direction.
A few caveats on the above findings:
The researcher panel is likely more interested in non-human welfare and s-risks than the broader alignment community. Some of the panel responses therefore may not be entirely representative of alignment researchers generally when it comes to their views on non-humans and s-risks.
A similar selection effect may be true of the community survey, which were only completed on the EA Forum. That said, we intentionally promoted the poll in wider alignment circles, including through the ACX Open Thread and on LessWrong; responses to some questions (e.g., on non-human neglect) suggested a diversity of interests and prioritizations. For future surveys, it would still be beneficial to conduct them through a separate form to receive respondents from EA Forum, LessWrong and other online spaces equally.
There was disagreement on how questions should be interpreted. We received a few comments mentioning how respondents interpreted certain terms (‘suffering’, ‘post-AGI’ ‘consciousness’, ‘neglected’), which would alter how these questions were answered. This is at least in part an indication to tighten our wording and scope for future polls; it also seems indicative that confusion over terms is leading to more disagreement and a difficulty in reaching consensus within the community. Both we as a team and the field as a whole must coordinate on our definitions, so that we can more effectively express where we stand and what matters.
We didn’t distinguish between neutrality and uncertainty. When encountering a question they weren’t sure on, it’s possible that some respondents would just skip the question, while others would select the middle value. Aside from comments, we don’t have a way of distinguishing these cases from a well-informed respondent who still has a middling credence.
Again, we’d be excited to see a more rigorous and in-depth study of opinion in the alignment space, similar to what Lucius Caviola has done for digital minds here.
Conclusions
Over the next few years, the alignment community will have to move quickly and respond to changing circumstances. Having a clear representation of what the community values and believes will be crucial in coordinating effectively, making convincing arguments to external stakeholders, and pushing on new research directions in areas of uncertainty.
This survey has suggested support for compassion to be instilled within AI systems, but also disagreement about whether non-human interests are neglected. CaML will continue to produce research to inform how and why compassion can be instilled within these systems, and make the case for why that compassion should be broad, accounting for the welfare and interests of humans, animals, and potentially sentient AI systems.
We’d like to thank BlueDot Impact for supporting this work, and Scott Alexander for sharing the community poll with the Astral Codex Ten audience.
Compassion Aligned Machine Learning (CaML) is conducting research and deploying benchmarks with the aim of implementing broad compassion for all sentient beings into AI systems. In so doing, we’re attempting to integrate ideas from across the broad field of alignment research, including those focussed on specific technical questions, and those considering more strategic questions around who or what AI should be aligned with. To do this, it's important for us to understand how the alignment community is conceptualizing and prioritizing different ideas that CaML cares about.
In this spirit, we recently conducted a survey on controversial questions in alignment, to encourage more discussion on topics important to CaML, and roughly gauge where both core researchers and wider community members stand. Questions ranged from the value of benchmarking as eval awareness increases, to the implications of AGI for animals, to thoughts about AI s-risks. We’ve shared some interesting findings from this survey below.
Our intent with this survey was not to conduct an exhaustive analysis of the community’s position on every issue. This was an initial probe primarily designed to generate discussion and feedback on issues important to CaML. That said, we’d love to see research more accurately tracking core areas of (dis)agreement, consensus and controversy within alignment research. If we know where there is uncertainty, we can better understand where our research should focus next.
What We Did
Our survey consisted of two stages.
In the first, we shared 14 polls with a panel of 14 researchers, where they privately responded to statements on an 11-point agree-disagree scale. They could also share the reasoning behind their answers. We received responses from the following 14 researchers:
Please note that respondents took these polls in an individual capacity, and responses do not represent the views of their employers
In the second stage, we released the survey totalling 15 polls (with one additional poll to the panel survey) to the broader community. Respondents answered via an EA Forum post (where all the polls are still viewable), and could publicly share their reasoning for different responses. After feedback from the expert panel, we included clarifications alongside the polls. We received a total of 1444 votes over all 15 polls, and over 100 comments providing further reasoning.
Once the survey was completed, we analysed two primary variables across the 15 questions: the lean (whether respondents on average voted agree or disagree), and standard deviation (SD) to see the spread of responses on a given question. Lean is measured from -5 meaning strongly disagree to +5 meaning strongly agree (with 0 being neither agree or disagree).
What We Found
Collective lean per statement, sorted by community lean. Grey connectors join the community and panel means for the same statement; rows where the dots sit on opposite sides of the dashed line are statements on which the two samples lean in opposite directions. The right-hand column gives the standard deviation of responses for the community (C) and panel (P). *Asked to the community only.
Most People see the s-risk field as underperforming
In two polls, we asked explicitly about the value of S-risk work:
On both of these questions, the community leant towards disagreeing with the statement (lean of -1.2 on poll six, and -2.0 on poll eight). Only 3 of 73 respondents agreed that s-risk progress was on track and not one strongly agreed (lean of more than 1.5).
Poll six saw the biggest distance between the community and the expert panel, with the panel leaning agree (+1.2) that s-risk work was more valuable at the margin. The panel was somewhat split on this though: this poll had the highest standard deviation (3.2) of the set. This may have been somewhat representative of the different research interests of the panel, with some explicitly more focussed on s-risks than x-risks or vice versa.
Nevertheless, both the panel and the community disagreed that s-risk work was on track. Some text responses gave more detail as to why:
Regardless of how respondents prioritise s-risks, there was at least agreement that, as it stands, what progress actually looks like in the s-risk space is difficult to identify, be that due to the field’s far-future focus, small size, or general uncertainty. This suggests a primary challenge for those currently focussed on s-risk research: identify and communicate how progress could be measured in concrete ways in the field, and consider what robustly good things can be done under uncertainty.
People want compassion promoted in AI systems
The final poll in the community survey asked ‘Insofar as alignment continues to promote overlapping sets of characteristics (e.g. helpful, harmless, honest, corrigible), should compassion be one of those characteristics?’ This question was not shared with the panel, but it had the highest rate of agreement on the EA Forum of any poll: 89% of respondents agreed (lean +2.4, SD 1.9). Few voters articulated their reasoning on this question, though some responders did add that care must be taken in how this value is conceived of and pursued:
This poll points directly at CaML’s core mission and suggests broad support among respondents for the kinds of values we’d like to instil in AI systems. What we need to show to labs and the alignment community now, is how broad compassion can be implemented alongside other important values in a way that is safe and robust.
Vote distributions for three featured statements, on the same -5 (strongly disagree) to +5 (strongly agree) scale. Blue bars are community votes; coral dots below the baseline are the 14 individual panel responses (the compassion statement was community-only). The shapes show the findings: no community votes above +1.5 on s-risk progress, two opposing camps on non-human neglect, and the majority leaning agree on compassion.
It’s controversial whether non-human suffering is neglected by AI Safety
‘Non-human suffering is neglected by the AI safety community’ has the highest standard deviation among the community polls (3.1). In agreement, one respondent said: ‘as I understand, many key players are explicitly anthropocentrists. Hell, some key players care more about money than humans, let alone non-humans.’ In disagreement, another respondent said work had already been done to make Claude care about non-human suffering, and that it wasn’t even clear if this was a valuable direction to pursue (due to potential backfire effects, and the generalisation of human alignment to non-humans). Some of the disagreement here may have been due to the respondents’ reference point: those voting agree were more likely to mean ‘non-human suffering is neglected relative to the moral stakes’; disagree voters saw it as ‘neglected relative to common sense views’’. This was splits was visible in some panel responses —
Given that CaML’s positioning is very strongly in favour of more non-human representation in AI Safety and development, the lack of consensus here represents a challenge for us and similar organisations: if we want to increase representation of non-human interests in AI safety, how do we demonstrate the importance of this work? This is something we hope to emphasise more in forthcoming research on scaling compassion, and our animal cruelty benchmark currently in development.
As above though, there was strong support for instilling compassion as a value within AI systems. CaML is specifically interested in broad compassion that is robust and generalisable across different classes of sentient beings.
The panel was more uncertain about deceptive AIs and mech interp
The community’s second strongest result overall was that 84% agreed that deceptive AIs will evade mechanistic interpretability tools to hide their features from humans and operators. The panel was more evenly split on this (lean of +0.4, SD 2.6), with some respondents saying it would depend on specifics:
Vote distribution for `Deceptive AIs will be able to hide unwanted behaviours from mechanistic interpretability tools''. Blue bars are community votes; coral dots below the baseline are the 14 individual panel responses. The community mass sits largely on the agree side (84% above neutral), while the panel is less confident.
Caveats
As mentioned above, these polls should not be interpreted as an exhaustive analysis of where the alignment community stands. We wanted to share the results primarily to encourage further discussion and work on these questions, as well as get feedback on CaML's research direction.
A few caveats on the above findings:
Again, we’d be excited to see a more rigorous and in-depth study of opinion in the alignment space, similar to what Lucius Caviola has done for digital minds here.
Conclusions
Over the next few years, the alignment community will have to move quickly and respond to changing circumstances. Having a clear representation of what the community values and believes will be crucial in coordinating effectively, making convincing arguments to external stakeholders, and pushing on new research directions in areas of uncertainty.
This survey has suggested support for compassion to be instilled within AI systems, but also disagreement about whether non-human interests are neglected. CaML will continue to produce research to inform how and why compassion can be instilled within these systems, and make the case for why that compassion should be broad, accounting for the welfare and interests of humans, animals, and potentially sentient AI systems.
We’d like to thank BlueDot Impact for supporting this work, and Scott Alexander for sharing the community poll with the Astral Codex Ten audience.