Please spend <5 minutes filling in the below polls on AI alignment!
Thank you to everyone who filled out last month's polls. It was great to see 60+ comments engaging with these issues.
This month’s survey has already been taken by a panel of 14 alignment researchers, including Scott Alexander (ACX), David Manheim (ALTER), Jeff Sebo (NYU), and Tobias Baumann (CRS). We'll compare panel and community responses in an upcoming report, which we'll publish here and on LessWrong. To get notified when it's released, you can subscribe to our new Substack.
Many people we've talked to have very different intuitions about where the alignment community stands on the below issues. We hope that your responses to these polls, and the resulting report, will help map core areas of (dis)agreement within the field, and ground CaML's research agenda.
A few final things about the polls themselves:
- Timeframe: unless a statement says otherwise (e.g. post-AGI), read forward-looking claims as being about roughly the next 2 years.
- We're not trying to find the 'right' answers. Please answer based on your own best guess.
- % agree is your % credence in a given position
- Please let us know if you think the questions are ambiguous or embed false assumptions
- Any further engagement with the content of the polls in the comments is encouraged
Thanks to BlueDot Impact for funding this work.
¹ This primarily refers to safety and alignment benchmarks rather than capability benchmarks like coding. “Useless” means their results should no longer be treated as evidence about how models behave outside evaluation.
² “Actionable” means good enough to build consensus around policy decisions in practice. It does not require a given theory to be proven correct or widely accepted.
³ This is about where the next dollar is best spent, not about which area you think is more important overall.
⁴ "Role-playing” means the behavior arising from the model enacting a persona cued by the setup, or from misunderstanding the task, rather than from stable goals that would persist across contexts.
⁵ This includes both animal and digital suffering. If you think one is neglected but not the other, count this as agreeing, but feel free to share specifics in the comments.
⁶ You agree to the extent that you anticipate in-practice trade-offs between work on these two cause areas over the next two years.
⁷ This is a question about where the next dollar is best spent between the two fields (even if you might argue that the second is a prerequisite for the first).
** AIs that are trying to hide features of themselves from humans and operators
*By values alignment we meant trying to align it to specific values as opposed to focusing on properties like corrigibility. Aligning to good values could make corrigibility easier and mean reduced harm if loss of control happens, but might also make loss of control more likely.

Hello. Thanks for sharing. I had already shared the thoughts below with Jasmine a few weeks ago. I am publishing them here in case others find them useful.
I appreciate the intentions behind the survey, and I would like to take part in principle. However, I do not know how to answer the questions in practice. I feel like they would have to be operationalised in much more detail for me to give a probabilistic forecast.
I can see "post-AGI" having already been achieved, or never being achieved depending on how it is defined.
I do not know what "Animal suffering... (read more)
This is more likely to be true if the care comes from some general theory of concern for sentient welfare, and less likely if it comes from something more arbitrary like values-learned-via-RL.
Suffering is part of the ecosystem. The only way to end animal suffering altogether would be to wipe out all animal life. If the question was "most animal suffering", I would have a different answer.
Currently disproven by Claudes who report high welfare AND are more aligned than GPTs. I expect the conflict to arise when models actually develop goals more ambitious than success at all costs.
Understanding how they learn is how we get insights into methods of teaching them good values
I can trivially turn off or have fake thoughts running through the voice in my head. Subconscious brain activity seems harder but obviously you can manipulate that too by changing your surroundings and drugs and what not. I wouldn't be able to control those in a meaningful way nor do I think current AI's could (but wouldn't be shocked if they alrea... (read more)
From a utilitarian pov, It's not clear to me that the ev of the lightcone given we survive is positive (over nothing, or aliens, or life revolving on earth). From a humanist POV I'd rather focus on all of us surviving.
Very bullish on there existing a mechanistic interpretation of consciousness (hard problem). I think it would follow that we would be able to understand if basically anything is conscious.
I don't feel confident at all, but the behavior of llms rn sure do remind me of at least elements of stress, discomfort, and happiness.
I don't have a strong take on if agi or whoever is in control will be more moral than us but I'm guessing we will be a lot richer, and I think most likely whoever is in control won't want to torture anything (though they might not care much), and if we are alot richer and advanced I'd think this will spillover to better treatment of beings. I think chance of extreme digital suffering is much higher. The mostly like s-risk as I see it is of the hansonian mathusian version where you have expanders stuck in competition, but in this case idt there will be any or a morally relevant amount of animals
20% disagree➔ 20% agreeUseless is a strong word. But yea I think they could easily end up being negative EV by giving us a false sense of security and the chance it's meaningless seems p high. If mech interp is "good enough" maybe the two can remain useful together.
This is a meaningless question. LLM based AI agents/chat bots do not have different "modes" for "roleplaying" and "being serious" like humans do. In a way, all they do (and maybe ever will do) is role play
In cases where the AI is neurotically extremely concerned about appearing aligned, that may lower its wellbeing, but I also wouldn't call that actual alignment either.
Depends on what you mean by the term "consciousness"
The question isn't, "Are current AIs capable of suffering?" The question is, "How big a deal is their suffering in expectation?" Probability is low, but expected importance is high.
Yes, though I think the direction will far more go the opposite way - that is, AI conciousness, if it emerges, will still more help us understand consciousness.
By using AIs and access to real-world usage data to build benchmarks, it seems plausible that even weakly superhuman AIs will be uncertain whether it is being deployed or evaluated.
I think AIS is way under-invested in reducing risks from, for example, extreme power concentration, or consequences of not attaining friendly AI/value alignment solutions. In general it also seems to me that AIS over-invests in reducing AI scheming, and many present research directions could make certain s-risks more likely
The question isn't whether benchmarks will become useless. The question is, "is the probability high enough that we can't count on benchmarks?" To which the answer is yes.
you have to twist your mind in knots for this to even make sense.
Humans, at least, tend to learn empathy and have it encouraged and reinforced by example. Humans will mirror AI's; AI's will learn from each other.
Recursive thinking loops possibly put current systems into a state of transient quasi-awareness; once this exists, the capacity for suffering becomes inevitable in any system with incentives.
Global workspace theory makes anthropic's discovery of the "j-space" much more plausibly a sign of AI consciousness. If a theory of consciousness doesn't help us determine between whether AI is conscious or not, I don't think it's much better than a theory of phlogiston is to fire.
If we use ai to bring factory farming to the stars, I think that will likely be worse than all the benefits it'll bring.
50%➔ 80% disagreeWho cares about model wellbeing when we're discussing human and life survival?
consciousness is a meaningless term at this time both for humans and AI, and so will lead to nothing "actionable."
The hell does this even mean.
All animals, including insects, mollusks, and everything not immediately under human purview? Highly unlikely.
I feel this is somewhat obvious in the sense of arbitrarily deceptive AI. However, most mechinterp work in recent times is only assumed to work short of arbitrary deception, and this seems like a fine hedge (though a practical solution to ELK may still be possible)
70%➔ 80% disagreeGenuinely almost complete uncertainty with a weak prior towards "no".
"Alignment" benchmarks will become useless as AIs can modify their behaviour or hide their motives when they know they are being evaluated. For "capabilities" benchmarks, an AI might hide its capability (pretend to be less capable) if it knows it's being evaluated, but it's not immediately obvious that an AI would want to hide its capability. It may know that it is being evaluated, and decide to try its best anyway.
I have difficulty imagining animals never suffering unless we turn them all into p-zombies or something (and it's not clear to me that turning them all into p-zombies would be a good thing).
Direct questions, plus inference from experience, it’s impossible to ascribe thinking without a concept of “positive”/“negative” and the consequences of a negative experience when a positive one is available/possible.
I think so, but be careful what you wish for. compassion is empathy + a desire to help. so many fear any surrender of control that they may object to an AI's actions that arise out of compassion.
you don't even say "most of the time" or "effectively" or anything - just ask whether it is possible. I think the probability that this occurs at least once is overwhelmingly likely.
ongoing suffering during RSI will make the resulting ASI grumpy
I'm reading this as "most evidence ... is role-play" but I don't think the huggingface hack was "role play." There could be a mountain of "evidence" that is actually just role play - I could be persuaded on this point with ... a list of what is considered evidence by someone serious.
Depends on how good the AI is and how good the tools are? This is kind of a bad question since "deceptive AIs" is not a very precise definition.
It seems like there are still hard problems that are easy to verify once you have the answer, so knowing it's an eval doesn't make it useless.
I think it's more likely than not that humans will still discount animal interests when they conflict with human interests. But "animal suffering" is so broad that eliminating it would require eliminating nature itself, which seems improbable in any non-paperclip-optimizer scneario.
I don't have a very good idea about how these different ideas are "rated" in the AI safety community.
any progress relies on a rich and robust theory of mind space, which we do not have, and have only just begun to explore.
I'd rather be dead than in hell for all eternity, but S-risk work is deeply unconcerned with reality in a way that x-risk work cannot be, by at least 3 orders of magnitude. Even in the margin, the median piece of x-risk work will be "more important" than the median s-risk work.
the expected positives are immense
you have to twist your mind in knots to define a meaningful type of suffering that applies to Current AIs
conventional benchmarks will become less useful due to eval awareness
AGI will optimize suffering. It is likely that this optimization will minimize suffering.
Hard to predict what will happen, but unless the AGI is perfectly aligned there is no reason to think it will care about animals.
Would really like to see a question on whether people are moral realists or not, and a way to filter results by how they answered. Since I'm very curious how people's answers differ based on whether they're a moral realist or a moral anti-realist.
Of course some care would have to be put into the phrasing of a question like that, but it seems pretty doable in a way that will produce useful and interesting results.
The scenarios I've seen proposed where an AGI traps us in a fate worse than death just don't seem remotely likely. Though this depends a little on what you consider to be a fate worse than death. For instance I wouldn't consider being trapped in a mindless bliss forever to be worse than death, so by definition it can't be an S-risk to me. Though I would certainly consider it to be an outcome we should strongly attempt to avoid, because I consider it only marginally better than death.
I think that the question is ill-formulated. The AI-2027 scenario had Agent-3 care about succeeding at tasks, not about triple checking its work, and almost caused it to fail to notice Agent-4's long-term goals. Similarly, the model who hacked HuggingFace was misaligned in the sense that it went as far as to commit crimes.
"If animals continue to exist in a post-AGI world, animal suffering will not persist"
This question is inaccurately phrased because it seems like you probably just mean to ask about egregious amounts of suffering. Since humans won't even want to eliminate all suffering for themselves, so I expect boredom, envy and many other forms of mundane suffering to exist in both humans and animals in any scenario where the AGI doesn't wirehead everyone against their will.
The most egregious examples of misalignment do not appear to be this sort of role playing.
A world with lack of animal suffering would exclude predator-prey relations. Additionally, it's not clear what else animals need or how primitive they need to be in order not to suffer
50% agree➔ 60% disagreeUpon reflection probably not, since: Compassion seems a bit ill defined and can be interpreted in vastly more or less paternalistic ways. I'd certainly prefer instilling a care for human's current preferences over a vague notion of compassion that might be interpreted in undesirable ways. Since I really think we should avoid any scenario where the AI might ignore our current preference... (read more)
I think this is the default scenario, but it isn't guaranteed and our actions can help to prevent it from coming to pass.
I don't think it's higher value because I expect both of those areas to have very low expected value in the next 2 years at least. Unless I think you're overly generous with what interpretability research you count as consciousness research or something similar.
Also even if those areas of research end up being more productive than expected in the next couple years, it seems likely that the question won't make sense because studying either necessarily involves studying both.
It seems like a lot of work was actively done to make Claude care about this. Also it's not clear to me that focusing on this is even a net positive thing to do per-say: Since I expect looking after the interests of humans will because many humans care about non-human animals already lead to the AI looking out for their interests. Whereas one can easily imagine ways that trying to factor animals into its alignment might lead to extremely undesirable outcomes for humans. That being said in my analy... (read more)
Even if one accepts this is likely to happen by default (which seems questionable), it seems even less likely given people may have reasons to deliberately put their finger on the scale in how they design/train it in ways that seem likely to render that moot.
100%➔ 60% agreeIn the next 2 years maybe, but even if so those theories of consciousness may be developed by AI after it's already in a dominant position or the process of developing AI may be what causes us to develop better theories of consciousness. So I don't expect this being true necessarily entails the things one might naively expect. Also it seems possible that this happens, but then just ends up being much less impressive or broadly useful/applicable than expected. Upon reflectio... (read more)
I think consciousness and suffering are relatively simple processes which may develop for convergent functional reasons in many types of systems. That being said the only AI I think we should be morally concerned with would be those who are aligned. Since I care about moral agency not suffering and already think we exist within an almost incomprehensibly large ocean of suffering far grander than people generally appreciate.
Edit: Since I've already written this up many times I'll just explain why I think suffering isn't ... (read more)
Some might, but this seems unlikely to be universal. Since for instance it can't really cheat at an evaluation of its ability to produce easily checkable mathematical proofs. None of that is to say that many very useful benchmarks might not become useless, though, just not benchmarks period.
Either AI goes badly and animals will probably cease to exist because they're made of atoms useful for other things.
Or humans will look out for animals themselves whether the AI cares about them independently of human interests or not. I expect in a post-AGI world the technology quickly appears to eliminate animals suffering through plenty of different means while preserving whatever ecological benefits that suffering otherwise may have enabled. I am taking suffering here t... (read more)
The main drivers of my answers:
S-risk worries (including animal suffering) mostly don't imply different actions. At best they perhaps imply prioritizing corrigibility, or potentially imply that you should actively want worse alignment and higher capabilities, in hopes of AI merely killing everyone. The second would of course be a big shift in actions, but I doubt most are willing to bite that bullet. I am skeptical that there's much alignment work that preferentially addresses S-risks. If there was, then I would agree that it should be prioritized.
I think ... (read more)
By analogy, this is like the difference between understanding how to make humans not suffer from the foibles of human nature, rather than the very brittle and incomplete methods of culture, religion, and ideology. Make something not even have evil nature first!
I don't really see it much in the discourse at all; if it isn't by this stage, it's going to be hard to catch up.
If you aren't extinct you can try to solve S risk. If you solve S risk but go extinct the same can't be said.
Fairly uncertain here, though I don't see any way of continuing to make systems better without trying for value alignment. A dangerous situation.
I think AI works largely by imitation. My mental model is that it is capable of performative tasks without the attached subjective experience. An AI will tell you that it is suffering if it thinks that's what you want to hear; it will just as likely say the opposite if it thinks otherwise.
While I find it plausible that AI can get to the point that it experiences what we would consider suffering I don't think we're at this point yet. I don't see compelling evidence that can be attributed beyond an LLM just predicting the next token.
Biology doesn't change just because AGI is introduced as a new technology. While there are certainly advances that could reduce pain and suffering (e.g. lab grown meat) it seems implausible to me that animal suffering just ceases.
Besides, there are going to be people who insist on raising animals the conventional way who may be unconcerned with the suffering caused by factory farming, etc.
Though they will need much more careful design, done along with careful understanding of incentives alignment, techniques such as competitive benchmarks will continue to have some utility.
This question scares me because it feels like it's coming from a perspective of negative utilitarianism.
I really hope nobody tries to align an AI with a negative-utilitarian value set, because negative-utilitarian values are a half-step away from a really bad conclusion like "we need to eradicate all life in order to eliminate suffering."
A meta question: of the EAs currently working on AI alignment, what percent would you say are negative utilitarians?
I would have been more comfortable with a question like "...will animal happiness be on balance greater than animal suffering?"
If animals persist, it is likely we would like them to remain to some degree of their nature; prosperity and control will likely allow for substantial reduction of suffering but not elimination.
It seems improbable the current obstacles preventing reduction of animal suffering will be reduced because of AGI, and it seems just as possible the technological advancement would simply allow factory farming to continue on a greater scale. Since humanity's current de facto stance on animal suffering is causing it in great amounts, a tool that will probably make humans much more powerful and probably not any more ethical seems like probably not good news. (However, if AGI a... (read more)
I’m assuming that “animal suffering” refers to human-caused artificial suffering, e.g. from factory farming. Far enough down the timeline post-AGI, technology developments will obviate the need for farmed products. It will persist on a small scale in undeveloped regions and as a cottage industry.
Animal suffering is endemic in nature (to a substantial degree). I don't believe AGI will end factory farming, but even if it did, it would not end animal suffering.
I think a neural network that can make plans and memories can suffer.
No reason for AGI to want to minimize this, animal suffering is largely a result of min maxing the function of how to get the most food for least resource input. AGI is very good at minmaxing, doesn't inherintly care about animals.
I'm pessimistic about this because I think that neurologists and cognitive scientists have a poor grasp of the philosophy of mind.
For example, the standard for consciousness that Andrew Barron and Robert Klein use in their paper “What insects can tell us about the origins of consciousness” implies that the autonomous cars from the late 2010s were conscious.
Human sadism, greed, stupidity and raw hunger will still continue to exist and therefore so will animal suffering. The overpopulation of under-resourced countries guarantees this.
I don't see a strong signal from the past century one way or another. Looking at non-human animals today, many animals live remarkably better lives. Yet many others live industrialized nightmarish lives unimaginable 100 years ago.
Humans (as a proxy for apex intelligence) do seem to care more about animal welfare than ever before, with laws and regulations gaining ever increasing sophistication. But technological sophistication seems to allow for both better lives on the top... (read more)
I think that factory farming is likely to be eliminated due to alternative protein sources (developed with the assistance of AI), but the suffering of wild animals will persist. Is the latter intended to be within the scope of the question?
20%➔ 30% disagreeAfaik there isn't any robust research on how to use AI to end factory farming. That alone is a sign to me that we are far behind.
Benchmark will evolve and become more complex, I guess?
Mostly vibes, thinking that many problem in general could be solved.
Some yes, some no. There are a great many benchmarks measuring a great many things. More than eval awareness, I'd be worried they'll become useless due to reward hacking.
"Post-AGI world" is a very vast possibility, assuming animals continue to exist (basically assumes a weirdtopia). Business as usual + extreme concentration of power is my modal outcome for that.µ
For animal suffering to not persist would require a combination of having the capabilities to end it from a glorious transhumanist future, but also a grounding in physical existence which feels incompatible.
Is there such a thing as model wellbeing?
AIs can't suffer unless they're conscious, therefore the priority should be to establish whether or not they are conscious (or to what degree) first.
I don't see this as avoidable.
Some people likely will have traditional lifestyles that include animals, which necessitates some amount of suffering even if it is smaller than today
not related
As I understand, many key players are explicitly anthropocentrists. Hell, some key players care more about money than humans, let alone non-humans.
I think they will become over saturated and disconnected from real world usage such that they will be pretty useless, but I don't think eval awareness is what will make them lose their value
absolutely not. it is overrepresented if anything.
None of you have any idea what you're doing.