Researcher at the Center on Long-Term Risk. All opinions my own.
That seems like a big leap to me, and I donât see how it follows or what justifies it
The sequence gives general arguments for this, especially sections 2.3, 3.2, and 4.1. I'm not exactly sure what you find uncompelling about them. The worry is that severe coarseness makes the degree of justification so severely vague that we have all the same qualitative problems as under the maximality analysis.
Of course, the claim is defeasible by arguments to the contrary for some specific A vs. B. But the burden of proof seems quite high to me.
I think no. Basically, when I really internalize how dwarfed every action's cosmic-scale consequences are by off-target effects, "no" feels very common-sensical to me. (Cf. this paper on how "simple cluelessness" is fake.)
I think I wouldn't be clueless about c-preferability in Elliott's "trapped in a box" example, but can't think of any real-world case analogous to this.
Thanks!
I donât think the arguments against the capacity-building strategies I discuss are as strong as those in their favor
My core objection to capacity-building in the sequence is: For any concrete capacity-building strategy, our understanding of that strategy's full range of possible consequences is extremely coarse. And it seems very plausible to me that these consequences will include large off-target effects on, e.g., lock-in events â in which case, capacity-building strategies inherit the non-robustness of strategies aimed at influencing lock-in events. This is for pretty similar reasons to how the off-target effects of AMF donations seem to dominate. I don't yet see why you think otherwise.
(So in particular, I don't think your responses in your appendix to specific backfire risks I mentioned in the post address this core objection.)
Moreover, much of the value of building capacity is the value of being able to act on considerations we arenât yet aware of. So, in my view, unawareness bears asymmetrically on these strategies rather than neutrally (I realize this latter point is stated very briefly and needs further development).
Yeah, I'd be interested in seeing this spelled out a lot more sometime. Per the above, even if a strategy might enable us to act on considerations we arenât yet aware of, this doesn't help us with cluelessness if the strategy's impact is still very plausibly dominated by off-target effects.
Iâm one of the judges of the competition. My comments shouldn't be taken as a full review of a post. And, unfortunately, I wonât have capacity to comment on every post or engage with all replies. Thanks so much to everyone who entered! Â
I'm confused by your responses to the vignettes and thought experiments you quote in this post.
(Just to be clear, that's a contingent matter. I don't find any of the counterexamples offered so far persuasive because I don't think they adequately engage with my arguments for P3.)
Yeah I still disagree for this 4th example as well, for the same reasons as the 1st.
I worry about a motte-and-bailey implicitly happening here, where the motte is "all things considered, we should prefer the second action" â very difficult to deny! â and the bailey is "we should c-prefer the second action". This matters because the corresponding motte seems actually pretty easy to deny (IMO) when the comparison is "study some altruistically irrelevant branch of academic philosophy" vs. "try to prevent AI misalignment". The latter only looks clearly preferable to me if it's c-preferable. (I guess this is what Ben's comment is getting at.)
Some of the contest entries do at least briefly attempt that, I think. E.g. this post gives an argument that one should c-prefer "low-footprint capacity-building" over "doing nothing". And this post more generally argues that "donate $5 to Make-A-Wish Foundation" is c-dispreferable to some mixed action (maybe that's not specific enough for what you have in mind). (I don't yet buy either of these arguments, though.)
Alternatively, if you posit incommensurable values then you should probably reject P1.
It's a fair point that deference principles come into conflict with prospective reasons in Hare's case, and the prospective reasons argument seems really plausible there. I don't feel very confident, but here's how I'm thinking about this:
In particular, this framing doesn't require that humans even have well-defined "hypotheses" in our epistemic state. I don't think the alternative framing I use in the sequence itself requires us to have unrealistically precisely defined hypotheses, either. But this was a bit of a sticking point when discussing the problem of unawareness with one thoughtful interlocutor â which is (AIUI) what led to Jesse writing the post on deference to the idealized self that I'm drawing on.
I think this is a fair point. But here's where I was coming from in the quote you respond to here.
First, in that context I was taking the fundamental contrastive reasons-givers to be person-moments, not individual persons. From that perspective, the problem is that we do have long-run contrastive reasons for each option, which we don't know how to weigh up. I'd agree that if the contrastive reasons-givers are persons, each person whose welfare we're clueless about gives us no contrastive reasons.
Second, I think we should separate two claims:
If the contrastive reasons-givers are persons, I agree with (1). But when I expressed doubt about the "moral urgency" of bracketing, I was doubting (2). To my metanormative intuitions, it just doesn't feel like a bracketing-based verdict in favor of A has as much weight as the analogous EV-based verdict. (I could try to say more on why, if that's helpful, though it's hard to articulate.)
Still, I find it hard to say whether bracketing-based verdicts have more weight than, say, "precise consequentialism says I should instead do [galaxy-brained thing]". Something that helps me probe my intuitions here is: