Some quick observations on the attacks on METR, EA, and other AI safety orgs (from my perspective as a political campaigner and former Communications Director for a union and mayor during COVID):
If you're explaining, you're losing: Every rebuttal repeats the accusation to people who haven't heard it yet. Answer the accusation once, link to it, and stop explaining. This is the "illusory truth" effect (information is true when it's repeated).
Silence doesn't work either: Not responding on the record suggests something is being hidden. Make one clear statement rather than amplifying in multiple twitter threads.
Concede what you can: A concept I know from political comms is the "admission against interest". This buys trust, disarms cynical voters, and establishes credibility. Saying "yes, AND here's what we're doing about it" will take some air out of the critique.
Messenger > message: Ideally people who know the organization being attacked, are from outside the AI safety world, have independent standing, or better yet, have disagreed with the organization publicly will land the defense better.
Going after people's personal lives or spouses is overreach: Ordinary people will obviously find this unfair and it's where their attacks will backfire.
Don't quote tweet and rebut: You're just distributing their message for them.
I want to write something longer on this but don't have time right now. Lots of opportunity for good crisis comms here. Let me know what's wrong about this take and if you'd like to hear more!
1 / 4 / 6 feel like they might make sense strategically, but I'm proud to be part of a community that often values epistemic humility and correct argumentation above strategic communication.
Points #1, 4 and 6 aren't asking anyone to really say anything differently, it's more about who says what, where, and when. If someone from outside the AI safety world (or better yet, someone who can show they've disagreed with METR before or are not a natural ally) says "the corruption story isn't true," that's not less epistemically correct, it's just likely more believable to the general public.
I'd gently push back on the bigger part of your argument though: The public conversation can't end up being about METR's independence, it should be about an outside entity checking the lab's work. That seems to me the tradeoff if we lose the bigger comms battle.
Says "if you're explaining you're losing". IMO explanations are good. You might be correct that needing to explain is a bad sign for whether you are winning a comms battle. But I would still rather offer people genuine explanations, so we can come to a shared understanding of the truth (exactly like this convo).
My reading of this is that you're claiming who says something is more important than what is said. Again, you may be correct here on how to optimally persuade. But I personally want to value good arguments, no matter who makes them.
Feels like it is straightforwardly asking folks to not engage in debate or disagreements. I don't really get how we come to a shared understanding of the truth if we don't offer clear reasons for why we disagree with particular points.
Again, I very much am not saying you are wrong about the implications of ignoring strategic communication norms. You are likely right. I would nonetheless prefer to be in a community that decided to communicate honestly and earnestly, even with people that it disagrees with.
If you're explaining, you're losing: Every rebuttal repeats the accusation to people who haven't heard it yet. Answer the accusation once, link to it, and stop explaining. This is the "illusory truth" effect (information is true when it's repeated).
Is that actually true of X's reply mechanism to a significant degree? My impression was that replies on X are mostly just seen by people who tap on the original post. Replies wouldn't necessarily have the effect of amplifying the original post?
Certainly writing Community Notes would not present much risk I presume??
I meant "if you're explaining, you're losing" for any platform, not X replies specifically.
I agree Community Notes are very likely net positive (they certainly have an impact on me). But where I still see some risk is volume: Replies are a signal and if a lot of people pile on to rebut a post, perhaps that pushes it to more people? (I don't know X's algorithm well enough to say that with much certainty.)
I also don't want to double down too hard on illusory truth. I defined it a bit too loosely in my quick take, and there are follow-up studies showing a good correction usually wins with the people who read it.
For me it's more about Zaller's Receive-Accept-Sample (RAS), which I find helpful for thinking about mass/political comms. Voters have to:
Actually hear or read the argument,
Accept it as fitting their frame or worldview, and
Have it available to sample when they're about to act or make a decision.
Continuing to respond to a critique that's wrong makes it more likely people 1) hear it and 3) sample it later.
Thinking aloud here. I don't think you necessarily have to respond by quoting the person.
You can just provide a frame or facts which counter or inoculate against whatever critiques are currently widespread. E.g. if someone accuses you of a crime, you can mention that you were someplace else at the time the crime was committed, without directly repeating the criminal accusation. However, your X replies might fill with accusation talk in that case.
Wanted to share a quick update on my comments on AI safety comms from a couple weeks ago. The NY Post is now reporting that the FTC is investigating OpenAI, Anthropic and METR, with formal demands (similar to subpoenas) reportedly coming in the next few weeks.
I think Brian Martin's Backfire framework could be useful in responding here. Martin found that power holders use five methods to suppress outrage, and there's something we can do in response to make it backfire:
They cover it up → We reveal: Most of what they do will happen out of the public view (though they seem to be more emboldened and this is less true). We can document their actions starting with the narrative they've been advancing on X to the Post stories to the eventual probe so the general public can see a pattern.
Devalue → Redeem: They've called METR a "handpicked watchdog," they've branded EA a cult, they've attacked Lighthaven etc. We can counter their attacks using people outside of the community, who the public respects, vouching for METR's work.
Reinterpret → Reframe: Their official line is that this is neutral fact-finding. We can name what's actually happening: they're investigating the group that published an independent report on AI agents getting out of a test environment because they want to continue accelerating A(S)I development.
Send it to official channels → Redirect: They'll say "it's under investigation", which sounds neutral and ends the conversation. METR should not go quiet while it waits for the official action. Keep making the case in public that this is important work and explain why (without responding to the attack).
Intimidate → Resist: This may have a chilling effect on every other evaluator, whether anyone intends it or not. Other evaluators and researchers can say publicly that they'll keep publishing their work, speaking out, and will stand with METR.
I've worked on a few similar cases in the past two years, mostly in the pro-democracy space, including the FBI's arrest of Milwaukee County Judge Hannah Dugan. We built a coalition beyond the usual suspects and brought her case to the public with polling that showed a clear message to the general public. We hosted a vigil before her trial and a "Court of Public Opinion" outside the courthouse each time she appeared before a judge.
Making this backfire will be easier to do now, before the attack lands. Crisis comms is 95% preparation and relationships, 5% the crisis itself.
Would love to talk with more people about this!
*AI use: I used AI to tighten up the backfire bullets and intend to lengthen this piece into a longer one.
Some quick observations on the attacks on METR, EA, and other AI safety orgs (from my perspective as a political campaigner and former Communications Director for a union and mayor during COVID):
I want to write something longer on this but don't have time right now. Lots of opportunity for good crisis comms here. Let me know what's wrong about this take and if you'd like to hear more!
3 & 5 make sense.
1 / 4 / 6 feel like they might make sense strategically, but I'm proud to be part of a community that often values epistemic humility and correct argumentation above strategic communication.
Here are some quick thoughts:
Points #1, 4 and 6 aren't asking anyone to really say anything differently, it's more about who says what, where, and when. If someone from outside the AI safety world (or better yet, someone who can show they've disagreed with METR before or are not a natural ally) says "the corruption story isn't true," that's not less epistemically correct, it's just likely more believable to the general public.
I'd gently push back on the bigger part of your argument though: The public conversation can't end up being about METR's independence, it should be about an outside entity checking the lab's work. That seems to me the tradeoff if we lose the bigger comms battle.
Says "if you're explaining you're losing". IMO explanations are good. You might be correct that needing to explain is a bad sign for whether you are winning a comms battle. But I would still rather offer people genuine explanations, so we can come to a shared understanding of the truth (exactly like this convo).
My reading of this is that you're claiming who says something is more important than what is said. Again, you may be correct here on how to optimally persuade. But I personally want to value good arguments, no matter who makes them.
Feels like it is straightforwardly asking folks to not engage in debate or disagreements. I don't really get how we come to a shared understanding of the truth if we don't offer clear reasons for why we disagree with particular points.
Again, I very much am not saying you are wrong about the implications of ignoring strategic communication norms. You are likely right. I would nonetheless prefer to be in a community that decided to communicate honestly and earnestly, even with people that it disagrees with.
Is that actually true of X's reply mechanism to a significant degree? My impression was that replies on X are mostly just seen by people who tap on the original post. Replies wouldn't necessarily have the effect of amplifying the original post?
Certainly writing Community Notes would not present much risk I presume??
Thanks for pointing that out.
I meant "if you're explaining, you're losing" for any platform, not X replies specifically.
I agree Community Notes are very likely net positive (they certainly have an impact on me). But where I still see some risk is volume: Replies are a signal and if a lot of people pile on to rebut a post, perhaps that pushes it to more people? (I don't know X's algorithm well enough to say that with much certainty.)
I also don't want to double down too hard on illusory truth. I defined it a bit too loosely in my quick take, and there are follow-up studies showing a good correction usually wins with the people who read it.
For me it's more about Zaller's Receive-Accept-Sample (RAS), which I find helpful for thinking about mass/political comms. Voters have to:
Continuing to respond to a critique that's wrong makes it more likely people 1) hear it and 3) sample it later.
Thanks for continuing to engage.
Thinking aloud here. I don't think you necessarily have to respond by quoting the person. You can just provide a frame or facts which counter or inoculate against whatever critiques are currently widespread. E.g. if someone accuses you of a crime, you can mention that you were someplace else at the time the crime was committed, without directly repeating the criminal accusation. However, your X replies might fill with accusation talk in that case.
Wanted to share a quick update on my comments on AI safety comms from a couple weeks ago. The NY Post is now reporting that the FTC is investigating OpenAI, Anthropic and METR, with formal demands (similar to subpoenas) reportedly coming in the next few weeks.
I think Brian Martin's Backfire framework could be useful in responding here. Martin found that power holders use five methods to suppress outrage, and there's something we can do in response to make it backfire:
I've worked on a few similar cases in the past two years, mostly in the pro-democracy space, including the FBI's arrest of Milwaukee County Judge Hannah Dugan. We built a coalition beyond the usual suspects and brought her case to the public with polling that showed a clear message to the general public. We hosted a vigil before her trial and a "Court of Public Opinion" outside the courthouse each time she appeared before a judge.
Making this backfire will be easier to do now, before the attack lands. Crisis comms is 95% preparation and relationships, 5% the crisis itself.
Would love to talk with more people about this!
*AI use: I used AI to tighten up the backfire bullets and intend to lengthen this piece into a longer one.
Pangram says this quick take is 51% AI: https://www.pangram.com/history/01232a86-af15-43a2-8e6a-047896c47336?ucc=2RxEGYyJeZr
Is it?
Thanks Yarrow, I used AI to tighten the backfire bullet points and just added a disclosure!