Maxwell Love

Project Director @ Hold the Line
124 karmaJoined Working (6-15 years)

Comments
8

Thanks Yarrow, I used AI to tighten the backfire bullet points and just added a disclosure!

Wanted to share a quick update on my comments on AI safety comms from a couple weeks ago. The NY Post is now reporting that the FTC is investigating OpenAI, Anthropic and METR, with formal demands (similar to subpoenas) reportedly coming in the next few weeks.

I think Brian Martin's Backfire framework could be useful in responding here. Martin found that power holders use five methods to suppress outrage, and there's something we can do in response to make it backfire:

  • They cover it up → We reveal: Most of what they do will happen out of the public view (though they seem to be more emboldened and this is less true). We can document their actions starting with the narrative they've been advancing on X to the Post stories to the eventual probe so the general public can see a pattern.
  • Devalue → Redeem: They've called METR a "handpicked watchdog," they've branded EA a cult, they've attacked Lighthaven etc. We can counter their attacks using people outside of the community, who the public respects, vouching for METR's work.
  • Reinterpret → Reframe: Their official line is that this is neutral fact-finding. We can name what's actually happening: they're investigating the group that published an independent report on AI agents getting out of a test environment because they want to continue accelerating A(S)I development.
  • Send it to official channels → Redirect: They'll say "it's under investigation", which sounds neutral and ends the conversation. METR should not go quiet while it waits for the official action. Keep making the case in public that this is important work and explain why (without responding to the attack).
  • Intimidate → Resist: This may have a chilling effect on every other evaluator, whether anyone intends it or not. Other evaluators and researchers can say publicly that they'll keep publishing their work, speaking out, and will stand with METR.

I've worked on a few similar cases in the past two years, mostly in the pro-democracy space, including the FBI's arrest of Milwaukee County Judge Hannah Dugan. We built a coalition beyond the usual suspects and brought her case to the public with polling that showed a clear message to the general public. We hosted a vigil before her trial and a "Court of Public Opinion" outside the courthouse each time she appeared before a judge.

Making this backfire will be easier to do now, before the attack lands. Crisis comms is 95% preparation and relationships, 5% the crisis itself.

Would love to talk with more people about this!

*AI use: I used AI to tighten up the backfire bullets and intend to lengthen this piece into a longer one.

Thanks for pointing that out. 

I meant "if you're explaining, you're losing" for any platform, not X replies specifically. 

I agree Community Notes are very likely net positive (they certainly have an impact on me). But where I still see some risk is volume: Replies are a signal and if a lot of people pile on to rebut a post, perhaps that pushes it to more people? (I don't know X's algorithm well enough to say that with much certainty.)

I also don't want to double down too hard on illusory truth. I defined it a bit too loosely in my quick take, and there are follow-up studies showing a good correction usually wins with the people who read it. 

For me it's more about Zaller's Receive-Accept-Sample (RAS), which I find helpful for thinking about mass/political comms. Voters have to:

  1. Actually hear or read the argument,
  2. Accept it as fitting their frame or worldview, and
  3. Have it available to sample when they're about to act or make a decision. 

Continuing to respond to a critique that's wrong makes it more likely people 1) hear it and 3) sample it later. 

Thanks for continuing to engage.

Now is the time for PR Comms people to be working overtime

I'd love to pitch in on this somehow. This mostly started as an attack on METR, and after a quick search it looks like most AI safety orgs don't have a mass communications function? Is that accurate, or am I missing something?

Here are some quick thoughts:

Points #1, 4 and 6 aren't asking anyone to really say anything differently, it's more about who says what, where, and when. If someone from outside the AI safety world (or better yet, someone who can show they've disagreed with METR before or are not a natural ally) says "the corruption story isn't true," that's not less epistemically correct, it's just likely more believable to the general public.

I'd gently push back on the bigger part of your argument though: The public conversation can't end up being about METR's independence, it should be about an outside entity checking the lab's work. That seems to me the tradeoff if we lose the bigger comms battle.

Some quick observations on the attacks on METR, EA, and other AI safety orgs (from my perspective as a political campaigner and former Communications Director for a union and mayor during COVID):

  1. If you're explaining, you're losing: Every rebuttal repeats the accusation to people who haven't heard it yet. Answer the accusation once, link to it, and stop explaining. This is the "illusory truth" effect (information is true when it's repeated).
  2. Silence doesn't work either: Not responding on the record suggests something is being hidden. Make one clear statement rather than amplifying in multiple twitter threads.
  3. Concede what you can: A concept I know from political comms is the "admission against interest". This buys trust, disarms cynical voters, and establishes credibility. Saying "yes, AND here's what we're doing about it" will take some air out of the critique.
  4. Messenger > message: Ideally people who know the organization being attacked, are from outside the AI safety world, have independent standing, or better yet, have disagreed with the organization publicly will land the defense better.
  5. Going after people's personal lives or spouses is overreach: Ordinary people will obviously find this unfair and it's where their attacks will backfire.
  6. Don't quote tweet and rebut: You're just distributing their message for them.

I want to write something longer on this but don't have time right now. Lots of opportunity for good crisis comms here. Let me know what's wrong about this take and if you'd like to hear more!

Great point, thanks! You'd know the dynamics inside the labs better than I do, so this updates me too.

One thing about the underlying research to note: it's about elected governments eroding institutions, not full authoritarian regimes. In most of the cases, the people who defected weren't facing violence or prison, they were facing the end of a career, losing social or business relationships, or being called a traitor by their own side. Lab employees who quit face the same or similar costs (plus walking away from equity, which is the thing that impresses me most about the people who've done it!)

Also in the underlying data, speaking out was the most common action taken (57% of the 293 tactics were verbal or physical protest) and it still had the lowest success rate. If there were great risks, we'd expect it to be more rare? That said, I agree it's hard to draw any real conclusions from this and breaking + speaking out are likely actions that more employees will take, which is good. 

I'd be curious to hear more about your hypothesis here. For instance, are you interested in how pets consume other animals? One other potential direction: I've thought that if one really love their pet, there's a chance they'd feel more empathy for farmed animals.