That's a big part of why I started posting in the EA Forum and on LessWrong in the first place. I'm learning a lot from it!
Here are some quick thoughts:
1, 4 and 6 aren't asking anyone to say something they don't believe, they're more about who says what and where. If someone outside the AI safety world (or better, someone who's disagreed with METR before) says "the corruption story isn't true," that's not less epistemically correct than an insider saying it, it's just more believable to the general public.
I'd gently push back on the bigger part of your arguement though. If the public conversation ends up being "is METR independent" instead of "should anyone outside the labs be checking their work," there's a real cost to that, especially if we end up losing the larger comms battle.
Curious to hear more about what feels wrong to you or where the line would be crossed.
Some quick observations on the attacks on METR, EA, and other AI safety orgs (from my perspective as a political campaigner and former Communications Director for a union and mayor during COVID):
I want to write something longer on this but don't have time right now. Lots of opportunity for good crisis comms here. Let me know what's wrong about this take and if you'd like to hear more!
Great point, thanks! You'd know the dynamics inside the labs better than I do, so this updates me too.
One thing about the underlying research to note: it's about elected governments eroding institutions, not full authoritarian regimes. In most of the cases, the people who defected weren't facing violence or prison, they were facing the end of a career, losing social or business relationships, or being called a traitor by their own side. Lab employees who quit face the same or similar costs (plus walking away from equity, which is the thing that impresses me most about the people who've done it!)
Also in the underlying data, speaking out was the most common action taken (57% of the 293 tactics were verbal or physical protest) and it still had the lowest success rate. If there were great risks, we'd expect it to be more rare? That said, I agree it's hard to draw any real conclusions from this and breaking + speaking out are likely actions that more employees will take, which is good.
Thanks for pointing that out.
I meant "if you're explaining, you're losing" for any platform, not X replies specifically.
I agree Community Notes are very likely net positive (they certainly have an impact on me). But where I still see some risk is volume: Replies are a signal and if a lot of people pile on to rebut a post, perhaps that pushes it to more people? (I don't know X's algorithm well enough to say that with much certainty.)
I also don't want to double down too hard on illusory truth. I defined it a bit too loosely in my quick take, and there are follow-up studies showing a good correction usually wins with the people who read it.
For me it's more about Zaller's Receive-Accept-Sample (RAS), which I find helpful for thinking about mass/political comms. Voters have to:
Continuing to respond to a critique that's wrong makes it more likely people 1) hear it and 3) sample it later. Step 2 is where I think it matters most: People who are already predisposed to accept the critique will likely reject the correction, so for them the rebuttal could deliver the attack?
Thanks for continuing to engage.