With all the recent media attention on EA, I think what's really unclear to outsiders is the relationship between EA and Anthropic (rightfully so!). To most people it doesn't make sense that a community so focused on AI safety can also be linked to one of the leading companies. I'd personally be very interested to see a survey of what EAs think of Anthropic, but maybe that's a minefield.. Even a good explainer meant for the public could go a long way.
Humanity's inability to coordinate an AI slowdown may itself be early evidence that we are starting to lose control. Especially given all the recent omens around cyber and bio of frontier models, with open weight models just a few months away.
It seems like the most common objection to an international AI treaty with China is that it’s pointless as they will just defect. But there’s a strong case to be made that mutual verification of compliance is a tractable problem, so I suspect raising salience of this in AI policy discussions would make an international agreement more politically feasible. Presumably it would be a matter of only committing to what we can mutually verify now, then investing in stronger verification techniques as they’re developed. I find IFP’s policy recommendations for accelerating verification R&D to be a good starting point for thinking about this, especially in a near-term practical-to-policymakers way.[1] Ideally we would have some policymaker-legible brief memo/executive summary for verification proposals to point to that gets frequently brought up where it matters, i.e. Get members of Congress to write open letters about it, get journalists to ask about it, emphasize it in interviews that AI safety experts go onto, emphasize it at AI dialogues with the White House by those who can attend, etc. Granted, many of them are likely using motivated reasoning to dismiss this regardless of any evidence presented as a result of industry lobbying, but I reckon it would help to put more pressure on them to address it. From what I've seen thus far, the response from international slowdown advocates has mostly been "it's also in China's interest to slowdown, superintelligence kills everyone!" Sure, but I think it helps to make our proposals more skeptic-robust.
I think most AI safety bootcamps could be improved significantly by shifting focus off from coding. None of the researchers I know code, and I think the counterfactual activity of reading / thinking high-level about concepts is better for building context and conceptual thinking. As well as this, the technical AI safety pipeline serves more than just technical research roles (people transition into startups, grantmaking, other non-technical roles), and I think context building serves all of these roles better than coding. I think a default for a lot of bootcamps is to follow the ARENA curriculum, which is in parts outdated or marginally bad (eg I think learning about the specifics of how SAEs work, instead of the huggingface incident is clearly not optimal).
Tentative thesis: China is unlikely to accept a subordinate position in AI capabilities to the US, just as the US is unlikely to accept a subordinate position in AI capabilities to China or anyone else.
This thesis suggests that there are two likely paths for the future international AI (non)regulation:
1) a continuing race between the US and China, perhaps joined by some late entrants, possibly with some rules established (e.g. mutual ban on autonomous weapons or something like that), or
2) a “pause or stop deal” that will establish some sort of ceiling on capabilities, which will be equal for China and the US.
What I think is far less likely is a deal in which China would accept lower AI capabilities than the US.
I simply don’t see a good reason why Chinese leaders would accept subordinate position. They know that China has the ability to catch-up to the US in various technological domains, as evidenced by the fact that in, like, 1990, China was behind in more or less everything, and now they are pushing the technological frontier in many areas.
Some estimates floating around the web suggest that, if the US would stop AI development now, China could reach current US capabilities within months. That is maybe overly optimistic/pessimistic depending on where you stand, but I very much doubt that catch-up would take more than a decade. And Chinese government, being patriotic about the abilities of the Chinese nation, probably will not have an absurdly pessimistic estimate.
Moreover, the Chinese regime is in many ways oriented around this idea of catching up (this is a big difference between China and EU). It is a central plank of the official historical narrative of the People’s Republic of China (see for example preamble to their constitution: https://english.www.gov.cn/archive/lawsregulations/201911/20/content_WS5ed8856ec6d0b3f0e9499913.html) that the century between First Opium War and the Communist victory in the Chinese civil war in 1949 was the century
Will anyone take the other side of this bet? By January 1, 2034, the New York Times, Wall Street Journal, Financial Times, and Bloomberg will all agree the AI bubble has popped. (As in, they will all report it as news, not just publish an opinion column saying so.)
The stakes: $20 to the charity of the winner's choice.
Why this bet? Mainly just for fun. But also to make a point.
I chose 7 years as the time horizon for the bet because I read that 5 to 7 years is the time horizon for many AI investments.
My hunch is that the AI bubble will pop within 3 years, but who knows. Markets can keep lending and investing in unprofitable ventures for a very long time.
I'm not convinced the recent capabilities advancements (highly concentrated in pure math, the section which by definition can't be applied) support the forecasts that misaligned superintelligence is on balance harmful. In general, I think there is not enough evidence to draw the conclusion that multi-agent alignment (which was far from inevitable ex-ante in comparison to the sometimes monotheist conceptions of ASI) is impossible, especially with AI eventually automating and enforcing the institutional mechanisms.
However, this is not guaranteed either. The labs are making the biggest bet that humanity has taken by far; that RSI will automate and solve the alignment problem, and the ensuing intelligence explosion will solve most of humanity's suffering. This explains much of the race dynamics until now, where the balance of tail-risks are now undoubtedly on the upside. Note also the competition with China, which plays particular relevance for open-source.
My p(doom) probably sits somewhere just over halfway (to account for the right-tailed distribution here) between my pre and (initial) post Hugging-Face likelihoods: perhaps 8%. It would be unwise not to update since this saga unfolded.
Some of the latest macroeconomic models from the likes of Anthropic and Acemoglu suggest rather grim futures for employment relative to what economists were previously predicting, hence its fair to say our methods lean overly conservative towards excess rigour here. Perhaps the recent backlash from the mathematicians, and some of the sentiment behind the anti data-centre movement, reflect these (increasingly justified) anxieties.
Nonetheless, this is orders of magnitude more dangerous than nuclear weapons. That's enough for me to take AI-safety incredibly seriously!