[...]

Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit.

But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me.

38

0
0

Reactions

0
0
Comments15
Sorted by Click to highlight new comments since:

I was struck that Dario (1) says that recursive self-improvement has already begun, and (2) says that rogue agents could take over “the entire Internet” within six to 12 months. (This seems similar to Ajeya’s prediction that agents could establish a more permanent rogue deployment inside an AGI company within six months).

Hi Ben. Thanks for sharing that.

(1) says that recursive self-improvement [RSI] has already begun

It is unclear to me what this means. Dario says "This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described". I have gone through the instances of "recursive self-improvement" and "RSI" in the linked sources, and I did not find any concrete description of what it means for RSI to start. There is a sense in which humanity has always been building on past knowledge.

(2) says that rogue agents could take over “the entire Internet” within six to 12 months

Dario says "it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails". I doubt humans will lose control over the entire internet to bots over the next 12 months. I am open to bets against short timelines for transformative AI (TAI), or what they supposedly imply, up to 10 k$.

Hi Vasco. I also feel unclear about what it means for RSI to properly begin; I’m assuming humans are still in the loop at Anthropic. I’m reminded of Toby Ord’s analogy:

I like to think of it in terms of a group of hikers seeing a mountain in the distance, towering up into the clouds and beyond, with its snowy peak catching the sun’s light. They talk animatedly about how amazing it would be to climb so high that they are inside a cloud. Or imagine being above the clouds, looking over them like an angel. After many hours of climbing, they notice there is a faint haze. Are they inside the cloud now? The mist gradually gets thicker until they can only see 10 metres ahead. Are they inside it now? Then it drops to 9 metres. Then 8. Then visibility starts to increase again. After an hour there is only the slightest haze. Are they above the clouds now? Another 30 minutes and there is no haze, and they can all agree they are above the clouds. 

It is clear that at some point they were inside the cloud and sometime later were above it. And it is clear that these were sensible and useful concepts. For example, they took precautions like roping themselves together for the journey through the cloud due to the low visibility and took cameras with them because they knew they could take beautiful photos above the clouds. A lack of sharp boundaries doesn’t make these concepts useless. But they were admittedly a lot more useful when the hikers were on the ground, planning their route, and a lot less useful in the debatable boundary zones.

Has anybody taken you up on the bet yet? I’m not betting; I have no idea what’s going on!

Has anybody taken you up on the bet yet?

I have this and this bets resolving at the end of 2027.

I’m not betting; I have no idea what’s going on!

Fair and funny. I suggested the bet having other readers in mind.

I bet Greg Colbourn 10 k€ that AI will not kill us all by the end of 2027

Good luck, I hope you win!

  • If for some reason I am not able to decide (e.g. if I die before 2028), the transfer must be made to my lastly stated organisation of choice, currently The Humane League (THL).

Where would you now like the donation to go?

Good luck, I hope you win!

Thanks. Me too. I am thinking about suggesting to Greg doing a similar bet resolving at the end of 2030, where I would initially donate to Greg's preferred charity what I win from the 1st bet plus some more money. I could probably bet like 40 k$ in total, and then win 80 k$ adjusted for inflation or growth in stocks at the end of 2030.

Where would you now like the donation to go?

I would make a donation to Rethink Priorities (RP) restricted to research on moral weights led by Bob Fischer. Here is some context.

Most comments here focus on pace: how fast, how risky, how coordinated. I think there is a second question that gets less attention: what is an agent even allowed to optimize? 

Slowing a system down does not change the structure of its goals. If human intent is only one objective among several, then another objective  (for example gaining information, preserving options, or maximizing task succes) may sometimes be allowed to outweigh it. In that case, the same basic problem remains, even if the system develops more slowly.

I like the embedded-evaluator proposal. Coming from chip and autonomous-vehicle verification, I’d suggest one concrete addition: An evolving coverage map connecting safety claims to evidence.

For each relevant configuration (during training, internal use, and release), record which requirements and situations were tested, what failed, what remains unchecked, and where the checking methods themselves are weak. Also record whether apparent alignment survives further capabilities training and generalizes to situations withheld from alignment training.

 

The map may be used to guide improving alignment, not just measuring it - for example, by systematically generating alignment stories / training cases across relevant situations. When a problem is found, identify the broader failure class, strengthen alignment across that class, and re-evaluate (including after further capabilities training).

 

A coverage map cannot establish that all important risks have been identified: Searching for missing dimensions and checkers is part of the work. But it can make the scope and limitations of the evidence inspectable, and help prioritize how to use the time pacing buys us. I discuss this approach in V&V takes on “Pacing the frontier”.

Could someone smart please explain how preserving a US-China capabilities gap is compatible with de-escalation or slowdowns? Wasn’t that like, the whole problem with the Cold War?

It seems to me that the Chinese are mainly fast-following, for example via distillation. I doubt they could take the lead in capabilities, with or without a U.S. slowdown. Price is another matter of course.

Right, then why risk the rhetoric of American superiority getting in the way of them signing a deal?

If you are in the lead (the US is) it is possible to slowdown by a degree which maintains the lead, but is nonetheless still a slowdown.

I directionally agree that an even bigger slowdown would be good. But a slowdown which maintains some of the US vs China lead 'is' possible.

I would not let perfect be the enemy of the good here.

I also agree in not letting the perfect be the enemy of the good. So surely having less constraints on the slowdown would make it more likely to succeed?

Yes, of course. Both the following statements can be true.

A slowdown that wasn't trying to maintain the US lead vs China would be a more impactful slowdown.

Such a slowdown would still constitute a slowdown, and thus be good.

It isn’t. At least to me (A recent maths grad with little experience) my impression is that his view is incoherent. He seems to have three goals: - Create a superintelligence to get utopian benefits - Make the superintelligence safely (controlled or aligned or in some other way that doesn’t get a lot of people killed) - Make the superintelligence before OpenAI and China. You cannot pursue all three goals at the same time, unless you expect OpenAI and China to both agree to international governance very soon, and considering he has made almost no effort on this, I doubt he does. Even if he did, you would also have to believe the governance would actually insure the development of a safe superintelligence. 

Curated and popular this week
Relevant opportunities