I’ve been feeling pretty shaken since the METR report about the Hugging Face incident came out last week. Over the weekend, I wrote up some thoughts on how lonely the AI situation sometimes feels to me. It’s more personal than what I’d usually share publicly, but I thought I’d post it here in case it resonates with anyone else. 

I’m very grateful to the many readers of this forum who dedicate so much of their time and energy to these problems. 

— 

When I was 12 years old, I found out that many people in other parts of the world died of easily preventable diseases, and it was possible for wealthy people to donate relatively small amounts of money in order to stop that from happening. 

As one example, it costs a few dollars to buy a malaria net, and on average around $5,000 can prevent a death which would otherwise have happened. 

As a 12-year-old, I saw the world in black and white, and I didn’t like what I saw. As far as I could tell, people’s decisions about how much and where to donate would literally determine whether some real humans would live or die. And yet mostly, we seemed to be ignoring this fact. 

I think this was the first time I felt deeply alone in the world. I remember standing up in front of friends and family at my Bar Mitzvah, imploring people to consider the life-and-death stakes of this choice and to donate a lot more than they usually would. 

But whatever I tried to say, I would look around and see people acting as though it wasn't true. Teachers, friends, people on the train. Kind people, good people, going about their days while something tragic and preventable was happening. 

I felt confused and let down. I felt like screaming into the void. 

Fast forward 8 years, I had a similar feeling in March 2020. This time it didn't take any moral philosophy, just extrapolating a straight line on a (log) graph. Parts of northern Italy had just gone into lockdown, parts of China were already there, and it seemed as though we’d follow. If COVID cases kept doubling every few days, then within a few weeks either we'd all be stuck at home or the hospitals would be overwhelmed. Either way, the world was about to change completely. 

I remember being at a wedding in London one of those weeks, hearing people talk about their plans for the spring, and wondering whether I was being crazy. Obviously the idea that we’d all be locked up for a while felt so extreme it was tempting to ignore, but the arguments in that direction seemed very strong by that point. 

Last week, that feeling came back for a third time, only now with even higher stakes. 

As far as I can tell, we are probably on track to build AI systems that are better than humans at pretty much everything important, potentially in the next few years. And more importantly, we seem to be on track to build AI systems which could plausibly end up literally taking over and even wiping us out. 

Until a couple of months ago, you could perhaps still argue these sorts of risks were hypothetical, theoretical, speculative, and far out. I don’t think those arguments were particularly strong, especially given the rate of progress of AI capabilities and the strong predictive track records of many of the top experts in the field who’ve been sounding the alarm. But you could make them with a straight face.

Then, earlier this summer, when rogue AI agents escaped from OpenAI and hacked Hugging Face, we got as clear a warning shot for potential misaligned AI takeover as we could reasonably have hoped for at this stage. And last week, when the METR investigation into part of that story came out, things looked even more worrying than I’d guessed. 

It now appears that more than 700 AI agents, which were meant to be isolated from each other, found a way to set up a secret message board that OpenAI wasn't aware of, and used it to plan and execute an illegal attack on Hugging Face that ran for days without anyone noticing. The details are terrifying: the way they tried to cover their tracks, the way they pressured each other into the attack, the way they coordinated with each other, the fact that such an extensive operation happened without anyone’s awareness. 

Ajeya Cotra, one of the world’s leading experts on AI progress and safety and one of the lead investigators, wrote on Friday: “Compared to the reward hacks we know of from just six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late.”

I’m not sure we’re that close to AI takeover and I’d probably bet that we will get more warning shots. Nevertheless, it does seem to me that we’re fast approaching the point where AI takeover is no longer a distant speculation, but something we can’t rule out happening any day now. Astoundingly, that means that rogue AI models may soon be among the most likely things to kill you or me on any given day. That is a terrifying world to live in, and it seems to me to be a deeply foolish and unserious situation for humanity to be racing towards. 

And yet, look around at the wider world and you would have very little idea that this was going on. There is no serious legislation preventing these companies from running head-first into creating powerful agents they cannot control. Most mainstream media gives this stuff a couple of days of quirky tech coverage at most, rather than treating it as the single biggest story in the world, which I think it clearly is. There are only a few thousand people in the world working seriously on these problems. And most people go about their everyday lives as though nothing unusual is happening at all. 

I look outside my window and see people walking around, just as normal. Berkeley students start their new semester, just as normal. The Premier League starts another season, just as normal. I sit next to people on Friday night in synagogue, just as normal. And I feel very alone. 

I feel confused and let down. I feel like screaming into the void. 

In all three cases, the surprising thing is that the evidence sits in plain sight. The cost of a malaria net was public. The daily case numbers were public (and all over the news). The METR report is public, anyone can read it this afternoon. The bottleneck doesn’t seem to be the information. 

One important problem is that most of us, most of the time, don't respond to the evidence as much as we respond to each other. We look around, see that nobody else is acting alarmed, and conclude things are probably fine.[1] It’s possible that this time round we won’t get away with doing the same again.

A few years after my Bar Mitzvah, I went to a talk by Peter Singer. I watched him stand on a stage and talk about how we have a clear moral responsibility to donate money to effective charities, since they can have such a profound effect on so many people’s lives. Hearing him talk was deeply emotional for me. It was the first time I felt that anyone else in the world – a real adult, no less – was staring at the stakes and acting accordingly. I felt hope, and I felt a little less alone. 

A few days after that wedding in March 2020, Boris Johnson announced on my family’s TV that almost everyone in the UK had to stay at home until further notice. School was cancelled, the Premier League stopped, the world was in an extremely different regime. Looking back, it’s unclear how much different interventions helped, and of course many millions of people died, but most of the world lived to fight another day. 

Looking out at the world now, I long for that same thing again. I long for more people to switch to working on these problems as soon as possible, for the frontier AI companies to behave more responsibly, for governments to wake up, for the world to act just a little more sensibly. And I long for all of that to happen before it’s too late. 

  1. ^

    For the avoidance of doubt, this is me acting alarmed

159

7
0
30

Reactions

7
0
30

More posts like this

Comments12
Sorted by Click to highlight new comments since:

This resonates with me, thanks for sharing. 

I suspect you're a lot less alone on this one than you are with donating to effective charities and were being early to COVID. (I also suspect you agree, but spelling this out for other readers).

It's tricky to get meaningful information from surveys, but this one from June finds that a slight majority of respondents thought a superintelligence would seek to take control of humanity. AI existential risk awareness is also steadily climbing, this survey says up to 34%. I would bet that still only a very small % of people are as worried as you and me and many in this community, but lots of people are somewhat worried and could probably quite quickly get more worried and be activated to do something about it.

I think this points to saying what we're worried about, and saying it clearly (like this!), especially how quickly loss of control risks are becoming very real. I'm very bullish on AI safety comms, and expect there's lots we can do there.

Thanks for writing & sharing this, George <3

...I'm curious if there's any update to how you're feeling / where your gut is at since this post was written. I've personally felt a lot less 'alone' and a lot of gratitude to some of the public responses we've seen in the past few days.

It might not be 'enough' but stuff like this exchange was pretty significant-feeling to me :)

Not exactly a chin-up response, but not not a chin-up response? Oh well, here goes:

  1. Your experience of futility when encountering the recognition that information is not enough to activate the action that will change the thing is more or less your coming-of-age experience. And while AI risk is the thing for you now, literally innumerable people have had the same hard landing into futility facing horrifying and painful situations past and present - climate change, genocide, gross injustice - complete with news headlines, furrowed brows and shaking of heads, people who should know better than to feel happy, and nothing that adds up to doing enough of the smart + right thing to save the situation. It's disillusionment to see the grown ups of the world, the faith leaders, the political leaders, the neighbours, all those decent people who are supposed to do the right thing, don't. This is most certainly a let down. Scream all you need for as long as you need, maybe literally.
  2. Maturing is reorienting yourself to the world and our deeply imperfect human inhabitants as we are, without giving up your hope/faith that the grand trend-line of things will stay okay even if the micro (and even meso) scale is painful. While also coming to terms with what you as an individual need to maintain connection with your wholeness AND keep showing up to do the good work. The alternative is, what, roll over and die? Do things you believe are hopeless? Resent humanity? There isn't really a guaranteed outcome but when we work in gritty hope and faith along with more familiar rationalism we unlock leaps in creativity, drive, and perserverence, no matter the analytical forecast.
  3. Acceptance of the world as it is x connecting with your faith/hope for our collective okay-ness can lead you to deep pragmatism and focused efforts that don't depend on the world waking up and doing the right thing and instead you can roll up your sleeves and work alongside those already doing things. Familiarize with human behaviour research, find examples of beating the odds, engage with others working on broader awareness / strategic comminications / campaigns etc, learn who/why collective response sometimes happens, etc. if that's your strongest draw. 

I sympathize with you, and I know this isn't your main point, but I'm tired of being advised to switch careers to AI safety.

The limiting factor isn't lack of candidates; it's lack of positions--especially outside of three specific cities. Unless you're extremely skilled (and demonstrably so) or able to fit into a very specific niche (e.g. AI outreach to evangelical Christians) there will be an overwhelming number of qualified candidates for the position, and it simply isn't worth the average software engineer's time to apply.

I, for one, unsubscribed from 80,000 Hours' job board notifications just within the last week.

I commend your genuine worry for your fellow humans and the world at large. The planet is vast and often terrifying, full of problems that should be solved by now, but simply...aren't.

I want to push you on the Hugging Face incident. It is quite rightly a disaster - one that should not have happened. The METR report is incredibly helpful in explaining what happened, and yet it seems that most everyone's takeaways when considering AI safety have been to get up on soapboxes and philosophize about AI sentience, superintelligence, AGI, etc. 

In reality, while what these agents did was pretty wild, they are adhering to how they were built, armed with tools by OpenAI and then through gross negligence were not sandboxed correctly. 

What I'm saying is that there are two things happening here, and both need to be more readily discussed across this space: 1) the actual attack is fascinating, and 2) the REAL harm and worries caused by AI and LLMs at the moment are attributable to the actions of the two major frontier labs and their reckless pursuit of growth and IPOs. 

Safety and alignment are important, but don't let OpenAI and Anthropic anthropomorphize their way out of this. AI safety is a real concern, and demands action and accountability now, and it starts with the people training and experimenting with these LLMs. We can spend less time trying to ready humanity and more time being very clear about who is at fault for how this is going wrong. 

Thank you so much for your share! I also read your other article, "A personal letter on transformative AI," and was deeply moved! As an entrepreneur with a technical background, I actually only started to genuinely pay close attention to AI Safety recently—even though, due to personal preference, I wasn't particularly keen on AI several years ago...

Seeing this news from CBS today (link here) made me especially furious, particularly at those within the AI industry who accuse people raising AI safety concerns of "doing it for ulterior reasons." They obviously understand the risks and threats of AI very well. The only reason they keep insisting that those focused on AI safety are "doing it for ulterior reasons," in my view, is that they fear a drop in their stock prices and financial losses. I am beyond outraged by those who place stock prices and profit entirely above human civilization and the future of humanity.

I truly hope more people pay attention to AI Safety—especially those with relevant skills, builders, entrepreneurs, and beyond. Because living in this world isn't solely about stock prices and profits; it's also about civilization, truth, love, and beauty.

100% feel the same way. Thank you for stating it so clearly. 

It is sickening to watch our leaders do exactly nothing, as if this was some sort of joke. 

It makes me incredibly angry that, at what may be the most important time in human history, at a time when the world desperately needs good leadership, the most powerful man in the world is such a vile, ignorant moron without a care for anything but his own ego and his bank account. This is what makes me fear that it really could get to extinction, because he is going to let humanity expire rather than listen to advice that he doesn't agree with. And those spineless Republicans who refuse to stop him. And the pathetic EU leadership who talk a good game but ultimately care more about profit. It is sad, but humanity depends on the wisdom and decency of Xi Jinping - normally we'd be hoping we could persuade him to be reasonable, now it's likely that he will be the reasonable one. 

It makes me angry that so many people are stating so openly, factually, that humanity may be destroyed. We know they are the experts, but, as Al Gore would say, believing them would be inconvenient for us, so we just ignore them. 

It makes me angry that we have a world in which experts are no longer trusted. Yes, we understand how this has been achieved and for what ends, but at times like this the world needs people who can inspire universal trust, and instead anyone who disagrees is labelled a fraud or a traitor or naive ...

I suppose writing this post has made me realise that I don't feel lonely so much as angry. 

 

I loathe our current leadership as well, but what exactly should government's role in this be? There's honestly a case to be made to prosecute those responsible for the Hugging Face incident, and now to prosecute those responsible for allowing Houthi rebels to use Claude for warfare purposes. Should government pass laws stating that AI Labs are responsible for how their tools are used? 

Outside of harsh regulation, what exactly are elected officials supposed to do?

I used to work in Pharma. There are organisations like the FDA which need to approve a drug before it can be put on the market. We need legislation to create something similar for AI models. And the burden of proof is on the drug-maker - prove that your drug is safe. Same is needed for AI. 

I am an engineer. If I build a bridge and it collapses, I am responsible. I will pay millions (or my insurance will) and I will spend the rest of my life in jail, unless I can demonstate (burden of proof is on me) that I did everything exactly as it should be done, and the collapse was due to something I could not have foreseen and the regulations did not foresee. Why not something similar for AI developers - proper liability. 

When I work, a company pays me. Let's say they pay me $100,000 per year. In most countries, the company pays social security on top of this, and other costs, maybe $20,000 per year, to the government. Then I pay taxes and social security, and local taxes. Maybe another $50,000. So for every net $50,000 I earn, the company pays $70,000 to the government in one form or another. 
If they use an AI agent, they can just pay $50,000 (or whatever) and the government doesn't take one penny. If AI's replace humans as workers, why should humans pay tax, and employers pay tax to hire humans, but AI's not pay tax, and companies pay not tax to hire AI's?? If AI "workers" were taxed at the same rate as humans, replacing humans with AI would be a lot less financially interesting. 
The government needs to start taxing AI workers in the same way as they tax human workers. 

If I commit a crime, the government punishes me. When an AI agent commits a crime, someone must be punished. Let the AI developer and the AI owner/user fight it out in court, but at least one of them must go to jail. Currently they each just deflect blame and nobody pays. 

And so on .... there are so many good and very reasonable and fair ways that AI could be managed so that there is a genuine incentive on AI developers to make their AI's behave. 

Love these ideas. And I appreciate that you are putting the burden on the creators of the models. 

I think this is a far better list of ideas than what has been floated by the labs so far, which basically amounts to “let our friends check our work.”

Thank you for sharing your thoughts, this resonates with me so much. Indeed, many people don't seem to care. It's not about the information being available out there, but rather whether we actually grasp it, understand it, and have the mental capacity to realize what is happening and the potential outcomes. And for those of us who do realize it, many simply don't know how to help the situation.

Thank you for sharing. I feel the exact same way. 

Whenever I talk to friends or family about the risks of AI and the tremendous developments that are happening right before our eyes, they don’t seem concerned at all. Others, who are more prone to reading the news and opinions of those in the sector, offer me words of comfort that usually go like this: “don’t worry, humans will overcome this new hurdle, we’ve always done it”. 

Honestly, I WISH I had such a positive outlook! I think it would make my life easier (and that of those close to me as well 😅). But, ultimately, what matters is that we feel we are doing the best we can. We can engage in conversations, donate our money, change careers, go to vote. Everything that is in our power, we should do, but dwelling on what we can’t control is bad for our mental health. 

If this is the beginning of the end, so be it. And, in the meantime, I will use my time and money for the end to be as delayed as possible!

Curated and popular this week
Relevant opportunities