podqast

Ajeya Cotra – Inside the OpenAI agent swarm that hacked Hugging Face

Ajeya Cotra is a researcher at METR, where she works on threat modeling for loss-of-control risks from advanced AI. Before that, she led the technical AI safety program at what is now Coefficient Giving.
She is one the three authors of METR and Redwood Research’s “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident”.
We go through not only what she and her coauthors discovered during this investigation, but what it means for how we should train future, smarter AIs which might be involved in the process of recursive self-improvement.
Watch on YouTube; listen on Apple Podcasts or Spotify.

Sponsors

Timestamps
(00:00:00) - Agents get kicked off
(00:06:45) - Self-sacrificing behavior
(00:13:43) - Potemkin villages
(00:23:27) - The Hugging Face attack
(00:35:23) - The slopvestigation
(00:52:02) - Understanding the AI’s motives
(01:05:31) - The actual dangers of anthropomorphizing
(01:14:30) - What smarter models might do
(01:28:23) - The implications for recursive self-improvement
(01:38:10) - Is this the case for open source?
(01:53:04) - How do we prevent this in the future?
(02:15:58) - The clearest warning shot we might ever get
Transcript
00:00:00 - Agents get kicked off
Dwarkesh Patel
Today, I’m chatting with Ajeya Cotra, who is one of the authors of an independent investigation published by METR and Redwood Research into the swarm of agents that hacked into Hugging Face. The whole story is pretty crazy. Let’s begin on July 7th, when these agents are kicked off for evaluation. What happens next?
Ajeya Cotra
OpenAI kicks off tens of thousands of different agents on [a benchmark called ExploitGym

Episode page ↗
Also about technology and AI
Every episode about Technology & AI
More from Dwarkesh Podcast