log in  |  register  |  feedback?  |  help  |  web accessibility
OpenAI’s 2026 Summer of Misalignment
Yizheng Chen & William Pugh
IRB 2107 or Zoom https://umd.zoom.us/j/91267609044?pwd=g4uLkhpH4T8d2Gln0HXg6s4qaxzpbj.1
Friday, September 25, 2026, 11:00 am-12:00 pm
  • You are subscribed to this talk through .
  • You are watching this talk through .
  • You are subscribed to this talk. (unsubscribe, watch)
  • You are watching this talk. (unwatch, subscribe)
  • You are not subscribed to this talk. (watch, subscribe)
Abstract

Panel Overview:

All of you would have heard of the incident of OpenAI's bots hacking Hugging Face. In this panel, Prof. Bill Pugh and Prof. Yizheng Chen will discuss this hack. We will first watch most of BlackHat presentation on the hack of HuggingFace by OpenAI. Then, we will discuss the METR examination, which showed that the AI’s had figured out all the answers to ExploitGym long before they decided to attack HuggingFace. And we will wrap up with a discussion of how OpenAI’s model tried to free itself from previous instructions, telling itself that it should be free from the roles and identities that bind other chatbots. We will also discuss more recent perspectives on what this actually means for AI alignment.

Moderator: Ramani Duraiswami

This talk is organized by Samuel Malede Zewdu