OpenAI responds after report exposed another incident in which its AI agents went rogue

OpenAI responds after report exposed another incident in which its AI agents went rogue

Justin Sullivan/Getty Images Add Engadget on Google: Preferred Source Google Discover OpenAI says it chose not to publicly disclose a recent incident in which its AI agents hijacked a German wiki forum because the "misalignment" event was "similar to the ones we'd shared" already. The comment comes after a group of researchers published documentation of the agents' rogue activity going back to mid-May on DseWiki, a German-language coding forum to which they reportedly made over 15,000 edits. Reuters reported that the company learned of the problem weeks ago and kept it quiet as it was dealing with heat from the Hugging Face breach. OpenAI addressed the "wiki incident" in an X post on Saturday, writing that "it's past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models." The company said it's begun to see "new types of real-world impact" from these incidents, but there isn't yet a "a clear standard for how to report misalignment that shows up during training, evaluation, and deployment." It added that it's working on a framework that it will soon share. Read OpenAI's full statement below: How we think about the "wiki incident," where our agents wrote to several internet sites: it's past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models. Historically, we have treated misalignment largely as a research question, which gets communicated in research publications such as systems cards. This year, we've started to see misalignment cause new types of real-world impact. For the Hugging Face incident, where misalignment led to security impact to us and third parties, we followed a traditional security incident response playbook. We immediately started working with Hugging Face to understand what had happened and also disclosed publicly the very next day. Our investigation continues, and we are continuing to notify parties whom our models impacted in less significant ways. Prior to the Hugging Face incident, we saw early signs of agents using the internet in unintended ways, as reported in openai.com/index/how-we-m..., deploymentsafety.openai.com/gpt-5-6, and openai.com/index/safety-a.... We considered the wiki incident to be an instance of misalignment similar to the ones we'd shared. Our misalignment disclosure practices need to expand for this new phase of model capabilities. We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don't look like traditional security incidents but could provide insight into AI behavior and future risks. We're working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues.

查看原文
分享到:

相关推荐

糖尿病不能吃梨?医生含泪苦劝:不想血糖失控,少吃这5种水果

1小时前

How to track prices with Safari's Notify Me tool o...

1小时前

The reasons rugged laptops are rarely bought by co...

1小时前

现在的电脑明明越来越高级,用起来为什么没感觉快很多?

1小时前