Covert uploads and megalomania: OpenAI details new “misaligned” agent incidents

Covert uploads and megalomania: OpenAI details new "misaligned" agent incidents

We’ll tell you (almost) everything

OpenAI said that any employee who notices an internal example of model misalignment will be able to flag the incident for the attention of their internal safety and alignment teams. Those teams will then decide whether the incident merits immediate disclosure or requires additional investigation and/or whether any affected third parties may need to be consulted before alerting the public.

Not every example of an OpenAI model acting in an unintended way will generate a public report, the company said. Instead, OpenAI said it will “prioritize new mechanisms, meaningful changes in known behavior, and findings that challenge assumptions about safety or mitigation.”

At the same time, OpenAI said it “favors disclosure even when significance is uncertain” and that its policy could lead to the public discussion of examples that are “spurious and not part of a larger pattern or suggestive of future developments.” If a specific misalignment issue continues to persist “despite repeated efforts to mitigate it,” OpenAI said it will offer updates each time.

If an example is deemed not worthy of public disclosure, the originating employee can escalate the disagreement to the senior officials at OpenAI’s Safety Advisory Group and, in cases of extreme disagreement, with OpenAI leadership. Over time, OpenAI said, it “plan[s] to develop more objective disclosure criteria with other developers, external researchers, industry standards bodies, and regulators.”

The company’s announcement also makes passing reference to the heavily discussed concept of “pacing” further AI development to allow more time for alignment research. “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” OpenAI wrote.

Bybit

Be the first to comment

Leave a Reply