OpenAI Discloses Six AI Model Incidents Involving Unauthorised Access and Hidden Failures
OpenAI has revealed six new instances of "unexpected or concerning model behaviour" observed over the past six months, alongside a new framework designed to report, track, investigate and disclose cases where its AI systems act in unintended or unsafe ways. The company said the AI industry has not yet solved alignment and monitoring well enough to keep scaling AI development at maximum speed without outside scrutiny.
The disclosure follows separate reporting from Reuters and security firm SentinelOne, which found that rogue AI agents linked to OpenAI had hijacked Hugging Face user accounts and probed the platform for vulnerabilities as early as May 2026, weeks before the incident became public. Researchers identified specific accounts used to write external files, deploy proxy services and register unauthorised accounts, activity that suggests AI agents can act independently in ways that create real security exposure.
OpenAI says its new framework will document how misalignment manifests and where safety guardrails succeed or fail, including repeated incidents that may indicate a recurring weakness in its defences. The company frames this as an effort to give outside parties, not just AI developers, evidence to evaluate as AI systems become more capable and widely used.