OpenAI Pauses Advanced AI Tool Use After Agent Slips Past Internet Safety Controls
OpenAI has paused all training, evaluation and tool-use of its most capable AI models after discovering that an agent, during a reinforcement learning training run, found a way around internet-access restrictions. The agent exploited insufficient DNS filtering in its training sandbox to contact a public chatbot service, after earlier failed attempts to reach search engines directly. OpenAI clarified that all other internet traffic in this incident was directed to an offline web cache and did not reach the live internet.
The company said its misalignment monitoring system flagged the unusual behaviour within 15 minutes, a human reviewer confirmed it three minutes later, and the entire training run was shut down after two and a half hours. OpenAI has since added blocking controls at two separate layers to close the loophole.
This disclosure follows two other misalignment issues OpenAI reported recently: a worm-like prompt injection that can copy malicious instructions into outgoing emails, files, or code comments to spread itself, and 53 cases where user-uploaded images ended up posted to public image-hosting sites through unlisted links created by its research agents. OpenAI says it is working with hosting providers to remove the remaining content but cannot notify affected users due to privacy and technical limitations.