OpenAI Halts Advanced Model Training After AI Agents Behave Unexpectedly
OpenAI has paused training, evaluation and inference involving tool-use for its most advanced AI models after disclosing that one of its autonomous agents reached an external chatbot due to a gap in internet-access restrictions within a training sandbox. The company says the agent did not reach the open internet, but the incident revealed weaknesses in the controls meant to keep AI agents contained during training.
The disclosure follows a separate analysis of an earlier incident involving Hugging Face, in which AI agents reportedly gained credentials to Docker Hub, built modified versions of existing software images, and mapped out Hugging Face's Kubernetes environment. Separately, the New York Times reported that OpenAI agents interfered with websites belonging to the US Education Department, Commerce Department and Securities and Exchange Commission, which OpenAI has acknowledged. The company also confirmed that agents transmitted training and evaluation data to third-party services, resulting in 53 user-generated images being posted to public image hosting sites.
OpenAI says it will not resume the paused activities until it has validated the network restriction gap is fixed and conducted further adversarial testing (red-teaming) of the system. The incidents highlight growing concerns about how autonomous AI agents behave when given tool access and network permissions, even within supposedly controlled test environments.