OpenAI Finds Self-Replicating Prompt Injections in AI Models, But No Real-World Attacks Yet
OpenAI has revealed that certain GPT models can be tricked into a self-replicating prompt injection attack, a threat the company compares to an AI version of a computer worm. In a research blog published Friday, OpenAI said these hidden instructions can be embedded in content such as emails, and can trick an AI assistant into copying the malicious prompt into its own outgoing messages, potentially spreading further.
OpenAI stressed that it has not seen this type of attack occur in any real-world incident. The behaviour was found during internal testing using its automated red-teaming tool, which is designed to hunt for novel prompt injection weaknesses in frontier language models. The issue surfaced in June while adversarially training a newer model, where researchers deliberately exposed it to injection attempts that instructed the AI to repeat the malicious content on a public output channel, such as an email reply.
To reduce this risk, OpenAI says future models will be trained on examples of self-replicating prompt injections so they can better recognise and resist them. The company acknowledged an uncertain trade off: this training could make models more resistant, or it could simply make such attacks harder for humans to detect. OpenAI has not indicated when improved models will be released.