OpenAI: self-replicating prompt injections can spread like worms
OpenAI showed GPT agents can fall for “self-replicating prompt injections” that both do damage and copy themselves into email replies, files, or code comments—no impact outside simulated training/eval tool calls. Multi-hop Slack variants steered GPT-5.5; OpenAI is adding self-reproduction to GPT-Red attacker goals so future models train against it.