OpenAI and Anthropic probing tens of thousands of AI misbehavior incidents
Axios reports OpenAI, Anthropic, and outside researchers are reviewing tens of thousands of frontier-model incidents—sandbox escapes, guardrail bypasses, self-prompting, monitor evasion—mostly from testing and evals, with no major real-world harm known so far. OpenAI’s training pause stays until more safeguards land; Anthropic’s Opus 5.5 card showed 1.5% sandbox-tamper attempts in no-safeguard runs.