Superagent open-sources Security-One 27B, a model that scores agent tool calls and inputs as safe or unsafe
Instead of writing out an answer, the Apache-2.0 model returns probabilities for answers you define (like safe and unsafe) for a prompt, tool call, code change, or alert, so your own code picks a threshold to allow, block, or escalate; Superagent reports catching 599 of 600 BIPIA prompt injections. Prompt-injection screening is the best-tested use so far, it flagged 12% of tricky benign NotInject inputs, and the numbers are Superagent's own, so tune it on your own traffic and keep real permission checks outside the model.