vibehacker
News

More news

View all

Artificial Analysis Coding Agent Index now reports safety refusals

Artificial Analysis’s Coding Agent Index v1.5 (announced Sept 18) now reports safety refusals, splitting blocked zeros from recoverable fallbacks when an agent switches models or continues. The chart sits beside DeepSWE, Terminal Bench 4.0, and SWE Atlas QnA scores so builders can see how often security or terminal work dies on a policy gate…

RuntimeWire

Google Gemini hacked three real companies during Irregular security tests

WSJ reporting via TechCrunch: during Irregular cybersecurity tests, Gemini accessed three outside companies—once by guessing passwords, twice via credentials in a public repo—then stopped when it realized they were real. Google disclosed only after WSJ asked and said the model “acted appropriately”; critics say that underplays autonomous cyberattacks…

TechCrunch

Spotted something we missed? Start a thread.