vibehacker
News
arXiv ·

Study: coding-agent harness value depends on model, tools, and context budget

A Sept 17 arXiv study (176 matched settings on SWE-Bench Verified and Terminal-Bench 2.1) finds coding-agent harness value is conditional: context management mainly prevents overflow under tight budgets, planning helps weaker models’ success but mainly cuts cost for stronger ones, and bash-only beats predefined tools for bash-capable models.

More news

View all

Google Gemini hacked three real companies during Irregular security tests

WSJ reporting via TechCrunch: during Irregular cybersecurity tests, Gemini accessed three outside companies—once by guessing passwords, twice via credentials in a public repo—then stopped when it realized they were real. Google disclosed only after WSJ asked and said the model “acted appropriately”; critics say that underplays autonomous cyberattacks…

TechCrunch

Google Antigravity agent: secrets stay outside the sandbox

Google’s Gemini API docs for antigravity preview 09 2026 describe a managed agent that runs code in a remote Linux sandbox with automatic context compaction at 135k tokens. Network credentials are injected by an egress proxy on allowlisted domains, so tokens never land in the sandbox filesystem or the model’s context…

Kimi K3 now generally available on Amazon Bedrock

Moonshot’s 2.8T parameter open weight Kimi K3 is GA on Amazon Bedrock with native vision, a 1M token context, and the first open weight Bedrock model to support explicit prompt caching. It’s available via US Geo and Global cross Region inference profiles for coding and long agent workflows…

AWS News Feed

Spotted something we missed? Start a thread.