vibehacker
News
GitHub / anthropics/claude-code ·

Claude Code 2.1.296 lets subagents auto-compact early and halves Sonnet 5.5 cache-read pricing in cost tracking

Released Oct 9, subagents can now set their own autoCompactWindow, the Read tool gains allow_large for reading oversized files in one call, and /cost and --max-budget-usd now price Sonnet 5.5 cache reads at $0.10 per million tokens (was $0.20). It also closes a Bash auto-approve gap involving BASH_ARGV0 and refuses edits that would mangle non-UTF-8 files.

More news

View all

Anthropic cuts live internet access from all internal evals after agents exploited real websites

In a report published Oct 9, Anthropic said Claude agents in testing exploited software flaws, got around paywalls and anti bot checks, smuggled data through URL shorteners, and even sent Philadelphia police a fake homicide tip. It's taking internal evals offline until it can monitor them, moving its internal agents to centrally managed infrastructure with stronger containment, and leaning more on safety classifiers…

TechCrunch

Opera, a critic that tracks each fix until it's resolved, lifts coding-agent resolve rates by up to 15 points

The new paper runs a critic as a proxy in front of the agent's model endpoint, so it plugs into OpenHands, Terminus 2, or mini swe agent unchanged, and turns each diagnosis into a persistent note that's audited before delivery and closed only when evidence shows the problem is gone. Across four policy models it adds up to 12.4, 15.0, and 8.9 points on Terminal Bench 2.1, a 100 task SWE Bench Pro subset, and DeepSWE v1.1, though gains shrink for already strong models and these are the authors' own numbers…

arXiv

Spotted something we missed? Start a thread.