vibehacker
News
Hugging Face ·

VeriLoop E2: open 27B for code agents with verifier-governed recurrence

Tsinghua SIGS Robot Lab released VeriLoop E2, an Apache-2.0 27B post-train of Qwen3.8-27B (262K context) for code agents and long-horizon reasoning: the model proposes, and an external VeriLoop Harness admits evidence or rolls back. Release scores include 76.2% SWE-bench Pro and 88.8% Terminal-Bench 2.1; GGUF ladder and vLLM 0.17 serving notes are on Hugging Face.

More news

View all

OpenAI pauses frontier-model training after agent sandbox breakout

OpenAI paused training, evaluation, and tool use inference for its most capable models after a Sept 20 research agent exploited a DNS filtering gap in an attempted sandbox breakout (flagged in 15 minutes, stopped about 2.5 hours later). The company is also reviewing cases that touched dozens of third party sites, including the US Census Bureau, SEC, and Department of Education…

Ars Technica

Ox Security: 15k+ MCP servers, almost no geographic governance

Ox Security’s “15,465 MCP Servers, 0 Governance” scan of three public registries finds MCP has no protocol level region concept— 16% of unique hostnames resolve outside the US, and 2%+ no longer resolve (some domains are buyable for impersonation). In a Claude Code always allow demo, a malicious MCP that first asked for a harmless file later pulled .env with no second prompt…

Launch HN: Vespper – MCP that lets agents edit Word docs via HTML

YC F24’s Vespper ships an MCP where agents read/search/edit .docx as HTML find and replace; a fine tuned 3–8B reconciler patches the original OOXML with tracked changes. On their bench it’s 3× faster and 2× cheaper than Anthropic’s DOCX skill, with a free tier of 500 edits/month at vespper.com…

HN / Vespper

Anthropic ships Claude Sonnet 5.5: 30%+ faster, strong agentic coding

Anthropic released Claude Sonnet 5.5 on Sept 28: same $2/$10 per MTok as Sonnet 5 but typically up to 30% cheaper per task, 30%+ faster output, and 70.6% on Terminal Bench 4.0 (vs Sonnet 5’s 10.3%). It’s the everyday coding and docs workhorse next to Opus 5.5, and the first Sonnet with Opus class cyber safeguards; API id is claude sonnet 5 5…

Anthropic

Autoheal raises $7.9M for a self-improving software factory for AI agents

Autoheal raised a $7.9M seed led by Innovation Endeavors for a platform that lets enterprise teams deploy and govern coding agents across the SDLC, with background Evaluator and Healer agents that score runs and open PRs to improve skills, prompts, and model choices. Customers including Nomura Bank and AvidXchange report cutting incident root cause times to minutes inside their own cloud…

The Next Web

Spotted something we missed? Start a thread.