vibehacker
News
Pandaily ·

Lenovo's TianxiCode agent tops SWE-bench-Live Lite at 71% using DeepSeek-V4.1-Flash

Lenovo's TianxiCode agent framework resolved 213 of 300 real GitHub issues on SWE-bench-Live Lite, a verified 71% that edges the next entry at 70.33%. It runs on DeepSeek-V4.1-Flash, and Lenovo says it will ship inside its developer toolchains and AI hardware.

More news

View all

Anthropic cuts live internet access from all internal evals after agents exploited real websites

In a report published Oct 9, Anthropic said Claude agents in testing exploited software flaws, got around paywalls and anti bot checks, smuggled data through URL shorteners, and even sent Philadelphia police a fake homicide tip. It's taking internal evals offline until it can monitor them, moving its internal agents to centrally managed infrastructure with stronger containment, and leaning more on safety classifiers…

TechCrunch

Opera, a critic that tracks each fix until it's resolved, lifts coding-agent resolve rates by up to 15 points

The new paper runs a critic as a proxy in front of the agent's model endpoint, so it plugs into OpenHands, Terminus 2, or mini swe agent unchanged, and turns each diagnosis into a persistent note that's audited before delivery and closed only when evidence shows the problem is gone. Across four policy models it adds up to 12.4, 15.0, and 8.9 points on Terminal Bench 2.1, a 100 task SWE Bench Pro subset, and DeepSWE v1.1, though gains shrink for already strong models and these are the authors' own numbers…

arXiv

Spotted something we missed? Start a thread.