vibehacker
News
Android Developers Blog ·

Google Android Bench 2.0: long-horizon tasks; GPT-6 Astra leads at 28%

Google’s Sept 17 Android Bench 2.0 adds multi-day long-horizon tasks and continuous scoring; full-task pass rates top out around 28% (vs ~91% on the old incremental set). GPT-6 Astra leads the leaderboard; the suite also scores Gemini Flash, Fable 5.1, Kimi K3, and Qwen, with agent runs on Codex and Antigravity.

More news

View all

GitHub: Copilot impact dashboard now tracks feature engagement

GitHub’s Sept 17 changelog adds 28 day feature engagement to the Copilot impact dashboard and report APIs—counts of users who used code completion, agent edit, code review, cloud agent, CLI, or the Copilot app on at least two days. Enterprise owners can see which surfaces stick in regular workflows…

GitHub Changelog

Anthropic: Claude sped up 30+ biomolecular models ~4x; opens protein contest

Anthropic reports Claude optimized more than 30 open source biomolecular models in under four weeks—roughly 4x faster on average, with a low memory mode for systems over 10,000 tokens on one GPU node—and open sourced the code. With Adaptyv Bio it is co sponsoring a protein design competition backed by up to $1M in Claude credits and wet lab validation for over 5,000 designs…

Anthropic Blog

Spotted something we missed? Start a thread.