Google Android Bench 2.0: long-horizon tasks; GPT-6 Astra leads at 28%
Google’s Sept 17 Android Bench 2.0 adds multi-day long-horizon tasks and continuous scoring; full-task pass rates top out around 28% (vs ~91% on the old incremental set). GPT-6 Astra leads the leaderboard; the suite also scores Gemini Flash, Fable 5.1, Kimi K3, and Qwen, with agent runs on Codex and Antigravity.