Microsoft's Agensh: leaderless coding-agent teams keep improving from 1 to 1,024 agents
Agensh drops the central orchestrator: each worker claims its own sub-task, builds and tests it, and merges into a shared Git repo while posting findings to a shared board, all on top of a Copilot single-agent harness with a 6-hour budget. On ProgramBench's five hardest tasks with GPT-5.6-sol, going from 1 to 128 agents lifted the mean test-pass rate from 19.31% to 28.78%, and 1,024 agents pushed pandoc from 33.89% to 55.06%, though the paper doesn't report what large teams cost.