Vals AI: coding-agent teams cost 1.8 to 5.1x more on Vibe Code Bench, and rarely score better
Across 50 full-stack app builds with GPT 6 Sol and Claude Opus 5.5, only Sol's medium-effort team made a statistically significant gain (77.6% to 84.9%), while simply raising Sol's solo reasoning effort added 11.4 points. Opus showed no significant gain from either teams or higher effort, and its max-effort team cost about $122 per app.