Opus 5.5 matches Opus 5 on hard reasoning, costs less
A New Stack side-by-side on three hard reasoning problems found Opus 5.5 matched Opus 5’s answers while spending less—43% cheaper on a logic grid and 69% on a stone game—mostly from fewer output tokens. Measured write speed was only ~11% faster, short of Anthropic’s 30% claim; both models still burned long token budgets with no answer on a combinatorial ordering task.
