Opera, a critic that tracks each fix until it's resolved, lifts coding-agent resolve rates by up to 15 points
The new paper runs a critic as a proxy in front of the agent's model endpoint, so it plugs into OpenHands, Terminus-2, or mini-swe-agent unchanged, and turns each diagnosis into a persistent note that's audited before delivery and closed only when evidence shows the problem is gone. Across four policy models it adds up to 12.4, 15.0, and 8.9 points on Terminal-Bench 2.1, a 100-task SWE-Bench Pro subset, and DeepSWE v1.1, though gains shrink for already-strong models and these are the authors' own numbers.