Why a second opinion works
This sounds like folk wisdom. It is not. It is among the most replicated findings in agent evaluation, and it has a name.
An agent is a model plus a harness — the scaffolding that decides what context the model sees, which tools it can call, how it retries, when it summarizes, when it quits. The scaffolding turns out to dominate.
Hold the model fixed, swap the harness, and the score moves further than a model generation upgrade moves it.
The finding that makes a product possible: harnesses do not merely differ in aggregate performance. They fail on different problems.
Research on LLVM issue resolution found strong complementarity across LLMs and agents, with different techniques resolving genuinely distinct subsets of issues and an ensemble beating every individual approach. Work on ensembles for code generation and repair found the same across model families — including cases where a smaller model solves what its larger sibling cannot.
If failure sets overlapped almost completely, a second opinion would be worthless. They do not overlap.
Anyone can run four harnesses. They can’t strip the previous agent’s wrong theory out of the context, and they don’t know which harness to try first.