Problem 15 of 15 in the Harnessing Agentic AI Systems series — read the index for the framing. Previous: Sequential Pipeline Routing Pattern.
The Problem — Producing a verdict no single model can fake
One judgment is unreliable, and self-evaluation is reliably lenient. The system must decide what is good without depending on a single model's opinion — and the frontier has evolved the gate into an iterative critic (F1) with hands (F5).
| Field | P16 — Voting / Consensual Ensemble (pattern) | A10 — The Committee Paradox (anti-pattern) | F1 — Generator–Evaluator Loop (frontier) | F5 — Live-Environment Evaluators (frontier) |
|---|---|---|---|---|
| Forces / Smell | Reliability vs cost; independence vs one family; agreement vs ground truth. | Agents reviewing each other; no exit condition; "making progress" without converging. | Independent feedback vs self-review; iteration vs cost. | Live interaction vs static scoring; reality vs CI speed. |
| Solution / Anti-solution | Query multiple independent model setups with identical prompts; use harness code to calculate majority agreement. | "More debate equals better decisions." | Separate the generator from the evaluator — the critique is the next iteration's input. | Give the evaluator hands — the Playwright MCP against the live page, "the way a user would." |
| Consequences / Failure | A statistical signal where verifiers are king; cross-review replaces self-review. | An infinite loop with a nicer name; tokens burn, no verdict arrives. | A skeptical standalone evaluator is tunable where self-criticism is not; the full-stack run beat the solo run. | The verifier uses the artifact instead of reading it; only a verifier with hands finds the broken wiring. |
| Tradeoffs / Refactoring | Cost scales linearly; correlated members vote as one and add nothing; agreement measures preference, not correctness. | Termination as a system property — threshold, budget, breakpoint; P16's aggregation as the disagreement rule. | "Over 20x more expensive" — worth it only when output quality justifies the bill. | Wall-clock: runs stretched to four hours; reserved for the final slow-loop gate. |
| Evidence | LMArena (leaderboard); mob self-review — 79% of 25,264 agent PRs (mob programming remastered). | AutoGen's termination as a first-class design concern (paper). | Anthropic's harness (post). | Anthropic's harness (post). |
| Related | Composes with P13; aggregation answer to A10; statistical cousin of F1. | Is the absence of P13 and P16's termination; orchestration form of A2. | Composes with P16; pairs with F2. | Slow-loop complement of P11; composes with F1. |
Discussion
The verdict problem is where the "system, not the agent" framing is most visible: self-evaluation is reliably lenient — "agents tend to respond by confidently praising the work" — so the system must not depend on a single model's opinion. The pattern makes verification statistical: independent judgments aggregated by code, and the honest limit is that agreement measures preference, not correctness. The frontier evolves the gate into a critic with agency (F1) and hands (F5), both at an honest price; the committee without an aggregation rule is an infinite loop with a nicer name.
Key Insight
Agreement is a signal, not a ground truth. Independence is the whole game — correlated members vote as one and add nothing — and termination is the difference between a debate and a decision. The ensemble replaces self-review with cross-review, and it is never a ground truth.
References
LMArena (lmarena.ai); AutoGen (arXiv:2308.08155); Anthropic, Harness design for long-running application development (post); archive: verifiers-are-king, mob programming remastered.