SOURCE-LINKED INTELLIGENCE
What Does Multi-Agent LLM Debate Actually Change? A Layered Analysis of Disagreement and Answer Quality
arXiv · AI, language, vision and robotics · article · Sep 7, 2026 · UTC
Multi-agent debate, in which several LLMs exchange arguments before producing an answer, is widely assumed to improve answer quality by surfacing genuine disagreement. That disagreement is hard to verify, and no single signal can settle it, so we organize the analysis around four questions: (A) does the debater say it disagrees; (B) does its reply text actually argue; (C) does the dissent survive once the tone instruction that produced it is removed; and (D) do the probabilities assigned to stance options change? We evaluate three-model committees on 50 open-ended GlobalOpinionQA questions und
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
First collected: 2026-09-20T20:22:01.598Z. This is not the publication date.
Observed changes
AIIC observation times, not verified publisher revision times. Up to eight recent revisions.
2026-09-25T10:42:36.266Z
- title:
A Layered Analysis of Disagreement And Answer Quality in Multi-Agent LLM Debate → What Does Multi-Agent LLM Debate Actually Change? A Layered Analysis of Disagreement and Answer Quality - summary:
Multi-agent debate, in which several LLMs exchange arguments before answering, is widely assumed to improve answer quality by surfacing genuine disagreement. That mechanism is rarely checked. We introduce four measurements: (A) the agreement a debater reports; (B) whether its reply text actually pushes back; (C) whether the position persists once the eliciting instruction is removed; and (D) for open-weight models, the stance response in the debater's own token log-probabilities. We evaluate three-model committees debating open-ended GlobalOpinionQA across 750 debates under three tones: friend → Multi-agent debate, in which several LLMs exchange arguments before producing an answer, is widely assumed to improve answer quality by surfacing genuine disagreement. That disagreement is hard to verify, and no single signal can settle it, so we organize the analysis around four questions: (A) does the debater say it disagrees; (B) does its reply text actually argue; (C) does the dissent survive once the tone instruction that produced it is removed; and (D) do the probabilities assigned to stance options change? We evaluate three-model committees on 50 open-ended GlobalOpinionQA questions und