Peer endorsement spreads wrong answers through clinical LLM committees (38% text contagion), and only a referee that privately re-queries the holdout separates true adoption from honest agreement.
Anthropic Engineering Blog (2026),https://www.anthropic.com/engineering/ demystifying-evals-for-ai-agents
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Peer endorsement spreads wrong answers through clinical LLM committees (38% text contagion), and only a referee that privately re-queries the holdout separates true adoption from honest agreement.