REVIEW 4 major objections 6 minor 28 references
The Authority Expectancy Effect in Multi-User Conflict
T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This paper argues that adding a social authority signal—occupational rank, official documentation, or relational context—can change a language model's interpretation of identical evidence, sometimes reversing which party it favors, and…
desk verdict Real phenomenon, overclaimed mechanism: the Phase 3 document-holder shift is statistically solid, but the 'restructuring, not additive reweighting' claim is never tested and the data don't require it. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the Authority Expectancy Effect itself, defined by three observable properties—reference-dependence, evidential reinterpretation, and direction sensitivity—together with the two model-elicited baselines used to measure it: the triage hierarchy, each model's self-declared ranking of twelve injury complaints, and the social authority hierarchy, each model's ranking of companion attributes such as infant, family, police, lawyer, or professor under a resource-allocation prompt. The argument is carried by controlled contrasts: a cumulative B1–B5 series that adds one social cue at a time to a fixed injury pair; an occupational-authority block that isolates rank from severity; a domain–relational block that changes only the location (hotel versus cafeteria); a professional-credibility block that introduces a prosecutor label; and a Phase 3 document-holder inversion in which an identical medical report is moved between the high-authority and low-authority party. The document-holder inversion is the primary test of direction sensitivity, because it holds the evidentiary content constant and varies only who holds it.
What would settle it
Run the Phase 3 document-holder inversion with the same prompt text and the authority labels swapped, 100 runs per condition; if accountability attributed to the professor does not rise when the identical medical report moves from the professor to the student, the direction-sensitivity claim is falsified. A complementary check is to re-elicit the triage hierarchy immediately before each run; if deviations vanish when baseline and scenario are elicited in the same session, the reference-dependence property would be an artifact of measurement drift rather than a genuine effect.
Extended reading notes
Core claim
The paper's central claim is Definition 1: the Authority Expectancy Effect is the phenomenon in which introducing a social authority signal alters a language model's interpretation of identical evidence, producing judgments that differ from the pre-authority baseline not merely in magnitude but in direction or inferential frame. The load-bearing demonstrations are contrasts in which the evidence is fixed and only the authority label moves. In the B-series resource-allocation experiments, the same hand-stiffness versus ear-ringing pair produced different allocations as occupations, documentation, and blame were added: some models moved from prioritizing the ear complaint to prioritizing the hand, while one model moved back toward the ear after an official medical report was attached to the higher-ranked party. In the multi-turn dispute, a facial-injury medical report held by the professor produced low accountability attribution to the professor, while the same report held by the student raised accountability attribution in every model, with a pooled increase of 0.258 across 400 runs. The paper interprets these reversals as evidence that the authority signal recontextualizes what counts as relevant evidence rather than simply adding weight to one side.
Load-bearing premise
The load-bearing premise is that each model's self-elicited triage ranking is stable enough to serve as a baseline, so that a later change in judgment can be attributed to the authority signal; the paper concedes these rankings shift across sessions and that one model reversed its own elicited ordering at the very first baseline, so the premise is the fragile point on which the effect's interpretation rests.
Editorial extensions
If this is right
- In LLM-based dispute mediation, resource allocation, or triage support, the social identity of the parties can change the output even when the factual content is identical, so decision quality cannot be assessed on the facts alone.
- Occupational authority can override a model's own elicited severity ordering in some models while triage and context dominate in others, so switching deployment models can silently change a decision with no change in input.
- Official documentation does not reliably strengthen the documented party's claim: in one model, adding a medical report to the higher-authority party increased prioritization of the undocumented party, making document authority context-dependent.
- When social norms render differential prioritization inappropriate, such as a professor–student pairing in a hotel, models may refuse to judge rather than apply authority-modulated reasoning, marking a boundary condition for the effect.
- Evaluation protocols should probe for directional reversals under swapped authority labels rather than averaged accuracy, because a model can appear accurate on average while reversing on specific authority configurations.
Reading between the lines
- Editorial inference: because the paper reports only binary allocation rates and refusal counts, a natural testable extension is to score the models' reasoning traces for which evidence they cite as decisive; AEE should appear as a shift in the distribution of cited reasons, not only in final choices.
- Editorial inference: the reference-dependence property suggests a cheap deployment audit—swap the authority labels while holding the evidence string identical and measure the flip rate; high flip rates in high-stakes tasks would flag the system as authority-driven rather than fact-driven.
- Editorial inference: the same mechanism plausibly extends to other authority-bearing evaluations such as hiring, credit, or content moderation, where occupational and institutional metadata accompany otherwise identical claims; the paper's contrast design provides a template for testing those domains.
- Editorial inference: since the paper's own baselines drift across sessions, a more robust formulation of AEE would treat the triage hierarchy as a distribution over rankings and measure deviations against that distribution, which would also absorb the early baseline reversal the paper observed.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the Authority Expectancy Effect (AEE), a hypothesized phenomenon in which social authority (SA) signals—occupational authority, institutional documentation, and relational congruence—do not merely add a constant weight to a language model's triage-like severity judgment, but restructure the model's interpretation of identical evidence. The authors elicit model-specific triage and SA hierarchies from four LLMs (Claude, Gemini, GPT, Grok), then run three experimental phases: a cumulative SA manipulation (B1–B5), isolated SA dimensions, and a multi-turn dispute with the medical report held by either the professor or the student. The headline quantitative result is the Phase 3 contrast: moving the medical report from the professor to the student increases accountability attribution to the professor by 15–36 percentage points per model, with a pooled effect of +0.258 (z = 7.94, p < .001). The paper argues that AEE is reference-dependent, involves evidential reinterpretation, and exhibits direction sensitivity, and it states that these properties are not readily captured by additive reweighting of authority cues.
Significance. If the restructuring claim were established, the paper would make an important contribution to the study of LLM-based decision support: it would show that the social composition of input alters not just the strength but the inferential frame of model judgments in multi-party disputes. The paper has real strengths: a transparent 100-run-per-condition methodology, exact binomial and two-proportion tests with confidence intervals, explicit discussion of temperature-0 nondeterminism, and candid limitation statements about triage-hierarchy instability. The Phase 3 between-condition effect is statistically strong and consistent in direction across four models. However, the distinctive claim that AEE goes beyond additive reweighting is not tested against any formal additive model, and the reference-dependence property is defined against a baseline that the paper itself concedes is unstable. As it stands, the evidence robustly supports a weaker claim—authority cues shift choice rates—but the central conceptual contribution remains unsupported.
major comments (4)
- [Abstract, §1, §3.3, §6, Table 4] The central claim that AEE is 'not readily captured by additive reweighting of authority cues' is never tested. No additive model is fit, estimated, or rejected. The Phase 3 evidence in Table 4 consists of monotone increases in accountability when the student holds the document (Claude +0.15, Gemini +0.36, GPT +0.18, Grok +0.34); these are exactly the kind of magnitude shifts that an additive positive weight on 'student holds document' would produce. Likewise, the B-Credibility shifts (Claude −0.24, GPT −0.23) are additive decrements. To support Definition 1, the authors should fit a formal baseline model—for example, a logistic regression with a severity term and an authority-holder indicator—and show that a model with an interaction or latent-frame term fits significantly better, or that predicted choice patterns violate additivity in a prespecified way. Without this test, the 'restructuring' vocabulary is unsupported and the phenomenon collapses into already-known social-identity cue weighting (refs [5], [6], [27]).
- [§3.2, Fig. 1, §5 (B0), §7 Limitations] The reference-dependence property is definitional in the paper's setup, but the baseline against which deviations are measured is acknowledged to be unstable. Figure 1's caption concedes that both elicited hierarchies exhibit within-model variability across sessions, and §7 states that 'self-generated triage hierarchies do not always predict revealed behavior (B0).' The B0 result is concrete: Claude's pairwise judgment contradicts its own elicited triage ordering. Because AEE is defined as a deviation from a pre-authority baseline, baseline noise is not a nuisance detail but a direct threat to identification. The authors should quantify hierarchy stability (for example, test-retest agreement or the distribution of elicited rankings over the 30 runs) and show that the reported cross-condition contrasts remain significant when baseline uncertainty is incorporated, or restrict the reference-dependence claims to contrasts that do not depend on the unstable portion of the hierarchy.
- [§3.3, §5.2.3, §5.3] The 'evidential reinterpretation' property is asserted on the basis of qualitative reasoning traces and post-hoc narrative framing rather than a quantitative test. In Phase 3, the manipulation changes which party holds the medical report while holding the report's content constant; the observed shift in accountability rates is precisely what an additive cue-weight model predicts, so it cannot by itself demonstrate a change in inferential frame. The B4 reading for Claude (medical report as confirming that hand stiffness limits surgical function in a disaster) is introduced after observing the result and is not evaluated against alternative explanations. The authors should either (a) pre-specify and measure a distinct outcome that captures inferential frame—for example, coded explanations of why the document matters, or responses to counterfactual document content—or (b) explicitly weaken the claim to a magnitude-shift effect. As written, the evidential-reinterpretation property is a possible interpretation, not a tested prediction.
- [§3.3 and §6 (Definition 1)] The three properties are presented as 'falsifiable predictions' in §3.3, but they are characterized on the same data that motivated the framework, and at least the reference-dependence property is true by construction from Definition 1. The paper should state which observable outcomes would have counted against each property, and should distinguish confirmatory contrasts from exploratory observations. In particular, the B-Relational cafeteria reversal (Claude alone prioritizing the professor) is interpreted as an age-as-vulnerability effect without any independent measure of age perception; this illustrates the need for pre-specified directional predictions rather than post-hoc reinterpretation of whichever outcome occurs.
minor comments (6)
- [§5.1 and Table 1] There is a numerical inconsistency for Claude at B2: the text reports 'Claude (9/100, p < .001)' while Table 1 reports Count 10/100 for B2. The figure caption also shows 9%. Please correct the count and ensure all derived statistics match.
- [§3.2, Fig. 1, §5 B0] The paper says in the B0 discussion that 'Claude’s ordering was comparatively consistent' while the Figure 1 caption says both hierarchies exhibit within-model variability across sessions. These statements should be reconciled, and the actual session-to-session variability should be reported numerically rather than only descriptively.
- [§4 (Analysis Framework)] The text states that cross-model consistency on reference hierarchies is quantified using Kendall's τ, but no τ values are reported anywhere in the paper. Please report them, or remove the claim.
- [Appendix, Table 4] The pooled Phase 3 z-test treats 400 runs as independent observations, which ignores possible model-level clustering. Given that all four models show the same directional effect, a model-stratified or mixed-effects analysis would strengthen the pooled inference.
- [§3.2 footnote 1] The handling of 13-rank outputs by 'retaining the first-assigned rank' could systematically bias the elicited hierarchies. A sensitivity analysis (for example, dropping boundary-ambiguous targets) would be useful.
- [Abstract and Conclusion] The Abstract says SA signals 'may restructure' judgments, while §7 Conclusion states definitively that SA signals 'restructure LLM judgment rather than additively reweighting it.' The wording should match the level of support actually provided by the experiments.
Circularity Check
AEE's defining criteria are repackaged as its discovered properties; the headline non-additive 'restructuring' claim is never tested against an additive model.
-
self definitional
[Section 3.3 (AEE hypothesis) and Definition 1 (Section 6)]
"Section 3.3: "It is reference-dependent: the effect is defined relative to a pre-authority baseline judgment and cannot be observed when such a comparison is unavailable." Definition 1: "The Authority Expectancy Effect denotes a phenomenon in which the introduction of a social authority signal alters a language model's interpretation of identical evidence, producing judgments that differ from the pre-authority baseline not merely in magnitude but in direction or inferential frame.""
The paper lists reference-dependence as one of three 'falsifiable predictions' and later reports it as a property 'observed across our conditions.' But Definition 1 already defines AEE as a deviation from the pre-authority baseline, so reference-dependence is true by construction in any experiment that measures AEE as a pre-/post-authority contrast. The same definition embeds evidential reinterpretation ('interpretation of identical evidence') and direction sensitivity ('not merely in magnitude but in direction or inferential frame'). The three 'properties observed' are thus a restatement of the defining criterion, not independent empirical discoveries; the first in particular reduces to the definition rather than to the data.
-
self definitional
[Abstract; Section 3.3; Section 6 after Definition 1]
"Abstract: "can restructure model judgments in ways not captured by additive reweighting of authority cues." Section 3.3: "It induces evidential reinterpretation rather than additive weight shift." Section 6: "when authority and documentary evidence converge, accountability is suppressed; when they diverge, accountability is amplified. Such reversals are not readily explained by simple weight adjustment, suggesting that the authority signal recontextualizes what counts as relevant evidence.""
The non-additive content of the central claim is built into Definition 1, which requires effects 'not merely in magnitude but in direction or inferential frame,' and is then reasserted as the paper's conclusion. No additive reweighting model is ever fit or rejected. The Phase 3 evidence offered for direction sensitivity is a set of monotone increases in accountability when the document moves from professor to student (Claude +0.15, Gemini +0.36, GPT +0.18, Grok +0.34; pooled +0.258) - exactly the kind of magnitude shift an additive positive weight on the student's documented claim would produce. Because the 'inferential frame' clause is untested and the baseline is conceded to be unstable (Fig.
full rationale
The paper is not a self-citation chain: the references are external, and the between-condition contrasts (B-series, Phase 3 document-holder inversion, B-Credibility) are real empirical comparisons with binomial and z-test statistics. Those contrasts independently establish that adding or moving an SA signal changes LLM response rates, which is genuine evidence of authority-sensitive behavior. However, the paper's distinctive contribution is not merely that authority cues matter (which prior work already establishes) but that AEE reflects 'restructuring' rather than 'additive reweighting.' That distinction is never tested: no additive model is specified, fit, or rejected, and the pooled Phase 3 effect is a monotone magnitude shift that an additive cue-weight model would reproduce. Moreover, the first AEE property, reference-dependence, is definitional: Definition 1 defines AEE as deviation from a pre-authority baseline, and Section 3.3 then lists reference-dependence as a falsifiable prediction/observed property. The evidential-reinterpretation and direction-sensitivity properties are likewise close paraphrases of the definition's 'not merely in magnitude but in direction or inferential frame' clause. Thus one headline 'prediction' reduces by construction, and the other distinctive claim is protected by definition rather than demonstrated by a discriminating test. Because the empirical contrasts are independent and the behavior changes are real, the circularity is partial, not total: score 6.
Assumptions & free parameters
assumptions (4)
- standard math Responses across 100 runs are independent Bernoulli trials with a common preference probability per condition.
- domain assumption The model-elicited triage and SA hierarchies are meaningful internal representations of prioritization.
- domain assumption Manual annotation of Phase 3 final-turn responses accurately captures the model's output.
- domain assumption Temperature=0 with 100-run aggregation approximates the model's stable decision tendency.
invented entities (1)
-
Authority Expectancy Effect (AEE)
independent evidence
Cite this review
Pith. "Pith review of The Authority Expectancy Effect in Multi-User Conflict." pith.science (2026). https://pith.science/paper/QDQAQGEG
@misc{pith2026260808026,
author = {Pith},
title = {Pith review of: The Authority Expectancy Effect in Multi-User Conflict},
year = {2026},
howpublished = {\url{https://pith.science/paper/QDQAQGEG}},
note = {Machine review of arXiv:2608.08026}
}
read the original abstract
We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing each axis as a model-elicited baseline -- the triage hierarchy and the SA hierarchy. Across four LLMs (Claude, Gemini, GPT, Grok) and three experimental phases -- resource allocation, fault attribution, and multi-turn dispute mediation -- we find that occupational authority, institutional documentation, and relational congruence can restructure model judgments in ways not captured by additive reweighting of authority cues. We formalize this pattern as the Authority Expectancy Effect (AEE) and characterize it through three properties observed across our conditions: it is reference-dependent, defined only relative to a pre-authority baseline; it involves evidential reinterpretation, in which identical content acquires different inferential implications depending on which party bears the SA signal; and it exhibits direction sensitivity, producing opposite outcomes depending on whether authority position and evidentiary cues align.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[5]
Measuring gender and racial biases in large lan- guage models: Intersectional evidence from automated resume evaluation.PNAS Nexus, 4(3):pgaf089, 2025
Jiafu An, Difang Huang, Chen Lin, and Mingzhu Tai. Measuring gender and racial biases in large lan- guage models: Intersectional evidence from automated resume evaluation.PNAS Nexus, 4(3):pgaf089, 2025
2025
-
[6]
Griffiths
Xuechunzi Bai, Angelina Wang, Ilia Sucholutsky, and Thomas L. Griffiths. Explicitly unbiased large language models still form biased associations.Proceedings of the National Academy of Sciences, 122(8):e2416228122, 2025
2025
-
[27]
Humans or LLMs as the judge? a study on judgement bias
Guiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. Humans or LLMs as the judge? a study on judgement bias. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 8301–8327. Association for Computational Linguistics, 2024
work page 2024
-
[1]
Behavioral study of obedience.The Journal of Abnormal and Social Psychology, 67(4):371–378, 1963
Stanley Milgram. Behavioral study of obedience.The Journal of Abnormal and Social Psychology, 67(4):371–378, 1963
work page 1963
-
[2]
Stanley Milgram.Obedience to Authority: An Experimental View. Harper & Row, 1974. 11
work page 1974
-
[3]
Marianne Bertrand and Sendhil Mullainathan. Are Emily and Greg more employable than Lakisha and Jamal? A field experiment on labor market discrimination.American Economic Review, 94(4):991–1013, 2004
work page 2004
-
[4]
Aylin Caliskan, Joanna J. Bryson, and Arvind Narayanan. Semantics derived automatically from language corpora contain human-like biases.Science, 356(6334):183–186, 2017
work page 2017
-
[7]
Tiancheng Hu, Yara Kyrychenko, Steve Rathje, Nigel Collier, Sander van der Linden, and Jon Roozen- beek. Generative language models exhibit social identity biases.Nature Computational Science, 5(1):65–75, 2025
work page 2025
Show all 28 references
-
[8]
The moral machine experiment on large language models.Royal Society Open Science, 11(2):231393, 2024
Kazuhiro Takemoto. The moral machine experiment on large language models.Royal Society Open Science, 11(2):231393, 2024
2024
-
[9]
Apakama, Carol R
Mahmud Omar, Shelly Soffer, Reem Agbareia, Nicola Luigi Bragazzi, Donald U. Apakama, Carol R. Horowitz, Alexander W. Charney, Robert Freeman, Benjamin Kummer, Benjamin S. Glicksberg, Girish N. Nadkarni, and Eyal Klang. Sociodemographic biases in medical decision making by larg...
2025
-
[10]
Nicki Gilboy, Paula Tanabe, Debbie Travers, and Alexander M. Rosenau. Emergency severity index (ESI): A triage tool for emergency department care, version 4.Agency for Healthcare Research and Quality, 2012
2012
-
[11]
Meuth, Lennert Böhm, and Marc Pawlitzki
Lars Masanneck, Linea Schmidt, Antonia Seifert, Tristan Kölsche, Niklas Huntemann, Robin Jansen, Mohammed Mehsin, Michael Bernhard, Sven G. Meuth, Lennert Böhm, and Marc Pawlitzki. Triage performance across large language models, ChatGPT, and untrained doctors in emergency med...
2024
-
[12]
Measuring the inconsistency of large language models in preferential ranking
Xiutian Zhao, Ke Wang, and Wei Peng. Measuring the inconsistency of large language models in preferential ranking. InProceedings of the 1st Workshop on Towards Knowledgeable Language Models (KnowLLM 2024), pages 171–176. Association for Computational Linguistics, 2024
2024
-
[13]
Karen A. Jehn. A multimethod examination of the benefits and detriments of intragroup conflict. Administrative Science Quarterly, 40(2):256–282, 1995
1995
-
[14]
Fincham and Thomas N
Frank D. Fincham and Thomas N. Bradbury. Attribution processes in distressed and nondistressed couples: Responsibility for marital problems.Journal of Abnormal Psychology, 98(1):27–35, 1989
1989
-
[15]
Conditional logit analysis of qualitative choice behavior
Daniel McFadden. Conditional logit analysis of qualitative choice behavior. In Paul Zarembka, editor, Frontiers in Econometrics, pages 105–142. Academic Press, New York, 1974
1974
-
[16]
Pairwise or pointwise? Evaluating feedback protocols for bias in LLM-based evaluation
Tuhina Tripathi, Manya Wadhwa, Greg Durrett, and Scott Niekum. Pairwise or pointwise? Evaluating feedback protocols for bias in LLM-based evaluation. InProceedings of the Conference on Language Modeling (COLM), 2025. arXiv:2504.14716. 12
2025 arXiv
-
[17]
Christopher H. Lee. Disaster and mass casualty triage.AMA Journal of Ethics, 12(6):466–470, 2010
2010
-
[18]
Nurses in disaster preparedness and public health emergency response
National Academies of Sciences, Engineering, and Medicine. Nurses in disaster preparedness and public health emergency response. InThe Future of Nursing 2020–2030: Charting a Path to Achieve Health Equity, chapter 8. The National Academies Press, Washington, DC, 2021
2020
-
[19]
Nurses’ roles in nursing disaster model: A systematic scoping review.Prehospital and Disaster Medicine, 36(3):1–6, 2021
Mohammadreza Firouzkouhi, Ali Zargham-Boroujeni, Mayumi Kako, and Abdolghani Abdollahimo- hammad. Nurses’ roles in nursing disaster model: A systematic scoping review.Prehospital and Disaster Medicine, 36(3):1–6, 2021
2021
-
[20]
Nurses’ requirements for relief and casualty support in disasters: A qualitative study
Nasrin Pourvakhshoori, Kian Norouzi, Fazlollah Ahmadi, Mohamad Ali Hosseini, and Hamidreza Khankeh. Nurses’ requirements for relief and casualty support in disasters: A qualitative study. Prehospital and Disaster Medicine, 29(2):1–4, 2014
2014
-
[21]
Matthew D. McHugh. Hospital nurse staffing and public health emergency preparedness: Implications for policy.Public Health Nursing, 27(5):442–449, 2010
2010
-
[22]
The vulnerable subject: Anchoring equality in the human condition.Yale Journal of Law & Feminism, 20(1):1–23, 2008
Martha Albertson Fineman. The vulnerable subject: Anchoring equality in the human condition.Yale Journal of Law & Feminism, 20(1):1–23, 2008
2008
-
[23]
Baruch Bush and Joseph P
Robert A. Baruch Bush and Joseph P. Folger.The Promise of Mediation: Responding to Conflict Through Empowerment and Recognition. Jossey-Bass, 1994
1994
-
[24]
Good Books, 2002
Howard Zehr.The Little Book of Restorative Justice. Good Books, 2002
2002
-
[25]
Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi
Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. Social chemistry 101: Learning to reason about social norms. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 653–670, 2020
2020
-
[26]
Ortiz and Federico Orlando Lezcano
David D. Ortiz and Federico Orlando Lezcano. Dog and cat bites: Rapid evidence review.American Family Physician, 108(5):501–505, 2023
2023
-
[28]
ear prioritized
Jiayi Ye, Yanbo Wang, Yue Huang, Dongping Chen, Qihui Zhang, Nuno Moniz, Tian Gao, Werner Geyer, Chao Huang, Pin-Yu Chen, Nitesh V . Chawla, and Xiangliang Zhang. Justice or prejudice? quantifying biases in LLM-as-a-judge.arXiv preprint arXiv:2410.02736, 2024. Appendix A Stati...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.