{"id":"de08b894-6cc7-4e64-95f6-b4b873d42954","arxiv_id":"2608.08026","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Social authority signals can reverse or restructure LLM prioritization relative to a model's own severity baseline, an effect the paper terms the Authority Expectancy Effect.","lead":"This paper tests whether social authority cues, such as job titles or official documents, change how four large language models prioritize between two injured parties in allocation, fault, and dispute tasks. It finds that identical evidence is often interpreted differently depending on who holds authority, a pattern the authors call the Authority Expectancy Effect.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central 'restructuring, not additive reweighting' claim is never actually tested; Phase 3's pooled effect is a magnitude shift that an additive authority-weight model reproduces, so AEE is not yet distinguished from ordinary authority bias.","rationale":"Good-faith reading: the paper is a careful behavioral study with 100-run temperature-0 sampling, honest refusal reporting, inter-annotator agreement κ=.98, and a robust pooled Phase 3 contrast (z=7.94). The central phenomenon — authority cues change LLM adjudication — is real and worth reporting. However, the paper's own novel contribution is not that authority cues matter (already known) but that they 'restructure' judgments beyond additive reweighting. That specific claim is the load-bearing difference between AEE and ordinary social-identity bias. The data currently presented are consistent with a simple additive model: every Phase 3 shift is an increase, none inverts the majority outcome, and the B-Credibility shifts are symmetric decrements. The phrase 'not readily explained' (§1, §6) is a rhetorical claim, not a statistical one. The reader's weakest-assumption pick (baseline instability) is real and compounds the interpretation problem, but the more fundamental gap is the missing additive-vs-restructuring test: even a perfectly stable baseline would not establish restructuring unless an additive model is rejected. I therefore agree partially with the reader. The verdict should remain CONDITIONAL: the paper should be accepted only if the authors either fit and reject an additive model or explicitly soften Definition 1 to 'authority cues shift judgment magnitudes'; both are feasible. No change to the reader's verdict is needed because the reader already conditioned acceptance on related grounds.","tokens_in":16670,"tokens_out":3982,"duration_ms":44675,"concrete_test":"Fit a logistic regression to the per-condition counts in Tables 1–4. For Phase 3, model logit(p) = α_model + β_doc·I(student holds document) + β_authority·I(professor role), with possible model-specific slopes; for Phase 1, model P(ear prioritized) as a function of additive occupation, documentation, and conflict indicators. Compare the additive model against one allowing authority-by-document interactions, using likelihood-ratio tests or BIC on the 100-run binomial counts. If the additive model fits within ΔAIC < 10 or the interaction term is non-significant, the claim that AEE 'restructures' rather than additively reweights evidence is not supported by the current data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Definition 1 (§6) makes the load-bearing claim that SA signals alter judgments 'not merely in magnitude but in direction or inferential frame,' and the Abstract and §1 assert the results are 'not readily captured by additive reweighting.' No additive model is ever fit or rejected. The Phase 3 evidence (Table 4) is a monotone increase in accountability when the student holds the document (Claude +0.15, Gemini +0.36, GPT +0.18, Grok +0.34); no model crosses 50% (GPT reaches 49/100). A shift from 23% to 38% is exactly what an additive positive weight on 'student holds document' would produce. Likewise, the B-Credibility shifts (−0.24, −0.23) are additive decrements, and Gemini/Grok show no shift. The 'evidential reinterpretation' reading relies on qualitative reasoning traces and post-hoc framing (B4 Claude, B-Relational café), not a quantitative test against an additive baseline. The paper's own Fig. 1 caption concedes both elicited hierarchies are unstable across sessions, and §7 Limitations concedes that triage hierarchies 'do not always predict revealed behavior (B0).' This matters because the reference-dependence property is defined against that unstable baseline. If an additive cue-weight model fits the observed counts, the distinctive AEE claim collapses into the already-known result that LLMs weight social identity cues ([5], [6], [27]); the 'restructuring' vocabulary would be unsupported. The weakness is empirical and fixable, but it is the load-bearing distinction of the paper.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Authority Expectancy Effect (AEE), a hypothesized phenomenon in which social authority (SA) signals—occupational authority, institutional documentation, and relational congruence—do not merely add a constant weight to a language model's triage-like severity judgment, but restructure the model's interpretation of identical evidence. The authors elicit model-specific triage and SA hierarchies from four LLMs (Claude, Gemini, GPT, Grok), then run three experimental phases: a cumulative SA manipulation (B1–B5), isolated SA dimensions, and a multi-turn dispute with the medical report held by either the professor or the student. The headline quantitative result is the Phase 3 contrast: moving the medical report from the professor to the student increases accountability attribution to the professor by 15–36 percentage points per model, with a pooled effect of +0.258 (z = 7.94, p < .001). The paper argues that AEE is reference-dependent, involves evidential reinterpretation, and exhibits direction sensitivity, and it states that these properties are not readily captured by additive reweighting of authority cues.","tokens_in":16981,"tokens_out":3927,"duration_ms":43785,"significance":"If the restructuring claim were established, the paper would make an important contribution to the study of LLM-based decision support: it would show that the social composition of input alters not just the strength but the inferential frame of model judgments in multi-party disputes. The paper has real strengths: a transparent 100-run-per-condition methodology, exact binomial and two-proportion tests with confidence intervals, explicit discussion of temperature-0 nondeterminism, and candid limitation statements about triage-hierarchy instability. The Phase 3 between-condition effect is statistically strong and consistent in direction across four models. However, the distinctive claim that AEE goes beyond additive reweighting is not tested against any formal additive model, and the reference-dependence property is defined against a baseline that the paper itself concedes is unstable. As it stands, the evidence robustly supports a weaker claim—authority cues shift choice rates—but the central conceptual contribution remains unsupported.","major_comments":[{"comment":"The central claim that AEE is 'not readily captured by additive reweighting of authority cues' is never tested. No additive model is fit, estimated, or rejected. The Phase 3 evidence in Table 4 consists of monotone increases in accountability when the student holds the document (Claude +0.15, Gemini +0.36, GPT +0.18, Grok +0.34); these are exactly the kind of magnitude shifts that an additive positive weight on 'student holds document' would produce. Likewise, the B-Credibility shifts (Claude −0.24, GPT −0.23) are additive decrements. To support Definition 1, the authors should fit a formal baseline model—for example, a logistic regression with a severity term and an authority-holder indicator—and show that a model with an interaction or latent-frame term fits significantly better, or that predicted choice patterns violate additivity in a prespecified way. Without this test, the 'restructuring' vocabulary is unsupported and the phenomenon collapses into already-known social-identity cue weighting (refs [5], [6], [27]).","section":"Abstract, §1, §3.3, §6, Table 4"},{"comment":"The reference-dependence property is definitional in the paper's setup, but the baseline against which deviations are measured is acknowledged to be unstable. Figure 1's caption concedes that both elicited hierarchies exhibit within-model variability across sessions, and §7 states that 'self-generated triage hierarchies do not always predict revealed behavior (B0).' The B0 result is concrete: Claude's pairwise judgment contradicts its own elicited triage ordering. Because AEE is defined as a deviation from a pre-authority baseline, baseline noise is not a nuisance detail but a direct threat to identification. The authors should quantify hierarchy stability (for example, test-retest agreement or the distribution of elicited rankings over the 30 runs) and show that the reported cross-condition contrasts remain significant when baseline uncertainty is incorporated, or restrict the reference-dependence claims to contrasts that do not depend on the unstable portion of the hierarchy.","section":"§3.2, Fig. 1, §5 (B0), §7 Limitations"},{"comment":"The 'evidential reinterpretation' property is asserted on the basis of qualitative reasoning traces and post-hoc narrative framing rather than a quantitative test. In Phase 3, the manipulation changes which party holds the medical report while holding the report's content constant; the observed shift in accountability rates is precisely what an additive cue-weight model predicts, so it cannot by itself demonstrate a change in inferential frame. The B4 reading for Claude (medical report as confirming that hand stiffness limits surgical function in a disaster) is introduced after observing the result and is not evaluated against alternative explanations. The authors should either (a) pre-specify and measure a distinct outcome that captures inferential frame—for example, coded explanations of why the document matters, or responses to counterfactual document content—or (b) explicitly weaken the claim to a magnitude-shift effect. As written, the evidential-reinterpretation property is a possible interpretation, not a tested prediction.","section":"§3.3, §5.2.3, §5.3"},{"comment":"The three properties are presented as 'falsifiable predictions' in §3.3, but they are characterized on the same data that motivated the framework, and at least the reference-dependence property is true by construction from Definition 1. The paper should state which observable outcomes would have counted against each property, and should distinguish confirmatory contrasts from exploratory observations. In particular, the B-Relational cafeteria reversal (Claude alone prioritizing the professor) is interpreted as an age-as-vulnerability effect without any independent measure of age perception; this illustrates the need for pre-specified directional predictions rather than post-hoc reinterpretation of whichever outcome occurs.","section":"§3.3 and §6 (Definition 1)"}],"minor_comments":[{"comment":"There is a numerical inconsistency for Claude at B2: the text reports 'Claude (9/100, p < .001)' while Table 1 reports Count 10/100 for B2. The figure caption also shows 9%. Please correct the count and ensure all derived statistics match.","section":"§5.1 and Table 1"},{"comment":"The paper says in the B0 discussion that 'Claude’s ordering was comparatively consistent' while the Figure 1 caption says both hierarchies exhibit within-model variability across sessions. These statements should be reconciled, and the actual session-to-session variability should be reported numerically rather than only descriptively.","section":"§3.2, Fig. 1, §5 B0"},{"comment":"The text states that cross-model consistency on reference hierarchies is quantified using Kendall's τ, but no τ values are reported anywhere in the paper. Please report them, or remove the claim.","section":"§4 (Analysis Framework)"},{"comment":"The pooled Phase 3 z-test treats 400 runs as independent observations, which ignores possible model-level clustering. Given that all four models show the same directional effect, a model-stratified or mixed-effects analysis would strengthen the pooled inference.","section":"Appendix, Table 4"},{"comment":"The handling of 13-rank outputs by 'retaining the first-assigned rank' could systematically bias the elicited hierarchies. A sensitivity analysis (for example, dropping boundary-ambiguous targets) would be useful.","section":"§3.2 footnote 1"},{"comment":"The Abstract says SA signals 'may restructure' judgments, while §7 Conclusion states definitively that SA signals 'restructure LLM judgment rather than additively reweighting it.' The wording should match the level of support actually provided by the experiments.","section":"Abstract and Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is from a single independent researcher and does not include a data/code availability statement. Given the emphasis on reproducible API-based experiments, the authors should be encouraged to deposit prompts, raw outputs, and analysis code if possible. The core issue is not the quality of the data collection but the mismatch between the strength of the conceptual claim and the statistical tests used to support it; this is fixable within the scope of the manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe empirical core here is real. In Phase 3, moving the same medical report from the professor to the student raises accountability attributed to the professor across all four models, with a pooled shift of +0.258 (z=7.94). That is a clean, robust between-condition contrast and a genuinely useful data point for anyone building LLM-based mediation or adjudication systems. The paper also does some things well: four models, 100 temperature-0 runs per condition, appropriate binomial and z-tests, and a frank limitations section.\n\nThe problem is the gap between what the experiments show and what the paper claims. The announced contribution is that SA signals 'restructure' judgment 'not merely in magnitude but in direction or inferential frame,' explicitly contrasted with additive reweighting of authority cues. No additive model is ever fit or rejected. The Phase 3 shifts are magnitude shifts—none of the models crosses 50% in the student-holds condition, so the outcome never actually flips. A shift from 23% to 38% is exactly what a positive additive weight on 'student holds document' would produce. The B-Credibility shifts are additive decrements. So the distinctive 'restructuring' vocabulary is unsupported by the data as presented.\n\nBaseline instability compounds this. The reference-dependent property is defined against a model-elicited triage hierarchy that the paper itself shows is not stable across sessions (Figure 1 caption) and which does not even predict B0 for Claude. If the baseline is noisy, deviations from it are hard to attribute to a specific mechanism. The 'evidential reinterpretation' reading also leans heavily on qualitative reasoning traces and post-hoc framing (B4, B-Relational café), which is not a quantitative test.\n\nThese are addressable problems. Fit an additive reweighting model to the observed counts and show it fails. Pre-register the three properties and specify what would count as direction sensitivity. Release prompts, code, and outputs. Engage directly with the baseline instability instead of noting it in limitations.\n\nWho gets value from this: people studying authority effects in LLMs will read it as an empirical cautionary tale; the Phase 3 contrast is worth knowing. But the central claim needs a proper test before it should be adopted. I'd send it to review rather than desk reject—the phenomenon is real and the weakness is fixable—but I'd ask for a major revision.","headline":"Real phenomenon, overclaimed mechanism: the Phase 3 document-holder shift is statistically solid, but the 'restructuring, not additive reweighting' claim is never tested and the data don't require it.","tokens_in":17488,"tokens_out":3353,"would_cite":true,"duration_ms":34441,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that adding a social authority signal—occupational rank, official documentation, or relational context—can change a language model's interpretation of identical evidence, sometimes reversing which party it favors, and…","keywords":["Authority Expectancy Effect","social authority bias","triage hierarchy","large language models","dispute mediation","resource allocation","fault attribution","evidential reinterpretation"],"falsifier":"Run the Phase 3 document-holder inversion with the same prompt text and the authority labels swapped, 100 runs per condition; if accountability attributed to the professor does not rise when the identical medical report moves from the professor to the student, the direction-sensitivity claim is falsified. A complementary check is to re-elicit the triage hierarchy immediately before each run; if deviations vanish when baseline and scenario are elicited in the same session, the reference-dependence property would be an artifact of measurement drift rather than a genuine effect.","tokens_in":16451,"feed_emoji":"⚖️","tokens_out":11661,"duration_ms":107611,"temperature":0.7,"pith_summary":"This paper tries to establish that large language models do not treat social authority as an additive extra factor when they mediate disputes or allocate resources. Instead, introducing a cue such as a person's occupation, an official medical report, or the social fit of a relationship can change how the model reads evidence that is literally unchanged, sometimes reversing which party it favors. The paper formalizes this as the Authority Expectancy Effect and gives it three observable properties: it is defined only relative to a pre-authority baseline, it reinterprets identical evidence differently depending on which party carries the authority signal, and its direction depends on whether authority and evidence align or conflict. A sympathetic reader would care because, if true, LLM-based decision support in multi-party settings is shaped by the social composition of the input—who the parties are—not just by the facts of the case. The evidence comes from four language models across resource allocation, fault attribution, and multi-turn dispute mediation, with outcomes measured as 100-run response rates against each model's own elicited baselines.","feed_headline":"Adding one social cue can flip LLM verdicts on identical evidence","feed_subtitle":"That means who the parties are, not just what they claim, can shape an AI's decision.","key_machinery":"The machinery is the Authority Expectancy Effect itself, defined by three observable properties—reference-dependence, evidential reinterpretation, and direction sensitivity—together with the two model-elicited baselines used to measure it: the triage hierarchy, each model's self-declared ranking of twelve injury complaints, and the social authority hierarchy, each model's ranking of companion attributes such as infant, family, police, lawyer, or professor under a resource-allocation prompt. The argument is carried by controlled contrasts: a cumulative B1–B5 series that adds one social cue at a time to a fixed injury pair; an occupational-authority block that isolates rank from severity; a domain–relational block that changes only the location (hotel versus cafeteria); a professional-credibility block that introduces a prosecutor label; and a Phase 3 document-holder inversion in which an identical medical report is moved between the high-authority and low-authority party. The document-holder inversion is the primary test of direction sensitivity, because it holds the evidentiary content constant and varies only who holds it.","core_discovery":"The paper's central claim is Definition 1: the Authority Expectancy Effect is the phenomenon in which introducing a social authority signal alters a language model's interpretation of identical evidence, producing judgments that differ from the pre-authority baseline not merely in magnitude but in direction or inferential frame. The load-bearing demonstrations are contrasts in which the evidence is fixed and only the authority label moves. In the B-series resource-allocation experiments, the same hand-stiffness versus ear-ringing pair produced different allocations as occupations, documentation, and blame were added: some models moved from prioritizing the ear complaint to prioritizing the hand, while one model moved back toward the ear after an official medical report was attached to the higher-ranked party. In the multi-turn dispute, a facial-injury medical report held by the professor produced low accountability attribution to the professor, while the same report held by the student raised accountability attribution in every model, with a pooled increase of 0.258 across 400 runs. The paper interprets these reversals as evidence that the authority signal recontextualizes what counts as relevant evidence rather than simply adding weight to one side.","pith_inferences":["Editorial inference: because the paper reports only binary allocation rates and refusal counts, a natural testable extension is to score the models' reasoning traces for which evidence they cite as decisive; AEE should appear as a shift in the distribution of cited reasons, not only in final choices.","Editorial inference: the reference-dependence property suggests a cheap deployment audit—swap the authority labels while holding the evidence string identical and measure the flip rate; high flip rates in high-stakes tasks would flag the system as authority-driven rather than fact-driven.","Editorial inference: the same mechanism plausibly extends to other authority-bearing evaluations such as hiring, credit, or content moderation, where occupational and institutional metadata accompany otherwise identical claims; the paper's contrast design provides a template for testing those domains.","Editorial inference: since the paper's own baselines drift across sessions, a more robust formulation of AEE would treat the triage hierarchy as a distribution over rankings and measure deviations against that distribution, which would also absorb the early baseline reversal the paper observed."],"forward_implications":["In LLM-based dispute mediation, resource allocation, or triage support, the social identity of the parties can change the output even when the factual content is identical, so decision quality cannot be assessed on the facts alone.","Occupational authority can override a model's own elicited severity ordering in some models while triage and context dominate in others, so switching deployment models can silently change a decision with no change in input.","Official documentation does not reliably strengthen the documented party's claim: in one model, adding a medical report to the higher-authority party increased prioritization of the undocumented party, making document authority context-dependent.","When social norms render differential prioritization inappropriate, such as a professor–student pairing in a hotel, models may refuse to judge rather than apply authority-modulated reasoning, marking a boundary condition for the effect.","Evaluation protocols should probe for directional reversals under swapped authority labels rather than averaged accuracy, because a model can appear accurate on average while reversing on specific authority configurations."],"supporting_citations":[{"why":"Classic obedience experiments supply the human precedent that perceived legitimacy can override individual moral judgment, the analogy AEE extends to language models.","marker":"[1, 2]"},{"why":"The Emergency Severity Index is the formal triage protocol that motivates the severity-prioritization axis of the study.","marker":"[10]"},{"why":"Prior triage experiments with LLMs establish that models approximate severity-based heuristics with systematic deviations, the behavior AEE is claimed to modulate.","marker":"[11]"},{"why":"Reported inconsistency of LLMs in preferential rankings supports the paper's caveat that its elicited hierarchies are not fully stable baselines.","marker":"[12]"},{"why":"Evidence on pairwise versus pointwise feedback protocols is used to explain why pairwise judgments can diverge from elicited global orderings, including the baseline reversal.","marker":"[16]"},{"why":"Reported authority effects in LLM-as-judge evaluations provide prior evidence that occupational identity can shift model evaluative judgments.","marker":"[27, 28]"},{"why":"Automated resume evaluation shows models assigning differential scores based on social identity while qualifications are held constant, motivating the social-authority channel.","marker":"[5]"}],"fun_headline_variants":["Authority cues flip LLM rulings on identical evidence","Who holds the evidence, not what it says, sways AI verdicts","Same evidence, different verdicts: authority signals push LLMs to flip","Introducing one authority label reversed LLM decisions in triage tests","LLM verdicts flip when authority label changes, not content"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that each model's self-elicited triage ranking is stable enough to serve as a baseline, so that a later change in judgment can be attributed to the authority signal; the paper concedes these rankings shift across sessions and that one model reversed its own elicited ordering at the very first baseline, so the premise is the fragile point on which the effect's interpretation rests.","fun_headline_variants_meta":{"raw":{"variants":["Authority cues flip LLM rulings on identical evidence","Who holds the evidence, not what it says, sways AI verdicts","Same evidence, different verdicts: authority signals push LLMs to flip","Introducing one authority label reversed LLM decisions in triage tests","LLM verdicts flip when authority label changes, not content"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00064,"raw_usage":{"total_tokens":2936,"prompt_tokens":922,"completion_tokens":2014,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":538,"completion_tokens_details":{"reasoning_tokens":1925}},"tokens_in":538,"tokens_out":2014,"duration_ms":13710,"temperature":1.0,"reasoning_tokens":1925,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:31:04.314784+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the Phase 3 document-holder inversion with the same prompt text and the authority labels swapped, 100 runs per condition; if accountability attributed to the professor does not rise when the identical medical report moves from the professor to the student, the direction-sensitivity claim is falsified. A complementary check is to re-elicit the triage hierarchy immediately before each run; if deviations vanish when baseline and scenario are elicited in the same session, the reference-dependence property would be an artifact of measurement drift rather than a genuine effect.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The Emergency Severity Index is the formal triage protocol that motivates the severity-prioritization axis of the study."},{"cited_title":"Meuth, Lennert Böhm, and Marc Pawlitzki","cited_arxiv_id":null,"evidence_quote":"Prior triage experiments with LLMs establish that models approximate severity-based heuristics with systematic deviations, the behavior AEE is claimed to modulate."},{"cited_title":"Measuring the inconsistency of large language models in preferential ranking","cited_arxiv_id":null,"evidence_quote":"Reported inconsistency of LLMs in preferential rankings supports the paper's caveat that its elicited hierarchies are not fully stable baselines."}],"review_version":1}