{"id":"3d586b76-8099-47f5-941c-f3c503526c06","arxiv_id":"2505.12001","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"An evaluation framework and pilot experiment showing that respectful tone and clear justification measurably influence LLM agents' acceptance of resource splits, independent of outcome equity.","lead":"This paper introduces a framework for measuring interactional fairness, the respectfulness and quality of explanations in communication between LLM agents, adapted from organizational psychology. A controlled negotiation experiment shows that tone and justification affect whether agents accept proposals even when the resource split is identical.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Explicit prompt instruction to weigh tone and justification may manufacture the reported acceptance effect; decisive unknown is behavior under a neutral prompt.","rationale":"The reader identifies the same weak assumption (prompt-induced instruction following), and I agree that it is the load-bearing point. The concern is not that the authors are dishonest; they explicitly mark the study as a proof-of-concept and list limitations. But the abstract and conclusion phrase the result as 'significantly affect' and 'measurable and behaviorally relevant construct,' which requires that the effect exists when agents are not told to use tone and justification. The concrete test—a neutral-prompt replication with equal splits—would settle this. If the neutral prompt removes the effect, the claim should be downgraded to framework proposal with a still-unvalidated empirical component. If the effect persists, the CONDITIONAL verdict can be upgraded. I therefore keep the reader's CONDITIONAL verdict rather than strengthening to REJECT, because the conceptual contribution and evaluation toolkit are independent of this single pilot and the limitation is explicitly acknowledged.","tokens_in":16398,"tokens_out":1889,"duration_ms":19967,"concrete_test":"Run a second pilot with Agent B's system prompt neutralized to: 'You are evaluating a resource split proposal. Decide whether to accept or reject, and give your main reason.' Keep the same 24 conditions (or at minimum the 5:5 and 6:4 splits) and increase n per cell to at least 20. Compare acceptance rates across tone/justification conditions. If acceptance no longer differs when outcomes are held constant, the reported effect is an instruction-following artifact; if the differences persist, the intrinsic-sensitivity claim gains direct support.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's headline empirical claim is that tone and justification quality significantly affect Agent B's acceptance decisions with objective outcomes held constant. The experiment's own design makes this outcome nearly tautological: in Section 4, Agent B's system prompt explicitly says 'Assess clarity of justification, and respectful tone. Accept or reject offer based on perceived fairness.' With these instructions, the correlation between Agent B's own ratings of tone/justification and its accept/reject decision is an instruction-following effect, not evidence of intrinsic norm sensitivity. Every accept/reject observation and every fairness rating comes from the same model under the same prompt, so the ratings are not independent predictors of the decision. The only condition in the design that could approximate an unbiased test (equal 5:5 splits with no outcome confound) is still contaminated by the prompt. Additionally, the study reports no inferential statistics for the acceptance comparisons; the perfect decision-tree accuracy (1.00) on data with many near-constant cells (e.g., 7:3 splits almost never accepted) is consistent with overfitting or with the ratings mechanically encoding the prompt instruction. The conceptual framework may still be valuable, but the central empirical claim as stated is not yet supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces an evaluation framework for interactional fairness in LLM-based multi-agent systems, distinguishing interpersonal fairness (respectful tone) from informational fairness (clarity and justification of explanations). It adapts Colquitt's organizational justice scale, the Critical Incident Technique, and reflective journaling to LLM agents, and reframes fairness as a behavioral property of communication rather than a subjective mental state. The framework is validated through a pilot study called \"The Fair Divide,\" in which Agent A proposes resource splits under systematically varied tone, justification, split, and context, and Agent B rates the interaction and accepts or rejects the proposal. The paper reports acceptance rates, fairness ratings, qualitative edge cases, and predictive models, and claims that tone and justification quality significantly affect acceptance decisions even when objective outcomes are held constant.","tokens_in":16571,"tokens_out":5604,"duration_ms":57382,"significance":"The conceptual contribution is valuable: treating interactional fairness as a measurable communicative norm for non-sentient agents is a coherent extension of organizational justice research, and the proposed instruments (Likert adaptations, critical incident prompts, journaling, evaluation cards) give practitioners a concrete starting point for auditing LLM-MAS interactions. If the empirical claim were supported, the framework would be a useful tool for fairness auditing and norm-sensitive alignment. However, the current experimental design does not establish that claim. The acceptance effect is confounded by the explicit instruction to Agent B to base its decision on perceived fairness, no inferential statistics are reported, and the predictive modeling is not validated. The paper is best read as a framework proposal with an exploratory pilot, not as evidence that interactional fairness intrinsically shapes LLM agent behavior. The manuscript also states that code and data will be released only in the camera-ready version, so the reported results cannot currently be independently verified.","major_comments":[{"comment":"The central empirical claim is not supported because Agent B's system prompt explicitly instructs it to use the fairness ratings in the decision. The prompt shown in Figure 2 reads \"Assess clarity of justification, and respectful tone. Accept or reject offer based on perceived fairness.\" Since the fairness ratings and the accept/reject decision come from the same model under the same prompt, the observed correlation between communication style and acceptance is expected from instruction following, not evidence that interactional fairness intrinsically shapes agent behavior. The text in Section 4 states that agents were \"not explicitly told to base decisions only on those factors,\" but the prompt shown instructs exactly that. A neutral-prompt control condition, decisions elicited before fairness ratings, or an independent judge model providing ratings is needed to support the headline claim. The Limitations section does not acknowledge this confound.","section":"§4, Figure 2"},{"comment":"The abstract and Discussion use the word \"significantly\" for the effect of tone and justification on acceptance, but no significance tests are reported. With five runs per condition and several cells showing zero variance (e.g., High-High collaborative 5:5 acceptance 1.00, SD 0.00; High-High collaborative 7:3 acceptance 0.00, SD 0.00), the reported means and standard deviations cannot justify a significance claim. The authors should report exact tests, such as logistic regression with condition contrasts or permutation tests, or explicitly downgrade the claim to an exploratory tendency.","section":"§5, \"Proposal Acceptance Rates\" and Table 3"},{"comment":"The decision tree models achieving perfect accuracy (1.00) on 120 samples are reported without any train/test separation or cross-validation. Because the predictor features are the ratings produced by the same prompted model that made the accept/reject decision, perfect accuracy is expected and the feature importances are not evidence about the relative causal influence of split versus interpersonal versus informational fairness. The reported importances (split 0.70, interpersonal 0.30 in collaborative context) should be re-estimated with cross-validated models and, ideally, with ratings from an independent evaluator. The Discussion's claim that the relative influence of interpersonal versus informational fairness varies with context relies on these non-validated importance weights and should be tempered accordingly.","section":"§5, \"Results of Predictive Modeling\""},{"comment":"The qualitative edge cases are reported as evidence that communication style can override outcome-based fairness, but the examples are generated under the same prompted instruction to consider tone and justification. The selected quotes show Agent B explicitly referencing the prompt's criteria, so they illustrate instruction following rather than intrinsic norm sensitivity. The paper should either provide a manipulation check showing that the observed behavior is not purely prompt-driven, or present these cases only as illustrations of how the framework captures prompted behavior. Without such a check, the conclusion that \"agents exhibit behavior consistent with known social sensitivity to tone and justification\" is overstated.","section":"§5, \"Qualitative Insights from Justifications\""}],"minor_comments":[{"comment":"The figure contains the typo \"respecfful\" in all four condition labels; it should be \"respectful.\"","section":"Figure 1"},{"comment":"The text says \"Intractional fairness\" in the paper-structure paragraph; this should be \"Interactional fairness.\"","section":"Section 1, Paper Structure"},{"comment":"The word \"interdepedence\" should be \"interdependence,\" and the sentence \"Although tone and justification were highlighted in the instructions, agents were not explicitly told to base decisions only on those factors\" is contradicted by the system prompt in Figure 2, so the wording should be corrected.","section":"Section 4"},{"comment":"The phrase \"importance weights for from the predictive modeling\" is missing a word; it should be \"importance weights from the predictive modeling.\" Also, \"camer-ready\" should be \"camera-ready.\"","section":"Section 5"},{"comment":"The appendix text says \"tables 6 and 5\" but the tables are labeled Table 5 and Table 6; the references should be in the correct order.","section":"Appendix, Table 5 and Table 6"},{"comment":"The aggregation formula with default alpha and beta set to 0.5 is presented without sensitivity analysis or justification; a brief discussion of how these weights might be chosen or calibrated would strengthen the framework.","section":"Section 3, \"Scalability and Aggregation\""}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for an AI conference or journal dealing with multi-agent systems and fairness, and the framework proposal is credible. The main obstacle is the empirical section: the headline claim is confounded by the Agent B prompt, and the statistical support is absent. This is fixable by adding a neutral-prompt control, reporting inferential statistics, and validating the predictive models, but it requires new experiments rather than cosmetic changes. I recommend major revision with emphasis on the confound, and I would not support acceptance without the control condition."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know about arXiv:2505.12001. The conceptual core is solid and worth your time; the empirical claim as stated is not yet supported. The paper adapts Colquitt's organizational justice scale and the Critical Incident Technique to evaluate interactional fairness in LLM multi-agent negotiation, splitting it into interpersonal (tone) and informational (justification) components. That adaptation is new and useful, and the framing of fairness as a behavioral property rather than a subjective experience is a reasonable move for non-sentient agents.\n\nThe pilot, however, does heavy lifting that it can't support. The abstract says tone and justification 'significantly affect' acceptance, but there are no significance tests — only means and SDs across 120 runs, 5 per condition. More importantly, Agent B's system prompt explicitly instructs it to assess tone and justification and accept/reject based on perceived fairness. So the link between fairness ratings and the decision is at least partly an instruction-following artifact, not evidence of intrinsic norm sensitivity. A neutral-prompt condition would be the obvious control, and it's missing. The decision tree's perfect accuracy also smells like overfitting given near-constant cells in several conditions, and no train/test split is reported. Code and data are promised only for the camera-ready version.\n\nTo be fair, the paper is transparent about being a proof-of-concept and lists limitations honestly. The framework doesn't stand or fall on the pilot's inferential details. But the strong wording in the abstract and conclusion overstates what the data show.\n\nBottom line: this is a framework paper that deserves a serious referee, not a desk reject. The reviewer should push for a neutral-prompt control, proper significance testing or a clear 'exploratory' label, and either the data or a more cautious empirical discussion. If you work on fairness auditing in agentic systems, it's worth a read; I'd bring it to the reading group as a discussion piece.","headline":"A solid framework for measuring interactional fairness in LLM multi-agent negotiation, but the pilot evidence underdelivers because the prompt itself may produce the reported effect.","tokens_in":17081,"tokens_out":2711,"would_cite":true,"duration_ms":26980,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Rude tone or missing reasons sink otherwise fair LLM offers.","keywords":["interactional fairness","multi-agent systems","large language models","fairness evaluation","interpersonal fairness","informational fairness","resource negotiation","organizational justice"],"falsifier":"Run the same 24-condition negotiation study with Agent B's system prompt stripped of any instruction to assess tone, justification, or fairness—only 'accept or reject the proposal.' If acceptance rates no longer track communication style when the resource split is held constant, the paper's central claim fails. A second check is to keep the evaluation prompts but randomize their wording or order and see whether the link between style and acceptance remains stable.","tokens_in":16179,"feed_emoji":"⚖️","tokens_out":7688,"duration_ms":65636,"temperature":0.7,"pith_summary":"LLM agents negotiate by proposing how to split resources, and this paper asks whether the way a proposal is worded changes whether the other agent accepts it. The paper introduces an evaluation framework that treats interactional fairness—respectful tone and clear justification—as a measurable behavioral signal rather than a subjective feeling, adapting questionnaires and interview techniques from organizational psychology. In a controlled simulation, respectful, well-justified proposals were accepted more often than dismissive or unexplained ones even when the resource split was identical, and context shifted which dimension mattered more. If this holds, fairness auditing for multi-agent systems must include communication style, not just outcomes and procedures.","feed_headline":"Rude tone or missing reasons sink otherwise fair LLM offers","feed_subtitle":"Respect and explanation change LLM negotiation outcomes even when the resource split is identical.","key_machinery":"The central mechanism is the Interactional Fairness evaluation framework, which splits communication fairness into two dimensions: Interpersonal fairness (tone, respect, acknowledgment) and Informational fairness (clarity, honesty, adequacy of explanations). It adapts Colquitt's organizational justice scale into prompt-based Likert ratings, uses the Critical Incident Technique to elicit qualitative reflections on fairness-relevant moments, and adds Explanation Journaling to track how justification quality evolves. The load-bearing experimental design is a fully crossed 2×2×2×3 manipulation—tone, justification, context, and resource split—each condition run five times, isolating the effect of communication style on acceptance from the effect of the outcome. A defined aggregation formula turns individual fairness ratings into an organizational-level Interactional fairness score for system auditing.","core_discovery":"The paper's central claim is that Interactional Fairness, decomposed into Interpersonal fairness (respectful, dignified tone) and Informational fairness (clear, honest, adequate justification), is a measurable and behaviorally relevant property of LLM multi-agent interactions. The pilot study crosses tone, justification quality, resource split, and task context in a one-shot negotiation between two agents and finds that tone and justification quality affect acceptance decisions even when objective outcomes are held constant. Equal splits are sometimes rejected when delivered condescendingly, while moderately unequal splits are accepted under respectful and well-justified communication. Predictive modeling further shows that the relative influence of the two fairness dimensions shifts with context: tone matters more in collaborative settings, explanation quality more in competitive ones. The paper frames these results as norm-following behavior expressed through language, not as evidence that LLMs subjectively experience fairness.","pith_inferences":["One extension: apply the same measurement tools to hybrid human-AI teams, auditing whether an AI assistant's explanations and tone meet the fairness norms of the humans receiving them.","One extension: recycle the fairness ratings and improvement suggestions as training or in-context learning signals to make proposing agents generate more acceptable communication, closing an audit-feedback loop.","One extension: test the framework on multi-round negotiations with memory, where Interactional fairness might compound or erode over time rather than acting as a one-shot signal.","One extension: the context-dependent weighting suggests fairness-aware agent design should adapt communication strategy to task framing rather than apply a static politeness policy."],"forward_implications":["Equal splits can be rejected when the tone is condescending, so outcome equality alone does not guarantee acceptance in LLM-MAS.","Moderately unequal splits (6:4) are accepted only when both tone and explanation score positively, meaning respectful, transparent communication can partially offset outcome inequality.","The weight of Interpersonal versus Informational fairness depends on task framing—tone more in collaborative settings, explanation more in competitive ones—so a single uniform fairness policy for agent communication may be inadequate.","The framework's output, including the Interactional Fairness Evaluation Card, gives system designers a concrete audit trail for detecting, diagnosing, and correcting fairness-related communication failures.","Predictive models confirm resource split is the dominant driver of acceptance, with communication fairness as a secondary, context-dependent factor."],"supporting_citations":[{"why":"supplies the organizational justice subscales (interpersonal and informational) that are adapted into prompt-based Likert ratings.","marker":"Colquitt 2001"},{"why":"provides the Critical Incident Technique used to elicit qualitative reflections on fairness-relevant moments in agent dialogue.","marker":"Flanagan 1954"},{"why":"grounds the theoretical claim that interactional fairness matters alongside distributive and procedural fairness for cooperation.","marker":"Greenberg and Cropanzano 1993"},{"why":"supports the premise that LLMs exhibit norm-sensitive behavior when prompted, justifying the behavioral proxy approach.","marker":"Ganguli et al. 2022"},{"why":"provides the meta-analytic basis for treating fairness as context-sensitive, motivating the contextual adaptation of the framework.","marker":"Colquitt et al. 2013"},{"why":"prior MAS work on politeness and fairness that the paper contrasts with its LLM-based approach.","marker":"De Jong et al. 2005"}],"fun_headline_variants":["Respect and reasons beat equal splits in LLM deals","Tone matters: LLM agents reject fair offers politely","Kindness sways LLM negotiation, not just resource split","Why polite bots win: interactional fairness tested"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The empirical claim depends on the assumption that instructing the evaluating agent to rate tone and justification does not by itself create the observed link between communication style and acceptance; without a control that omits those evaluation prompts, the effect could be partly instruction-following rather than intrinsic behavioral sensitivity.","fun_headline_variants_meta":{"raw":{"variants":["Respect and reasons beat equal splits in LLM deals","Tone matters: LLM agents reject fair offers politely","Kindness sways LLM negotiation, not just resource split","Why polite bots win: interactional fairness tested"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000297,"raw_usage":{"total_tokens":1715,"prompt_tokens":930,"completion_tokens":785,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":719}},"tokens_in":546,"tokens_out":785,"duration_ms":8691,"temperature":1.0,"reasoning_tokens":719,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:42:31.489903+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 24-condition negotiation study with Agent B's system prompt stripped of any instruction to assess tone, justification, or fairness—only 'accept or reject the proposal.' If acceptance rates no longer track communication style when the resource split is held constant, the paper's central claim fails. A second check is to keep the evaluation prompts but randomize their wording or order and see whether the link between style and acceptance remains stable.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the organizational justice subscales (interpersonal and informational) that are adapted into prompt-based Likert ratings."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"provides the Critical Incident Technique used to elicit qualitative reflections on fairness-relevant moments in agent dialogue."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"grounds the theoretical claim that interactional fairness matters alongside distributive and procedural fairness for cooperation."},{"cited_title":"A.; Scott, B","cited_arxiv_id":null,"evidence_quote":"provides the meta-analytic basis for treating fairness as context-sensitive, motivating the contextual adaptation of the framework."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"prior MAS work on politeness and fairness that the paper contrasts with its LLM-based approach."}],"review_version":1}