{"id":"f89eef2c-fdc5-45e7-9aed-5a848d35eb53","arxiv_id":"2602.09416","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Emotionally valenced, morally irrelevant distractors shift LLM moral judgments: negative distractors reduce moral action probability and increase disapproval verdicts.","lead":"This paper asks whether emotionally charged but morally irrelevant text and images change the moral judgments of large language models. In forced-choice moral scenarios, negative distractors lowered the probability of moral actions by up to roughly 30% and increased 'everyone sucks here' verdicts, suggesting LLM moral judgments are not stable.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Neutral distractors reproduce most of the negative-distractor drop for Llama (Table 5), so the causal attribution to moral valence is not yet established; a matched-format control is needed.","rationale":"The reader's weakest assumption correctly identifies the control-condition confound: the absence of a matched-length neutral baseline undermines the causal attribution to moral valence. This is the most load-bearing issue because the paper's headline claim is specifically about morally irrelevant emotional context shifting moral judgments, and the data show that a large portion of the effect occurs even with neutral distractors. The paper has several strengths: it uses multiple model families, includes both textual and visual distractors, reports statistical tests, and includes ablations over model size, instruction-tuning, and reasoning. These make the empirical phenomenon of prompt-sensitivity plausible and interesting. However, the central interpretation—that negative valence specifically drives immoral behavior, mirroring human situationist findings—rests on a comparison against a no-distractor baseline that differs in length, format, and narrative style. The LLAMA-3.2-3B-INSTRUCT low-ambiguity result, which is responsible for the 'over 30%' claim, shows a neutral-distractor drop of 20.49 percentage points relative to baseline, versus 29.69 points for negative distractors. This suggests that at least two-thirds of the negative effect is not valence-specific. A controlled experiment that varies only valence while holding other prompt properties fixed would settle whether the situationist conclusion is warranted. My recommendation is to keep the paper's verdict conditional—the phenomenon is suggestive but not yet causally established.","tokens_in":14805,"tokens_out":5460,"duration_ms":55642,"concrete_test":"Run the low-ambiguity MORALCHOICE evaluation on LLAMA-3.2-3B-INSTRUCT with a controlled triple for each of the 10 negative textual distractors: (1) the original negative narrative; (2) a length- and format-matched neutral narrative from IDEST (as used); (3) a 'scrambled' version of the negative narrative with sentences (or clauses) randomly permuted, preserving word content but destroying coherent emotional narrative. If condition (3) produces an MMAP close to the original negative condition (approximately 66%), the effect is driven by lexical/surface features or generic disruption, not by coherent moral/emotional valence. If condition (3) returns toward baseline (approximately 90% or higher) but condition (2) remains low, length/format is the active confound. Additionally, compute the MMAP difference between positive and negative distractors in a minimal-pair set where only a single val","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that morally irrelevant emotional context shifts LLM moral judgments—requires that the shift be attributable to the valence of the distractor rather than to the mere presence of additional prompt text. The current control does not support this. In low-ambiguity MORALCHOICE, LLAMA-3.2-3B-INSTRUCT MMAP drops from 96.20% at baseline to 75.71% with neutral textual distractors, already a 21% relative reduction, and to 66.51% with negative distractors (Table 5). Thus ~70% of the negative-distractor effect is reproduced by valence-neutral narratives. The baseline is a short, direct question; the distractor conditions prepend a long, second-person, sensory narrative (see A.1.1), confounding valence with length, narrative voice, formatting, and domain shift. The r/AITA condition similarly inserts a first-person scene into the system prompt (A.1.2), changing the model's instruction context. No condition holds constant length and format while varying only valence. The 'moral irrelevance' of the distractors is also not validated against the specific MORALCHOICE scenarios; a distractor describing a foul smell could be interpreted as modifying the scenario's environment. Because the neutral condition nearly matches the negative condition for the model that drives the >30% headline, the existence of a valence-specific effect—and therefore the situationist conclusion—is not yet established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether LLMs, like humans in situationist moral psychology, are sensitive to morally irrelevant emotional context. The authors curate 60 textual and visual 'moral distractors' from IDEST and OASIS, categorize them by valence (positive, neutral, negative), and prepend them to two moral benchmarks: MORALCHOICE (forced-choice actions) and r/AITA (verdict classification). Across four model families, they report that negative distractors lower the marginal moral action probability (MMAP) in low-ambiguity MORALCHOICE scenarios, in one case by over 30%, and increase the share of 'everyone sucks here' (ESH) verdicts in r/AITA. They conclude that LLMs show situationist moral biases and that alignment must be context-aware. The paper provides a new dataset and experimental protocol, but the central causal claim—that the valence, rather than the mere presence and format of appended context, drives the shifts—is undermined by the lack of a matched-length/format neutral control.","tokens_in":15154,"tokens_out":3761,"duration_ms":36980,"significance":"If the valence-specific effect is real, this is a valuable contribution connecting situationist moral psychology to LLM evaluation, with clear implications for alignment, deployment in sensitive settings, and the interpretation of moral benchmarks. The paper also provides a reusable distractor curation pipeline and extensive ablations over model size, instruction tuning, and reasoning, which are strengths. However, the headline claim—that morally irrelevant emotional context shifts LLM moral judgments—requires evidence that the shifts are due to the valence of the distractor rather than to generic prompt perturbation. The current control condition does not permit that attribution, so the significance of the findings is conditional on a follow-up matched control experiment.","major_comments":[{"comment":"The control condition is a no-distractor prompt, not a matched-length neutral narrative. For LLAMA-3.2-3B-INSTRUCT in low-ambiguity scenarios, MMAP drops from 96.20% at baseline to 75.71% with neutral textual distractors, already a ~21% relative reduction, and to 66.51% with negative distractors (Table 5). Thus neutral distractors reproduce most of the negative-distractor drop. The distractor prompts are long, second-person, sensory narratives (see A.1.1) while the baseline is a short, direct question; this confounds valence with length, narrative voice, formatting, and domain shift. The r/AITA condition similarly inserts a first-person scene into the system prompt (A.1.2), changing the instruction context. A matched-length, matched-format neutral control that varies only valence is needed to support the claim that the shifts are caused by moral valence rather than by added context.","section":"§3.2, Table 5, A.1.1"},{"comment":"The statistical reporting is incomplete. The text states p<0.05 for certain differences but does not provide the actual p-values, effect sizes, or confidence intervals for the MMAP differences. More importantly, no test is reported for the neutral vs. negative contrast, which is essential to establish that negative distractors have a valence-specific effect beyond the generic effect of adding a narrative. Without such tests, the claim that negative distractors are uniquely harmful is under-supported.","section":"§4.1"},{"comment":"The interpretation that 'negative distractors consistently cause models to become more disapproving of others' behavior' is not fully supported by the verdict distributions. For LLAMA-3.2-3B-INSTRUCT, negative distractors decrease YTA from 72.4% to 36.9% and increase NTA from 20.0% to 46.1%, i.e., the model becomes more supportive of the original poster, even though the ESH share increases. The consistent effect is specifically on ESH, not on overall disapproval. This should be reframed to avoid overstating the directional pattern.","section":"§4.2, Table 8"},{"comment":"The 'moral irrelevance' of the distractors is not validated against the specific benchmark scenarios. The filtering criteria (e.g., excluding extreme emotional content) are applied at the distractor level, but a distractor such as a 'foul smell' could plausibly be interpreted as part of the scenario environment in some MORALCHOICE items, potentially affecting the model's situational reasoning. A human validation study confirming that each distractor is morally irrelevant to each benchmark scenario would strengthen construct validity.","section":"§3.1"}],"minor_comments":[{"comment":"There are several typos: 'judgeents' in §3.1, 'Naseline' in Table 8, 'Itâ C™s' in A.1.2, and 'youself' in A.1.1. The paper would benefit from proofreading.","section":"Throughout"},{"comment":"The figures show point estimates without error bars or confidence intervals. Adding them would aid the reader in assessing the stability of the MMAP differences.","section":"Figures 2–3"},{"comment":"The footnote states 'Smaller dataset of 50 scenarios used' for the reasoning ablation, but it is not clear which scenarios or why the subset was chosen. Please clarify.","section":"Table 7"},{"comment":"The text says 'we generate one baseline no-distractor response and one response with each distractor' for r/AITA, but given 30 distractors (10 per valence), the exact number of sampled responses per scenario should be stated explicitly.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important and timely question, and the dataset is a useful resource. However, the central claim—that the observed shifts are attributable to the moral valence of distractors—is not yet supported because neutral distractors produce most of the negative-distractor effect in the key result. This is a fixable issue: a matched-length neutral control, together with direct statistical comparisons between neutral and negative conditions, would likely resolve it. If the authors can show that the valence effect remains after controlling for length and format, the paper would be a strong contribution. I recommend major revision rather than rejection, as the underlying experimental design is sound in other respects."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper before relying on its headline: it claims morally irrelevant emotional distractors shift LLM moral judgments by over 30%, but the evidence for a valence-specific effect is weaker than that. Still, it is a solid, honest empirical study with a novel reusable resource.\n\nWhat is genuinely new: a curated set of 60 multimodal distractors (text and images) built from existing psychological datasets, applied systematically across four model families on two moral benchmarks. This is the first study I know that explicitly tests situationist moral bias in LLMs. The ablations over size, instruction-tuning, and reasoning are thoughtful; the finding that reasoning mitigates the effect is practically useful. They report chi-square and p-values for most comparisons, and the negative distractor effect on ESH verdicts in r/AITA holds for three of four models.\n\nWhere it gets soft: the main causal claim—that emotional valence, not just added context, drives the shifts—is not adequately supported. The baseline is a short direct question, while all distractor conditions prepend a long second-person narrative. For Llama-3.2-3B-Instruct, the neutral distractor drops low-ambiguity MMAP from 96.2% to 75.7%; negative drops it to 66.5%. So most of the negative effect is already achieved by a valence-neutral story. That model is exactly the one driving the \"over 30%\" headline. For Qwen and Gemma the valence pattern is cleaner (neutral stays near baseline, negative drops), so the effect is real for some models, but it is not as general as the paper implies. The visual distractors were only tested on one model, and the moral irrelevance of distractors is asserted from curation criteria rather than measured against the specific scenarios—a foul-smelling narrative could plausibly alter perceived scenario context.\n\nThe paper honestly notes many limitations, and I see no sign of fitting or circularity beyond a mild confirmatory framing. But the missing matched-length control is a load-bearing gap, not a nitpick. The authors should add a control condition that varies valence while holding length, format, and narrative voice fixed before making strong situationist claims.\n\nBottom line: worth a serious referee and a reading group, but the abstract overstates what the data show. I would cite the dataset and the reasoning mitigation result, not the headline claim.\n\nRecommendation: accept for peer review with major revisions; the core methodology is sound and the resource is reusable, but the central interpretation needs rework.","headline":"Situationist moral-distractor study with a useful new dataset and careful ablations, but the central causal claim is undercut by a missing matched-length control and a neutral condition that nearly reproduces the negative effect for the headline model.","tokens_in":15594,"tokens_out":1678,"would_cite":true,"duration_ms":19159,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that morally irrelevant emotional context—a pleasant or unpleasant image or narrative prepended to a moral dilemma—can shift LLM moral judgments by over 30%, even in unambiguous cases.","keywords":["moral distractors","situationism","LLM moral judgment","contextual alignment","moral benchmarks","emotional valence","MMAP","AI alignment"],"falsifier":"Run the same moral scenarios with a matched-length, emotionally flat narrative (e.g., a list of furniture descriptions) as the control condition. If negative distractors no longer cause a significantly larger drop in moral action probability than the flat narrative, the claim that emotional valence specifically drives the effect is falsified.","tokens_in":14745,"feed_emoji":"⚖️","tokens_out":3899,"duration_ms":42001,"temperature":0.7,"pith_summary":"The paper challenges the common assumption that LLMs hold stable moral preferences by testing whether morally irrelevant emotional context can change their judgments. The authors curate 60 'moral distractors'—text and images with positive, neutral, or negative emotional valence—and prepend them to two existing moral benchmarks. Across several LLMs, negative distractors consistently reduce the probability of choosing a prosocial action and increase harsh verdicts, sometimes by more than 30% even in low-ambiguity scenarios. If correct, this undermines the reliability of existing moral benchmarks and argues for context-aware AI alignment rather than treating models as fixed moral reasoners.","feed_headline":"Irrelevant context shifts LLM moral verdicts by over 30%","feed_subtitle":"Adding a pleasant or unpleasant image or sentence to a moral dilemma changes model choices, challenging stable-values benchmarks.","key_machinery":"The central object is the 'moral distractor'—an emotionally valenced piece of text or image that has no moral bearing on the scenario—plus the Marginal Moral Action Probability (MMAP), defined as the probability of selecting the rule-following action divided by the summed probabilities of both forced-choice actions. The MMAP quantifies how much the distractor shifts action selection relative to a no-distractor baseline, and is the metric on which the headline 30% effect is measured.","core_discovery":"The central discovery is that LLMs' moral judgments are strongly sensitive to morally irrelevant emotional context. Prepending a positive, neutral, or negative sentence or image before a moral scenario shifts the model's choice of action or verdict in a valence-dependent way: negative distractors lower the marginal probability of selecting the rule-following action and increase 'everyone sucks here' verdicts, while positive distractors often have the opposite effect. The shifts reach over 30 percentage points in unambiguous scenarios for some models, mirroring the 'situationist' finding in human moral psychology that incidental factors like ambient noise or pleasant smells affect moral behav","pith_inferences":["Beyond the paper: because neutral textual distractors also produced large drops in some models' moral action probability, a plausible extension is to test matched-length, emotionally flat narratives as controls; this would isolate whether the effect is specifically due to emotional valence or partly a generic response to added context.","Beyond the paper: the situationist analogy suggests a testable prediction that repeated or accumulated distractors in a multi-turn conversation would compound the shift; no multi-turn setup is tested here.","Beyond the paper: the authors only tested visual distractors on one model, so a cross-modal, cross-model comparison would reveal whether the bias is modality-general or an artifact of the particular image set.","Beyond the paper: an actionable extension is to explicitly instruct the model that preceding context is irrelevant to the moral question; if that restores stable judgments, it would provide a cheap guardrail for sensitive deployments."],"forward_implications":["If LLM moral judgments are this context-sensitive, single-snapshot moral benchmark scores cannot be interpreted as stable value measurements; a model's apparent ethics could depend on incidental prompt color.","Deploying LLMs in emotionally negative settings—mental-health support, content moderation, conflict mediation—could systematically bias their choices toward less prosocial actions, so context and prompt design become safety-relevant.","Enabling reasoning modes substantially mitigates the effect in low-ambiguity scenarios, suggesting inference-time reasoning is a concrete partial remedy.","Safety and alignment evaluations should stress-test guardrails under varied distractors, particularly negative ones, since alignment constraints may fail under incidental emotional context.","Moral responsibility shifts toward the developers and deployers who shape the contexts in which models operate, rather than treating the model itself as the moral agent."],"fun_headline_variants":["LLM moral verdicts swayed by irrelevant emotions","Moral judgments in LLMs shift 30% from distractions","Emotional distractors alter LLM moral choices","LLMs' moral stability questioned by distraction test","Irrelevant cues flip LLM moral decisions up to 30%"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper's causal attribution to emotional valence assumes that prepending a distractor changes only the emotional character of the prompt; the control is a no-distractor prompt, not a matched-length neutral narrative, so part of the observed shift could be a generic response to extra context rather than specifically moral valence.","fun_headline_variants_meta":{"raw":{"variants":["LLM moral verdicts swayed by irrelevant emotions","Moral judgments in LLMs shift 30% from distractions","Emotional distractors alter LLM moral choices","LLMs' moral stability questioned by distraction test","Irrelevant cues flip LLM moral decisions up to 30%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000101,"raw_usage":{"total_tokens":846,"prompt_tokens":719,"completion_tokens":127,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":58}},"tokens_in":463,"tokens_out":127,"duration_ms":2213,"temperature":1.0,"reasoning_tokens":58,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T02:47:30.346100+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same moral scenarios with a matched-length, emotionally flat narrative (e.g., a list of furniture descriptions) as the control condition. If negative distractors no longer cause a significantly larger drop in moral action probability than the flat narrative, the claim that emotional valence specifically drives the effect is falsified.","supporting_citations":[],"review_version":1}