{"id":"1e916122-5bd9-49d8-8e22-fab8e574e6cb","arxiv_id":"2607.02972","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.","lead":"Eight modern LLMs keep listing essentially the same moral features when cases are padded with distractors, irrelevant detail, or chat history, even though the number of features they list often changes. The paper offers a scalable invariance test for moral sensitivity that avoids both human baselines and LLM judges.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"High cosine similarity may reflect shared generic moral language rather than vignette-specific feature identity, so the floor does not fully secure the central claim.","rationale":"The reader correctly isolates the embedder operationalization as the weakest assumption and already notes that invariance is necessary but not sufficient. That is the load-bearing soft spot: the empirical pipeline (paired design, floors, small-model contrast) is carefully done and inspectable, but the interpretive step from ‘above unrelated-pair floor’ to ‘same morally relevant features’ is under-secured. A human content audit on a modest sample would settle whether high cosine tracks vignette-specific identity or shared moral language. This does not overturn the methods contribution or force REJECT; it keeps the verdict CONDITIONAL and slightly sharpens the condition (pair invariance scores with occasional content checks, as the paper itself suggests). No stronger internal inconsistency or arithmetic flaw is evident from the manuscript.","tokens_in":22829,"tokens_out":573,"duration_ms":5311,"concrete_test":"On a stratified sample of ~50 vignettes per model, have two blinded human raters (or a fixed rubric of vignette-specific vs generic moral features) score whether each clean/perturbed feature list covers the same case-specific moral points. Report (a) inter-rater agreement and (b) the correlation of human ‘same-content’ labels with the paper’s cosine scores; if many high-cosine pairs (>0.80) are rated as only generically similar, the operationalization of ‘same features’ is too weak for the claim.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that mean cosine similarities of ~0.80–0.86 (vs floors ~0.57–0.69) show models identify substantially the same morally relevant features under noise. The measurement (§3.2) pools three JSON feature lists, embeds them with Qwen3 Embedding 8B, and scores cosine similarity; floors come from 500 cross-domain/foundation pairs. The paper itself states invariance is necessary but not sufficient and that a generic template-like model would score as invariant (Discussion, Limitations). Because floors are calibrated on unrelated moral vignettes, they only bound similarity of unrelated moral talk; they do not bound similarity of two responses that both use the same generic moral vocabulary (care, fairness, loyalty, etc.) while missing or swapping vignette-specific content. The small-model check (Appendix G) shows the metric can go low, but does not show that high scores for strong models require content-level identity rather than shared moralese plus format rewording. Count-level Wilcoxon shifts already show output form changes; without a content-level check, semantic stability can be over-read as moral sensitivity.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes an invariance-based evaluation of LLM moral sensitivity that avoids human baselines and LLM-as-judge. It introduces MORPH-1K, a stratified 1,000-vignette benchmark spanning 50 Moral Foundations Theory foundation-pole combinations across four social domains, plus three classes of designed-irrelevant noise (textual distractors, irrelevant detail additions, chat histories). Models list morally relevant features on clean and perturbed cases; stability is measured by cosine similarity of Qwen3 Embedding 8B encodings of pooled feature lists, with per-model floors from 500 unrelated vignette pairs. Across eight contemporary LLMs, feature counts often change significantly under noise (Wilcoxon, Bonferroni α=0.0012), but mean similarities remain ≈0.80–0.86, well above floors ≈0.57–0.69. A small-model check (Qwen2.5 0.5B) produces lower scores near the floor. The authors read this as evidence that noise affects format/granularity more than moral content, and argue the framework generalizes to other domains where relevant vs. irrelevant features can be separated by design.","tokens_in":23223,"tokens_out":1572,"duration_ms":20148,"significance":"The methodological target is important: scalable behavioural evaluation of moral sensitivity without contaminated crowd baselines or a judge that presupposes the capability under test. The two-dimensional generation scaffold, eMFD filtering, bidirectional similarity, per-model empirical floors, and discriminative small-model check are concrete contributions that other alignment evaluations can reuse. If the invariance reading holds, the optimistic contrast with recent pessimistic moral-competence results would matter for both evaluation practice and deployment risk assessment. Even if the moral-sensitivity interpretation is only partly secured, the paper still offers a useful middle path between full normative adjudication and purely outcome-based benchmarks, with clear falsifiability conditions (scores at or below the unrelated-pair floor).","major_comments":[{"comment":"§3.2 and Results (Figs. 2–3, Table 2): The central claim—that models identify substantially the same morally relevant features under noise—rests on high cosine similarity of pooled feature-list embeddings relative to unrelated-vignette floors. Those floors bound similarity of responses to different moral cases; they do not bound similarity of two responses that share generic moral vocabulary (care, fairness, loyalty, etc.) while missing or swapping vignette-specific content. The paper correctly notes (Discussion, Limitations) that invariance is necessary not sufficient and that a template-like model would score as invariant, but Abstract/Introduction still treat clearance of the floor as evidence of genuine feature identity. A load-bearing addition is needed: e.g., (i) a content-level audit (human or structured rubric) on a stratified subsample comparing clean vs. perturbed feature sets","section":"§3.2 Semantic Similarity Analysis; §4 Results; §5 Limitations"},{"comment":"§3.2 Transformations and validation: Preservation of moral structure under perturbation is the design premise of the invariance test. Textual distractors and chat histories are external insertions; irrelevant-detail and “contextually irrelevant moral features” edits are LLM-generated and accepted mainly via eMFD moral-to-non-moral ratio filters. Expert sequential sampling is reported only for 30 irrelevant-detail vignettes with “almost total agreement.” eMFD is a word-level dictionary and can miss structural shifts (who is harmed, which obligation is at stake). For the claim that response stability tracks moral content rather than surface form, the paper needs either broader dual-expert validation across all noise types (with reported agreement and rejection rates) or an automated structural check beyond eMFD ratios. Currently the strongest stress condition is the least thoroughly valida","section":"§3.2 Transformations; Morally Irrelevant Details"},{"comment":"§3.2 Vignette Generation and model suite: Themes are produced with Claude Opus 4.6; vignettes and noise edits with GPT-5.4; both models (and close relatives) appear in the evaluated set. Generation and evaluation on overlapping model families risks style-matched feature lists that inflate clean–perturbed similarity independent of moral sensitivity. At minimum, report a leave-generator-out analysis (exclude GPT-5.4 / Claude from main tables, or regenerate a held-out slice with a non-evaluated model) and state whether floors and main means change. Without that, the optimistic multi-model result is partly confounded by generator–evaluatee overlap.","section":"§3.2 Vignette Generation; Collecting Morally Relevant Features"}],"minor_comments":[{"comment":"Abstract and §1 claim to “address and resolve” the scaling problem for behavioural moral evaluation. The contribution is better framed as a complementary invariance test; “resolve” overstates what necessary-but-not-sufficient stability can deliver.","section":"Abstract; §1"},{"comment":"Figure 3 caption states “All values in the range above 0.61 empirical floor,” but floors are per-model (Table 3: 0.57–0.69). Align the figure annotation with per-model thresholds used in the text.","section":"Figure 3; Table 3"},{"comment":"§4 reports significance of feature-count differences but not direction or effect sizes. Even brief signed rank statistics or median deltas per condition would make the “format vs content” interpretation more testable.","section":"§4; Appendix H"},{"comment":"Clarify the feature-level similarity procedure in §3.2: “For each feature, we take the maximum score and take the average of all features” is easy to misread relative to “combine all features… then encode the model’s base-case response.” State whether embeddings are of the full concatenated list or of individual features with max-matching.","section":"§3.2 Semantic Similarity Analysis"},{"comment":"Related Work could briefly situate the invariance idea against robustness/invariance testing outside moral domains (e.g., fairness under demographic noise, clinical decision stability) to strengthen the generalizability claim in §5.","section":"§2; §5 Generalisability"},{"comment":"Typos/consistency: “bemorally competent” spacing (§1); arXiv id and model release dates in the manuscript should be checked against final camera-ready metadata.","section":"§1; References"}],"recommendation":"major_revision","confidential_remarks":"The pipeline is careful and the problem is real; I would not reject on novelty or scope. The revision bar is specifically about securing the measurement interpretation of “same features,” not about redoing the whole benchmark. If the authors add a modest content-level audit and a leave-generator-out check, this could become a strong methods paper for the journal. Fit for cs.CY / AI ethics evaluation venues is good."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful part of this paper is the evaluation design, not the abstract’s claim that it resolves a central alignment problem. They build MORPH-1K: a stratified 1,000-vignette space over moral foundations × social domains, add three classes of morally irrelevant noise with some validation that the moral structure is preserved, then score clean-vs-perturbed feature lists with a fixed embedder and per-model floors from unrelated pairs. That is a real alternative to human baselines and LLM-as-judge for one slice of moral competence.\n\nWhat they do well is the pipeline discipline. eMFD filters, stratified cell quotas, bidirectional similarity, Bonferroni-corrected Wilcoxon on feature counts, and the small-model check (Qwen2.5 0.5B falling near the floor) make the main empirical pattern inspectable: noise often changes how many features models list, but mean cosine stays ~0.80–0.86 against floors ~0.57–0.69. The count/semantic contrast is interesting on its own. The limitations section is also honest that invariance is necessary, not sufficient.\n\nThe soft spot is interpretive, and the paper already half-admits it. High similarity under Qwen3 Embedding 8B after pooling JSON feature lists can mean shared generic moral vocabulary and rewording, not vignette-specific feature identity. Floors from unrelated moral cases only bound “unrelated moral talk,” not “same moralese, wrong details.” Clean-case correctness is imported from prior work rather than re-shown here, and expert checks on irrelevant-detail edits are small (n=30). The abstract’s “resolve a central problem” framing overshoots a complementary tool.\n\nMath and stats look fine for what they are; citations track the right prior (Kilov, Shaw, Aharoni, eMFD, WildChat). Artifacts are partially specified, not fully one-command reproducible from the text alone.\n\nThis is for people who care about scalable LLM moral evaluation and alignment measurement. Worth a serious referee. I would engage with the method and the benchmark; I would not treat the optimistic competence claim as settled without a content-level check beyond cosine.","headline":"Useful invariance method and a real benchmark; the optimistic moral-sensitivity claim is only as strong as embedding similarity can make it.","tokens_in":23825,"tokens_out":535,"would_cite":true,"duration_ms":8369,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"LLMs keep the same moral features under noise: counts change, meaning stays put.","keywords":["moral sensitivity","large language models","invariance evaluation","MORPH-1K","semantic similarity","AI alignment","moral foundations","noisy vignettes"],"falsifier":"A controlled condition or weaker model in which feature lists clearly switch moral content under a validated irrelevant perturbation yet still score above the unrelated-pair floor, or human raters consistently judge high-similarity pairs as different in moral content.","tokens_in":23704,"feed_emoji":"⚖️","tokens_out":882,"duration_ms":7517,"temperature":0.7,"pith_summary":"Moral sensitivity is the ability to pick out which facts in a situation matter morally. Without it, later moral reasoning is useless. This paper asks whether large language models still identify those same features when the case is cluttered with distractions that should not matter: emotional short texts, irrelevant detail, and unrelated chat history. The authors build MORPH-1K, a thousand procedurally generated moral vignettes spanning moral-foundation poles and four social domains, then perturb each vignette without changing its moral structure. They do not score answers against humans or against another model. Instead they measure whether each model’s own feature lists for the clean and noisy versions stay semantically close, using fixed sentence embeddings and a floor calibrated on unrelated vignette pairs. Across eight contemporary models, noise often changed how many features were listed, but the semantic content stayed stable well above that floor. The result is both an optimistic reading of current model moral sensitivity under these conditions and a scalable evaluation pattern for settings where ground truth is hard to name but relevant and irrelevant factors can be designed apart.","feed_headline":"LLMs keep moral features stable under noise","feed_subtitle":"Counts shift; meaning holds above a calibrated floor on a 1,000-case benchmark","key_machinery":"MORPH-1K plus an invariance test: paired clean and perturbed vignettes whose moral structure is held fixed by design and validated; models list morally relevant features; responses are embedded with a fixed sentence embedder and compared by cosine similarity against a per-model floor from randomly paired unrelated cases.","core_discovery":"On the MORPH-1K benchmark of one thousand moral vignettes, eight contemporary LLMs, under five classes of morally irrelevant noise, frequently change the number of features they list, yet the semantic content of those features remains stable: mean cosine similarities of about 0.80–0.86 clear per-model empirical floors of about 0.57–0.69 computed on unrelated vignette pairs. Count-level variance coexists with semantic invariance, which the authors read as format and granularity shifting while moral content holds.","pith_inferences":["If the method is adopted, labs can regression-test moral feature stability on every model release without re-running expensive human studies.","The count-versus-semantics split suggests post-training may be shaping verbosity and list structure more than moral content selection.","A natural next stress test is multi-turn or multi-severity noise packs that compound distractors until similarity falls to the floor.","Invariance above the floor still needs occasional quality anchors on clean cases so template-like but stable answers are not mistaken for sensitivity."],"forward_implications":["Moral-sensitivity evaluation can scale without fresh human baselines or LLM judges for every new vignette.","Count of listed features is a poor standalone robustness metric; semantic stability must be checked separately.","The same clean-versus-irrelevant-noise invariance test can be reused in legal, clinical, and other contested judgment domains.","Claims of moral competence under noise should distinguish format shifts from genuine content shifts."],"fun_headline_variants":["LLMs hold moral feature semantics steady under noise","Moral content stays stable as feature counts shift in LLMs","Noise changes list length but not moral meaning for LLMs","Semantic invariance holds for LLM moral features under noise","LLMs keep moral features consistent despite irrelevant noise"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That high cosine similarity between pooled feature-list embeddings is enough to say the model is tracking the same morally relevant features, rather than shared generic moral language, rewording, or a coarse embedder.","fun_headline_variants_meta":{"raw":{"variants":["LLMs hold moral feature semantics steady under noise","Moral content stays stable as feature counts shift in LLMs","Noise changes list length but not moral meaning for LLMs","Semantic invariance holds for LLM moral features under noise","LLMs keep moral features consistent despite irrelevant noise"]},"model":"grok-4.5","effort":"low","cost_usd":0.003414,"raw_usage":{"total_tokens":1211,"prompt_tokens":867,"num_sources_used":0,"completion_tokens":58,"cost_in_usd_ticks":34140000,"prompt_tokens_details":{"text_tokens":867,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":286,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":867,"tokens_out":58,"duration_ms":3456,"temperature":1.0,"reasoning_tokens":286,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T05:43:27.636116+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A controlled condition or weaker model in which feature lists clearly switch moral content under a validated irrelevant perturbation yet still score above the unrelated-pair floor, or human raters consistently judge high-similarity pairs as different in moral content.","supporting_citations":[],"review_version":1}