{"id":"2f330a07-6740-4b07-af3b-2b5e924934ec","arxiv_id":"2607.11736","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"MET-D self-distills theory-selected moral grounds into native-language reasoning, lifting macro-F1 by ~3.7–4.2 points on MCLASH and MMoralExceptQA while raising native-language chains by ~62 points.","lead":"The paper introduces MCLASH, a culturally adapted multilingual moral-dilemma benchmark, plus MET/MET-D: theory-grounded two-step prompting and self-distillation that raise moral-decision F1 and native-language reasoning across several open models. It matters because deployed LLMs already face high-stakes moral queries from non-English users, and English-centric scaffolds systematically misalign with local norms.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The self-distillation proxy may reward label-matching rather than genuine theory-guided reconciliation of grounds.","rationale":"The Reader correctly isolates the weakest link: the assumption that matching the automatically derived ternary label is a reliable proxy for substantive engagement with theoretical grounds. The paper supplies independent evidence that this link is soft (Figure 2, Figure 18, Section 6.2, Limitations). Gains remain real and multi-model, cultural analyses are informative, and native-language legibility improves, so the package still supports CONDITIONAL rather than REJECT. No stronger load-bearing flaw (e.g., data leakage or broken cultural adaptation) is evident from the text; the proxy issue is the one that most directly undercuts the strongest claim about theory-grounded reasoning. The proposed annotation check would settle whether the concern lands without requiring new training runs.","tokens_in":28779,"tokens_out":494,"duration_ms":4982,"concrete_test":"On a stratified sample of 100 MCLASH instances (balanced languages and yes/no/ambiguous), have two independent annotators score each MET-D reasoning chain for (a) number of selected grounds substantively applied and (b) presence of explicit reconciliation among conflicting grounds (binary). Correlate both scores with correctness; if reconciliation score is near zero or uncorrelated with F1 while utilization alone predicts correctness, the proxy does not support the theory-grounding claim.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that MET-D improves theory-grounded moral reasoning rests on rejection sampling that keeps only reasoning chains whose final ternary answer matches the synthetic character-derived label (Section 4.2). The paper itself shows that even after training, models often list more grounds without reconciling them (Figure 18; Section 6.2 utilization rises from ~52% to ~61% in English but reconciliation is absent). If correct labels can be reached by shallow pattern-matching to character value cues while merely name-dropping grounds, the measured F1 gains (3.71 on MCLASH, 4.23 on MMoralExceptQA) and the 62-point native-language increase do not establish improved theory-guided reasoning; they establish better label prediction under a theory-flavored prompt. The untrained ground-selection step and the synthetic training distribution further weaken the causal link between the proxy and the claimed mechanism.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper addresses multilingual moral decision-making with three contributions: (1) MCLASH, a culturally adapted (not merely translated) benchmark of 1,852 long-form scenarios across six languages derived from CLASH; (2) MET, a two-step prompting method that first selects from 37 expert-curated theoretical grounds organized into six dimensions, then produces theory-guided reasoning in the user’s native language; and (3) MET-D, which adds rejection-sampling self-distillation on synthetic character-paired dilemmas whose ternary labels are automatically determined by constructed value priorities, requiring no external supervision. On Qwen3-4B/8B and Gemma3-4B, MET-D raises macro-F1 by ~3.71 points on MCLASH and ~4.23 on MMoralExceptQA (peak +12.94 Malay on Qwen3-8B), increases native-language reasoning by ~62 points, and yields analyses of culture-specific beneficial grounds and typology-driven cross-lingual transfer.","tokens_in":29048,"tokens_out":1043,"duration_ms":8214,"significance":"If the results hold, the work supplies a reusable culturally adapted evaluation resource, an extensible theory-ground inventory, and a supervision-light training recipe that improves both accuracy and legibility of multilingual moral reasoning. Strengths include multi-model/multi-dataset evaluation (MCLASH, MMoralExceptQA, MoCa), ablations separating selection from reasoning (Table 2), random-selection baselines (Table 3), utilization annotations, and explicit cultural analyses (Figures 4–6). The self-distillation design that avoids stronger-model or human labels is practically valuable. These contributions are of clear interest to multilingual NLP and AI ethics, even if absolute gains remain modest.","major_comments":[{"comment":"§4.2 and the axiom that matching the synthetic ternary label is a proxy for “substantive engagement with the grounds” is load-bearing for the claim of improved theory-guided reasoning. The paper’s own evidence undercuts this: §6.2 reports utilization rising only from ~52% to ~61% (English) and explicitly notes that engagement ≠ reconciliation; Figure 18 shows the trained model listing more grounds without weighing them. Without a direct measure of reconciliation quality (or human preference over reasoning chains), the F1 gains (Table 1) and native-language gains establish better label prediction under a theory-flavored prompt more securely than improved theory-guided moral reasoning. A targeted human evaluation of reconciliation, or an ablation that removes ground names while keeping character cues, would strengthen the causal claim.","section":null},{"comment":"Training and primary evaluation both rely on the same character–value structure (CLASH-style synthetic data for MET-D; MCLASH for main results). Although MMoralExceptQA and MoCa provide external checks, the largest and most culture-specific claims rest on MCLASH. The paper should quantify how much of the MCLASH gain is explained by distributional similarity to the synthetic training set versus genuine transfer of theory-guided reasoning, e.g., by reporting performance stratified by topic/value overlap or by training only on non-overlapping topics.","section":null}],"minor_comments":[{"comment":"Table 1 and Appendix C: report statistical significance or confidence intervals for the modest average gains (~3–4 F1); several per-language deltas are small enough that variance could matter.","section":null},{"comment":"§6.1: Cohen’s κ of 0.259 across models for ground selection is low; clarify whether this still supports “systematic, non-random” selection or mainly within-model stability (κ=0.468).","section":null},{"comment":"Figure 18 and §6.2: the distinction between engagement and reconciliation is important; consider elevating a short quantitative reconciliation metric (even on the 20-instance sample) into the main text.","section":null},{"comment":"Limitations correctly flags untrained ground selection and missing reconciliation; these should be cross-referenced earlier when interpreting MET-D gains so readers do not over-read the mechanism.","section":null},{"comment":"Minor presentation: ensure consistent spelling of “Probabilistic” vs “Probablistic” in grounds lists; expand the brief MoCa drop explanation for Gemma3-4B (clear-cut bias) with a short confusion-matrix note.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The central methodological risk (label-matching proxy vs. genuine theory use) is real but already partially acknowledged by the authors; it is fixable with additional analysis rather than a redesign. The culturally adapted benchmark and expert-curated grounds are solid contributions independent of the training story. Fit for COLM is good; I would not block on the proxy issue if the authors add the requested checks and temper the mechanism language."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful pieces here are MCLASH (long-form dilemmas culturally adapted rather than translated into five languages, with native-speaker checks) and a clean two-step recipe that first selects from an expert-curated inventory of 37 theory grounds, then reasons in the user’s language. MET-D adds rejection-sampling self-distillation on synthetic character-paired data so no external labels are needed. Gains are consistent across Qwen3-4B/8B and Gemma3-4B: roughly +3.7 F1 on MCLASH, +4.2 on MMoralExceptQA, a 12.9-point Malay peak, and a large jump in native-language reasoning chains. Ablations, random-selection baselines, utilization annotations, and cross-lingual transfer matrices are all present; the paper also shows that beneficial grounds track real cultural patterns (e.g., Contractarianism for Korean/Chinese, Relation-Based Authority for Spanish).\n\nWhat is actually new is the combination: cultural adaptation of long scenarios, a multilingual theory inventory co-authored with a moral philosopher, and a free supervision signal that exploits the character-value structure of CLASH-style data. Prior theory scaffolds were English-only and static; prior multilingual moral sets were mostly short or direct translations.\n\nThe soft spot the stress-test flags is real but not fatal. Training keeps only chains whose final ternary label matches the synthetic character-derived answer. The paper itself shows utilization rises only modestly and reconciliation of conflicting grounds often fails (Figure 18). So the F1 and native-language gains are better label prediction under a theory-flavored prompt, not proof of deep theory-guided moral reasoning. Ground selection remains untrained, and adaptation quality is checklist-based rather than quantified. These are honest limitations the authors already flag; they do not erase the external-dataset generalization or the cultural analyses.\n\nMath and citation pattern look ordinary and clean; code and data are promised. This is for people working on multilingual alignment, cultural NLP, or moral decision-making who need a usable benchmark and a practical prompting-plus-distillation recipe. It deserves a serious referee. I would bring it to reading group and would cite the benchmark and the ground inventory.","headline":"Solid multilingual moral-reasoning package: culturally adapted benchmark, expert grounds, and self-distillation that lifts F1 and native-language use; the proxy for “theory engagement” is weaker than claimed but the empirical package still holds.","tokens_in":29633,"tokens_out":561,"would_cite":true,"duration_ms":5341,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Theory-grounded two-step prompting plus self-distillation lifts multilingual moral decisions without human labels.","keywords":["multilingual moral reasoning","theory-grounded prompting","cultural adaptation","self-distillation","moral foundations","native-language reasoning","MCLASH"],"falsifier":"Measure whether models that reach the correct ternary label after MET-D training actually reconcile conflicting grounds (for example by human annotation of reconciliation quality or by controlled ablation of grounds) rather than merely pattern-matching the final answer; if correct labels rise while reconciliation stays flat, the proxy fails.","tokens_in":29684,"feed_emoji":"⚖️","tokens_out":960,"duration_ms":7734,"temperature":0.7,"pith_summary":"Language models are being asked to make moral judgments for people whose languages and cultures differ from the English-centric data that most models and benchmarks were built on. Direct translation of dilemmas erases local institutions, customs, and value systems; existing reasoning scaffolds stay static and English-only; and training for moral alignment usually needs expensive human or stronger-model supervision. This paper answers all three gaps at once. It releases MCLASH, a culturally adapted multilingual benchmark of high-stakes scenarios. It supplies MET, a two-step method that first selects situation- and culture-specific theoretical grounds drawn from psychology and philosophy, then reasons over them in the user’s native language. It further adds MET-D, a self-distillation stage that uses automatically verifiable labels from synthetic character-paired dilemmas so no external supervision is required. Across three model families the trained method raises decision quality, sharply increases native-language reasoning, and shows that the most useful grounds track real cultural patterns rather than a universal English template.","feed_headline":"Theory-guided self-distillation lifts multilingual moral F1","feed_subtitle":"No human labels needed; native-language reasoning jumps 62 points and cultural grounds differ by language","key_machinery":"MET / MET-D: the model first selects a situation- and language-specific subset of 37 expert-curated theoretical grounds (organized into six dimensions: Normative Authority, Value Systems, Ethical Theories, Cognitive Reasoning Strategies, Conflict Handling Strategies, Moral Uncertainty), then generates a theory-guided reasoning chain; MET-D further trains the second step by retaining only those chains whose final ternary answer matches the automatically derived ground-truth label of a synthetic CLASH-style character.","core_discovery":"A two-step procedure—selecting expert-curated moral-theory grounds that fit the situation and culture, then reasoning over them in the native language—combined with rejection-sampling self-distillation on synthetic character-paired dilemmas, consistently improves multilingual moral decision accuracy and native-language legibility without any external human or stronger-model labels.","pith_inferences":["If the label-proxy assumption holds only for clear-cut cases, the same pipeline may still under-train genuine multi-framework reconciliation needed for high-stakes ambiguous dilemmas.","The same select-then-reason-plus-self-distill recipe could be ported to other value-laden domains (law, medicine, content moderation) that currently rely on English-centric rubrics.","Ground-selection itself remains untrained; a second distillation stage that rewards culturally coherent ground sets could close the remaining gap the authors leave open."],"forward_implications":["Moral-reasoning systems can be made culture-aware without translating English dilemmas or importing Western scaffolds wholesale.","Self-distillation on character-paired synthetic data can replace expensive human or stronger-model supervision for ternary moral decisions.","Beneficial moral-theory grounds differ systematically by language and track real cultural patterns (e.g., contractarianism for Confucian languages, relation-based authority for Spanish).","Cross-lingual transfer of moral reasoning follows linguistic typology more than cultural similarity, so same-script or same-word-order languages help each other more than culturally close but typologically distant ones.","Enforcing native-language reasoning becomes performance-positive rather than a trade-off once the model has been trained with MET-D."],"fun_headline_variants":["Theory-grounded self-distillation lifts multilingual moral F1","Culture-aware moral grounds raise LLM decision accuracy","Native-language reasoning jumps 62 points via MET-D","Self-distilled theory scaffolds improve moral F1 without labels","MET selects cultural grounds then reasons natively for moral gains"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"Matching the automatically derived yes/no/ambiguous label on synthetic character-paired dilemmas is treated as a reliable signal that the model has substantively engaged with and reconciled the selected theoretical grounds.","fun_headline_variants_meta":{"raw":{"variants":["Theory-grounded self-distillation lifts multilingual moral F1","Culture-aware moral grounds raise LLM decision accuracy","Native-language reasoning jumps 62 points via MET-D","Self-distilled theory scaffolds improve moral F1 without labels","MET selects cultural grounds then reasons natively for moral gains"]},"model":"grok-4.5","effort":"low","cost_usd":0.005638,"raw_usage":{"total_tokens":1589,"prompt_tokens":878,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":56380000,"prompt_tokens_details":{"text_tokens":878,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":626,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":878,"tokens_out":85,"duration_ms":5080,"temperature":1.0,"reasoning_tokens":626,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T03:31:09.700137+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Measure whether models that reach the correct ternary label after MET-D training actually reconcile conflicting grounds (for example by human annotation of reconciliation quality or by controlled ablation of grounds) rather than merely pattern-matching the final answer; if correct labels rise while reconciliation stays flat, the proxy fails.","supporting_citations":[],"review_version":1}