{"id":"66b5f493-dc7a-4b5d-9b90-4981d4e3047a","arxiv_id":"2605.11632","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Macro uses DPO on composite preference pairs to raise validity of multilingual self-generated counterfactual explanations by 12.55% on average over chain-of-thought while preserving minimality.","lead":"Macro applies Direct Preference Optimization with a composite scoring function to generate self-generated counterfactual explanations that are more valid in seven languages without losing minimality. A smart generalist might read it to see how preference alignment can make LLM explanations more usable outside English.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Composite scoring function may not reliably produce unbiased preference pairs for DPO without hidden artifacts","rationale":"The reader's weakest assumption directly identifies the same load-bearing point. Because the full manuscript is referenced but the scoring-function details remain the critical unverified step, the UNVERDICTED status is retained; the concrete test above would resolve whether the assumption holds.","tokens_in":1715,"tokens_out":356,"duration_ms":16846,"concrete_test":"From the methods section, extract the exact composite score formula (including any weights, thresholds, or language-specific adjustments). Re-run the preference-pair construction on a held-out subset of the seven languages using (a) the reported formula and (b) an alternative monotonic combination (e.g., validity first, then lexicographic minimality). Train two DPO models and compare validity/minimality deltas on the test set; if the alternative pairing yields comparable or better results, the original composite score is not uniquely responsible for the claimed gains.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the composite scoring function (validity + minimality) produces preference pairs whose ordering is faithful to the intended trade-off and free of new biases or language-specific artifacts. If the function's weighting, normalization, or validity proxy (e.g., prediction flip detection across languages) is misspecified, DPO will optimize a distorted objective; the reported 12.55% validity gain and preserved minimality could then be artifacts of the scoring rule rather than genuine alignment improvement. This assumption is the least secure because the abstract provides no equation or validation for the composite score, and the experimental gains are measured against the same scoring function used to create the training signal.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces Macro, a preference alignment framework applying Direct Preference Optimization (DPO) to multilingual self-generated counterfactual explanations (SCEs). It uses a composite scoring function to build preference pairs from the validity-minimality trade-off and reports that this yields a 12.55% average validity gain over chain-of-thought baselines across four LLMs and seven typologically diverse languages, without degrading minimality and while outperforming translation-based and supervised fine-tuning approaches.","tokens_in":1837,"tokens_out":463,"duration_ms":17560,"significance":"If the results hold after full methodological disclosure and validation of the scoring function, the work would be significant for multilingual XAI: it provides evidence that explicit preference optimization can resolve the validity-minimality tension in counterfactual generation where supervised methods fall short, and it demonstrates cross-lingual perturbation alignment improvements.","major_comments":[{"comment":"Abstract: the composite scoring function that produces the DPO preference pairs is described only qualitatively; no equation, weighting scheme, normalization procedure, or validity proxy (e.g., cross-lingual prediction-flip detection) is supplied. Because the reported 12.55% validity gain and preserved minimality rest directly on the ordering induced by this function, its misspecification would render the gains artifacts rather than genuine alignment improvements.","section":"Abstract"},{"comment":"Abstract / Experiments: the 12.55% validity improvement is stated without statistical details (standard errors, confidence intervals, number of runs, or significance tests) or per-language/per-model breakdowns; this prevents verification that the gain is robust rather than driven by a subset of languages or models.","section":"Abstract"}],"minor_comments":[{"comment":"Abstract: the phrase 'further analyses reveal that Macro increases cross-lingual perturbation alignment' is asserted without naming the metrics or showing the supporting figures/tables.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The provided abstract and reader's notes indicate insufficient methodological detail for independent verification of the central experimental claim; this aligns with the low soundness rating and suggests the manuscript may require substantial expansion of the methods section before the results can be assessed."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments on our manuscript. We address each major comment below and will make the requested changes to improve methodological transparency and statistical reporting.","responses":[{"response":"We agree that the composite scoring function requires a formal mathematical description to ensure reproducibility and to confirm that the preference ordering is well-specified. In the revised manuscript we will add the explicit equation for the composite score, the weighting scheme between validity and minimality terms, the normalization procedure, and the precise validity proxy (cross-lingual prediction-flip detection) in the Methods section.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the composite scoring function that produces the DPO preference pairs is described only qualitatively; no equation, weighting scheme, normalization procedure, or validity proxy (e.g., cross-lingual prediction-flip detection) is supplied. Because the reported 12.55% validity gain and preserved minimality rest directly on the ordering induced by this function, its misspecification would render the gains artifacts rather than genuine alignment improvements."},{"response":"We agree that statistical details and breakdowns are essential for verifying robustness. In the revised Experiments section we will report standard errors, confidence intervals, the number of runs, significance tests, and full per-language and per-model tables so that readers can confirm the consistency of the 12.55% average gain.","revision_made":"yes","referee_comment":"[Abstract] Abstract / Experiments: the 12.55% validity improvement is stated without statistical details (standard errors, confidence intervals, number of runs, or significance tests) or per-language/per-model breakdowns; this prevents verification that the gain is robust rather than driven by a subset of languages or models."}],"tokens_in":1354,"tokens_out":383,"duration_ms":20692,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main move is to treat the validity-minimality trade-off as a preference signal and run DPO on it for self-generated counterfactuals across languages.\n\nThis is a straightforward extension of existing alignment work to a setting that has mostly been handled by translation or English-only methods.\n\nThe experiments cover four LLMs and seven typologically varied languages, and the reported 12.55% validity gain over chain-of-thought while holding minimality steady is the clearest result. The comparison to both a translation baseline and supervised fine-tuning is reasonable and shows why preference optimization is claimed to matter here.\n\nThe composite scoring function that creates the preference pairs is the actual novelty. If the weighting and normalization are sound, the approach gives a practical lever for balancing the two goals.\n\nThe soft spot is exactly that function. The abstract gives no equation or validation step, so it is still possible the gains partly reflect how the score was constructed rather than a clean alignment improvement. The stress-test concern about hidden artifacts in the preference pairs is the one that needs the methods section to settle.\n\nMention of post-hoc analyses on cross-lingual alignment and error types is noted but not shown, which leaves those claims thinner.\n\nThis is for people working on multilingual model explanations. A reader already in that area will get the experimental comparisons and the direction toward preference optimization.\n\nThe setup has enough concrete runs and baseline controls to deserve a serious referee, though the scoring details will be the main point to press.","headline":"Macro uses DPO plus a composite score to lift multilingual counterfactual validity, but the score itself is the part that still needs checking.","tokens_in":2317,"tokens_out":379,"would_cite":false,"duration_ms":20749,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Macro applies direct preference optimization to multilingual self-generated counterfactual explanations, raising validity by 12.55 percent on average over chain-of-thought baselines while preserving minimality.","keywords":["counterfactual explanations","multilingual LLMs","direct preference optimization","validity minimality trade-off","self-generated explanations","model interpretability","alignment methods"],"falsifier":"Running the same four models on an eighth typologically distant language and finding that validity gains are accompanied by statistically larger minimality violations or new error patterns would falsify the central claim.","tokens_in":2622,"feed_emoji":"🌐","tokens_out":576,"duration_ms":13574,"temperature":0.7,"pith_summary":"The paper presents Macro, a framework that turns the validity-minimality trade-off in counterfactual explanation generation into explicit preference pairs for direct preference optimization. It tests this approach on four large language models across seven typologically diverse languages and reports consistent gains in validity without the minimality losses seen in translation baselines. The work matters because valid and minimal explanations are needed to interpret model decisions in languages other than English, where current methods either fail to flip predictions or produce overly altered inputs. By showing that preference optimization outperforms both chain-of-thought and supervised fine-tuning on the combined metrics, the authors argue that explicit alignment is required to balance the two objectives in multilingual settings.","feed_headline":"Macro raises multilingual explanation validity by 12.55%","feed_subtitle":"Preference optimization balances validity and minimality for counterfactuals across seven languages and four models.","key_machinery":"A composite scoring function that converts the validity-minimality trade-off into measurable preference signals for direct preference optimization.","core_discovery":"Macro constructs preference pairs for direct preference optimization from a composite scoring function that rewards validity (prediction flip) and minimality (small input change) in self-generated counterfactual explanations. When applied to multilingual generation, this alignment step raises average validity by 12.55 percent relative to chain-of-thought prompting, keeps minimality intact, and outperforms supervised fine-tuning on both metrics. The same method also increases cross-lingual perturbation alignment and reduces common generation errors.","pith_inferences":["The preference-pair construction could be reused for other explanation formats that face similar validity-minimality tensions.","The gains observed across typologically diverse languages suggest the method may transfer to additional low-resource languages not tested here.","If the composite scoring function generalizes, similar alignment pipelines might improve other multilingual generation tasks that require controlled edits."],"forward_implications":["Macro produces higher cross-lingual perturbation alignment than the tested baselines.","It reduces common generation errors that appear in chain-of-thought and translation-based outputs.","It outperforms supervised fine-tuning on both validity and minimality simultaneously.","It avoids the severe minimality violations observed in the translation-based baseline."],"fun_headline_variants":["Macro: 12.55% multilingual SCE validity gain via preference optimization","Macro attains 12.55% higher validity for multilingual counterfactuals","12.55% SCE validity increase across seven languages via Macro","Macro shows 12.55% multilingual SCE validity gain via DPO alignment"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The composite scoring function produces reliable preference signals that direct preference optimization can follow without creating new biases or artifacts in the generated explanations.","fun_headline_variants_meta":{"raw":{"variants":["Macro: 12.55% multilingual SCE validity gain via preference optimization","Macro attains 12.55% higher validity for multilingual counterfactuals","12.55% SCE validity increase across seven languages via Macro","Macro shows 12.55% multilingual SCE validity gain via DPO alignment"]},"model":"grok-4.3","cost_usd":0.014926,"raw_usage":{"total_tokens":6412,"prompt_tokens":668,"num_sources_used":0,"completion_tokens":75,"cost_in_usd_ticks":149262000,"prompt_tokens_details":{"text_tokens":668,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":5669,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":668,"tokens_out":75,"duration_ms":39899,"temperature":1.0,"reasoning_tokens":5669,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T22:41:44.442294+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the same four models on an eighth typologically distant language and finding that validity gains are accompanied by statistically larger minimality violations or new error patterns would falsify the central claim.","supporting_citations":[],"review_version":2}