REVIEW 2 major objections 1 minor 4 references
Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization
T0 review · 2 major / 1 minor · reviewed 2026-06-30 · grok-4.3
Pith's one-line read Macro applies direct preference optimization to multilingual self-generated counterfactual explanations, raising validity by 12.55 percent on average over chain-of-thought baselines while preserving minimality.
desk verdict Macro uses DPO plus a composite score to lift multilingual counterfactual validity, but the score itself is the part that still needs checking. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
A composite scoring function that converts the validity-minimality trade-off into measurable preference signals for direct preference optimization.
What would settle it
Running the same four models on an eighth typologically distant language and finding that validity gains are accompanied by statistically larger minimality violations or new error patterns would falsify the central claim.
Extended reading notes
Core claim
Macro constructs preference pairs for direct preference optimization from a composite scoring function that rewards validity (prediction flip) and minimality (small input change) in self-generated counterfactual explanations. When applied to multilingual generation, this alignment step raises average validity by 12.55 percent relative to chain-of-thought prompting, keeps minimality intact, and outperforms supervised fine-tuning on both metrics. The same method also increases cross-lingual perturbation alignment and reduces common generation errors.
Load-bearing premise
The composite scoring function produces reliable preference signals that direct preference optimization can follow without creating new biases or artifacts in the generated explanations.
Editorial extensions
If this is right
- Macro produces higher cross-lingual perturbation alignment than the tested baselines.
- It reduces common generation errors that appear in chain-of-thought and translation-based outputs.
- It outperforms supervised fine-tuning on both validity and minimality simultaneously.
- It avoids the severe minimality violations observed in the translation-based baseline.
Reading between the lines
- The preference-pair construction could be reused for other explanation formats that face similar validity-minimality tensions.
- The gains observed across typologically diverse languages suggest the method may transfer to additional low-resource languages not tested here.
- If the composite scoring function generalizes, similar alignment pipelines might improve other multilingual generation tasks that require controlled edits.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Macro, a preference alignment framework applying Direct Preference Optimization (DPO) to multilingual self-generated counterfactual explanations (SCEs). It uses a composite scoring function to build preference pairs from the validity-minimality trade-off and reports that this yields a 12.55% average validity gain over chain-of-thought baselines across four LLMs and seven typologically diverse languages, without degrading minimality and while outperforming translation-based and supervised fine-tuning approaches.
Significance. If the results hold after full methodological disclosure and validation of the scoring function, the work would be significant for multilingual XAI: it provides evidence that explicit preference optimization can resolve the validity-minimality tension in counterfactual generation where supervised methods fall short, and it demonstrates cross-lingual perturbation alignment improvements.
major comments (2)
- [Abstract] Abstract: the composite scoring function that produces the DPO preference pairs is described only qualitatively; no equation, weighting scheme, normalization procedure, or validity proxy (e.g., cross-lingual prediction-flip detection) is supplied. Because the reported 12.55% validity gain and preserved minimality rest directly on the ordering induced by this function, its misspecification would render the gains artifacts rather than genuine alignment improvements.
- [Abstract] Abstract / Experiments: the 12.55% validity improvement is stated without statistical details (standard errors, confidence intervals, number of runs, or significance tests) or per-language/per-model breakdowns; this prevents verification that the gain is robust rather than driven by a subset of languages or models.
minor comments (1)
- [Abstract] Abstract: the phrase 'further analyses reveal that Macro increases cross-lingual perturbation alignment' is asserted without naming the metrics or showing the supporting figures/tables.
Simulated Author's Rebuttal
We thank the referee for the constructive comments on our manuscript. We address each major comment below and will make the requested changes to improve methodological transparency and statistical reporting.
read point-by-point responses
-
Referee: [Abstract] Abstract: the composite scoring function that produces the DPO preference pairs is described only qualitatively; no equation, weighting scheme, normalization procedure, or validity proxy (e.g., cross-lingual prediction-flip detection) is supplied. Because the reported 12.55% validity gain and preserved minimality rest directly on the ordering induced by this function, its misspecification would render the gains artifacts rather than genuine alignment improvements.
Authors: We agree that the composite scoring function requires a formal mathematical description to ensure reproducibility and to confirm that the preference ordering is well-specified. In the revised manuscript we will add the explicit equation for the composite score, the weighting scheme between validity and minimality terms, the normalization procedure, and the precise validity proxy (cross-lingual prediction-flip detection) in the Methods section. revision: yes
-
Referee: [Abstract] Abstract / Experiments: the 12.55% validity improvement is stated without statistical details (standard errors, confidence intervals, number of runs, or significance tests) or per-language/per-model breakdowns; this prevents verification that the gain is robust rather than driven by a subset of languages or models.
Authors: We agree that statistical details and breakdowns are essential for verifying robustness. In the revised Experiments section we will report standard errors, confidence intervals, the number of runs, significance tests, and full per-language and per-model tables so that readers can confirm the consistency of the 12.55% average gain. revision: yes
Circularity Check
No circularity: method uses external DPO on independently scored pairs; results are experimental comparisons
full rationale
The paper introduces Macro as DPO applied to preference pairs built from a composite validity+minimality scorer. No equations, derivations, or 'predictions' are shown that reduce by construction to author-defined inputs or self-citations. Validity and minimality are measured on held-out test instances across models and languages, independent of the training signal construction. The central claim rests on empirical deltas (12.55% validity gain) rather than any self-referential loop. This is the normal non-circular case for an applied alignment paper.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization." pith.science (2026). https://pith.science/paper/7P4MQPRZ
@misc{pith2026260511632,
author = {Pith},
title = {Pith review of: Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization},
year = {2026},
howpublished = {\url{https://pith.science/paper/7P4MQPRZ}},
note = {Machine review of arXiv:2605.11632}
}
read the original abstract
Self-generated counterfactual explanations (SCEs) are minimally modified inputs (minimality) generated by large language models (LLMs) that flip their own predictions (validity), offering a causally grounded approach to unraveling black-box LLM behavior. Yet extending them beyond English remains challenging: existing methods struggle to produce valid SCEs in non-dominant languages, and a persistent trade-off between validity and minimality undermines explanation quality. We introduce Macro, a preference alignment framework that applies Direct Preference Optimization (DPO) to multilingual SCE generation, using a composite scoring function to construct preference pairs that effectively translate the trade-off into measurable preference signals. Experiments across four LLMs and seven typologically diverse languages show that Macro improves validity by 12.55\% on average over the chain-of-thought baseline without degrading minimality, while avoiding the severe minimality violations of the translation-based baseline. Compared to supervised fine-tuning, Macro achieves superior performance on both metrics, confirming that explicit preference optimization is essential for balancing this trade-off. Further analyses reveal that Macro increases cross-lingual perturbation alignment and mitigates common generation errors. Our results highlight preference optimization as a promising direction for enhancing multilingual model explanations.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Measuring massive multitask language under- standing. InInternational Conference on Learning Representations. Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec,...
work page Pith review arXiv 2025
-
[2]
InAdvances in Neural Information Processing Systems, volume 36, pages 53728–53741
Direct preference optimization: Your language model is secretly a reward model. InAdvances in Neural Information Processing Systems, volume 36, pages 53728–53741. Curran Associates, Inc. Lucas Resck, Isabelle Augenstein, and Anna Korhonen
-
[3]
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Explainability and interpretability of multilin- gual large language models: A survey. InProceed- ings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 20465–20497, Suzhou, China. Association for Computational Lin- guistics. Alexis Ross, Ana Marasovi´c, and Matthew Peters. 2021. Explaining NLP models via minimal contrastiv...
work page Pith review arXiv 2025
-
[4]
andMMLU(Hendrycks et al., 2021) datasets. MMLUis a widely used benchmark for measuring multitask language understanding across 57 sub- jects. In contrast,MMLU-ProXextends the more challengingMMLU-Probenchmark, which fea- tures reasoning-focused questions and ten answer choices instead of four. Figures 11, 12 and 13 demonstrate that, in gen- eral, models a...
Reviewed June 30, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.