Pith. sign in

REVIEW 2 major objections 1 minor 4 references

Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization

T0 review · 2 major / 1 minor · reviewed 2026-06-30 · grok-4.3

Pith's one-line read Macro applies direct preference optimization to multilingual self-generated counterfactual explanations, raising validity by 12.55 percent on average over chain-of-thought baselines while preserving minimality.

desk verdict Macro uses DPO plus a composite score to lift multilingual counterfactual validity, but the score itself is the part that still needs checking. read the letter →

arxiv 2605.11632 v2 pith:7P4MQPRZ submitted 2026-05-12 cs.CL cs.AI

classification cs.CLcs.AI
keywords counterfactualexplanationsmultilingualLLMsdirectpreferenceoptimizationvalidityminimalitytrade-offself-generatedmodelinterpretabilityalignmentmethods
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents Macro, a framework that turns the validity-minimality trade-off in counterfactual explanation generation into explicit preference pairs for direct preference optimization. It tests this approach on four large language models across seven typologically diverse languages and reports consistent gains in validity without the minimality losses seen in translation baselines. The work matters because valid and minimal explanations are needed to interpret model decisions in languages other than English, where current methods either fail to flip predictions or produce overly altered inputs. By showing that preference optimization outperforms both chain-of-thought and supervised fine-tuning on the combined metrics, the authors argue that explicit alignment is required to balance the two objectives in multilingual settings.

What carries the argument

A composite scoring function that converts the validity-minimality trade-off into measurable preference signals for direct preference optimization.

What would settle it

Running the same four models on an eighth typologically distant language and finding that validity gains are accompanied by statistically larger minimality violations or new error patterns would falsify the central claim.

Watch

Extended reading notes

Core claim

Macro constructs preference pairs for direct preference optimization from a composite scoring function that rewards validity (prediction flip) and minimality (small input change) in self-generated counterfactual explanations. When applied to multilingual generation, this alignment step raises average validity by 12.55 percent relative to chain-of-thought prompting, keeps minimality intact, and outperforms supervised fine-tuning on both metrics. The same method also increases cross-lingual perturbation alignment and reduces common generation errors.

Load-bearing premise

The composite scoring function produces reliable preference signals that direct preference optimization can follow without creating new biases or artifacts in the generated explanations.

Editorial extensions

If this is right

  • Macro produces higher cross-lingual perturbation alignment than the tested baselines.
  • It reduces common generation errors that appear in chain-of-thought and translation-based outputs.
  • It outperforms supervised fine-tuning on both validity and minimality simultaneously.
  • It avoids the severe minimality violations observed in the translation-based baseline.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The preference-pair construction could be reused for other explanation formats that face similar validity-minimality tensions.
  • The gains observed across typologically diverse languages suggest the method may transfer to additional low-resource languages not tested here.
  • If the composite scoring function generalizes, similar alignment pipelines might improve other multilingual generation tasks that require controlled edits.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper introduces Macro, a preference alignment framework applying Direct Preference Optimization (DPO) to multilingual self-generated counterfactual explanations (SCEs). It uses a composite scoring function to build preference pairs from the validity-minimality trade-off and reports that this yields a 12.55% average validity gain over chain-of-thought baselines across four LLMs and seven typologically diverse languages, without degrading minimality and while outperforming translation-based and supervised fine-tuning approaches.

Significance. If the results hold after full methodological disclosure and validation of the scoring function, the work would be significant for multilingual XAI: it provides evidence that explicit preference optimization can resolve the validity-minimality tension in counterfactual generation where supervised methods fall short, and it demonstrates cross-lingual perturbation alignment improvements.

major comments (2)
  1. [Abstract] Abstract: the composite scoring function that produces the DPO preference pairs is described only qualitatively; no equation, weighting scheme, normalization procedure, or validity proxy (e.g., cross-lingual prediction-flip detection) is supplied. Because the reported 12.55% validity gain and preserved minimality rest directly on the ordering induced by this function, its misspecification would render the gains artifacts rather than genuine alignment improvements.
  2. [Abstract] Abstract / Experiments: the 12.55% validity improvement is stated without statistical details (standard errors, confidence intervals, number of runs, or significance tests) or per-language/per-model breakdowns; this prevents verification that the gain is robust rather than driven by a subset of languages or models.
minor comments (1)
  1. [Abstract] Abstract: the phrase 'further analyses reveal that Macro increases cross-lingual perturbation alignment' is asserted without naming the metrics or showing the supporting figures/tables.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive comments on our manuscript. We address each major comment below and will make the requested changes to improve methodological transparency and statistical reporting.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the composite scoring function that produces the DPO preference pairs is described only qualitatively; no equation, weighting scheme, normalization procedure, or validity proxy (e.g., cross-lingual prediction-flip detection) is supplied. Because the reported 12.55% validity gain and preserved minimality rest directly on the ordering induced by this function, its misspecification would render the gains artifacts rather than genuine alignment improvements.

    Authors: We agree that the composite scoring function requires a formal mathematical description to ensure reproducibility and to confirm that the preference ordering is well-specified. In the revised manuscript we will add the explicit equation for the composite score, the weighting scheme between validity and minimality terms, the normalization procedure, and the precise validity proxy (cross-lingual prediction-flip detection) in the Methods section. revision: yes

  2. Referee: [Abstract] Abstract / Experiments: the 12.55% validity improvement is stated without statistical details (standard errors, confidence intervals, number of runs, or significance tests) or per-language/per-model breakdowns; this prevents verification that the gain is robust rather than driven by a subset of languages or models.

    Authors: We agree that statistical details and breakdowns are essential for verifying robustness. In the revised Experiments section we will report standard errors, confidence intervals, the number of runs, significance tests, and full per-language and per-model tables so that readers can confirm the consistency of the 12.55% average gain. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: method uses external DPO on independently scored pairs; results are experimental comparisons

full rationale

The paper introduces Macro as DPO applied to preference pairs built from a composite validity+minimality scorer. No equations, derivations, or 'predictions' are shown that reduce by construction to author-defined inputs or self-citations. Validity and minimality are measured on held-out test instances across models and languages, independent of the training signal construction. The central claim rests on empirical deltas (12.55% validity gain) rather than any self-referential loop. This is the normal non-circular case for an applied alignment paper.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Only the abstract is available; no explicit free parameters, axioms, or invented entities are described.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization." pith.science (2026). https://pith.science/paper/7P4MQPRZ

@misc{pith2026260511632,
  author       = {Pith},
  title        = {Pith review of: Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7P4MQPRZ}},
  note         = {Machine review of arXiv:2605.11632}
}
read the original abstract

Self-generated counterfactual explanations (SCEs) are minimally modified inputs (minimality) generated by large language models (LLMs) that flip their own predictions (validity), offering a causally grounded approach to unraveling black-box LLM behavior. Yet extending them beyond English remains challenging: existing methods struggle to produce valid SCEs in non-dominant languages, and a persistent trade-off between validity and minimality undermines explanation quality. We introduce Macro, a preference alignment framework that applies Direct Preference Optimization (DPO) to multilingual SCE generation, using a composite scoring function to construct preference pairs that effectively translate the trade-off into measurable preference signals. Experiments across four LLMs and seven typologically diverse languages show that Macro improves validity by 12.55\% on average over the chain-of-thought baseline without degrading minimality, while avoiding the severe minimality violations of the translation-based baseline. Compared to supervised fine-tuning, Macro achieves superior performance on both metrics, confirming that explicit preference optimization is essential for balancing this trade-off. Further analyses reveal that Macro increases cross-lingual perturbation alignment and mitigates common generation errors. Our results highlight preference optimization as a promising direction for enhancing multilingual model explanations.

Figures

Figures reproduced from arXiv: 2605.11632 by the authors.

Figure 1
Figure 1. Overview of our three-stage framework (MACRO). Stage 1 samples counterfactual candidates via word￾level perturbations across multilingual inputs. Stage 2 ranks candidates using Rflip, Raug, and Redit to construct preference pairs. Stage 3 applies DPO to align the model toward generating minimal, effective counterfactuals. achieved without degrading minimality, marking a pronounced distinction from the translation-ba… view at source ↗
Figure 2
Figure 2. The validity-minimality trade-off across lan [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Relative performance change across languages for [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (11 more)
Figure 4
Figure 4. Figure 4: Cross-lingual edit similarity score changes [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 6
Figure 6. Figure 6: Total score distributions before and after ap [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Label distributions of the two evaluation [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: Prediction prompts used for the two evaluation datasets: [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 9
Figure 9. Figure 9: Counterfactual generation prompts used for [PITH_FULL_IMAGE:figures/full_fig_p021_9.png]
Figure 10
Figure 10. Figure 10: Dataset examples [PITH_FULL_IMAGE:figures/full_fig_p022_10.png]
Figure 11
Figure 11. Figure 11: Impact of MACRO on multilingual general capability measured on MMLU [PITH_FULL_IMAGE:figures/full_fig_p027_11.png]
Figure 12
Figure 12. Figure 12: Impact of MACRO on reasoning capability measured on MMLU-ProX from the category perspective . Subfigures (a) and (b) present the category-wise performance of Qwen3-4B and Gemma3-4B, respectively [PITH_FULL_IMAGE:figures/full_fig_p028_12.png]
Figure 13
Figure 13. Figure 13: Impact of MACRO on cross-lingual generalization measured on MMLU-ProX from the language perspec￾tive. Subfigures (a) and (b) present the language-wise performance of Qwen3-4B and Gemma3-4B, respectively [PITH_FULL_IMAGE:figures/full_fig_p029_13.png]
Figure 14
Figure 14. Figure 14: The validity-minimality trade-off across languages across all models on [PITH_FULL_IMAGE:figures/full_fig_p030_14.png]
Figure 15
Figure 15. Figure 15: Cross-lingual edit similarity scores [PITH_FULL_IMAGE:figures/full_fig_p030_15.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

4 extracted references · 4 canonical work pages

  1. [1]

    Gemma 3 Technical Report

    Measuring massive multitask language under- standing. InInternational Conference on Learning Representations. Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec,...

  2. [2]

    InAdvances in Neural Information Processing Systems, volume 36, pages 53728–53741

    Direct preference optimization: Your language model is secretly a reward model. InAdvances in Neural Information Processing Systems, volume 36, pages 53728–53741. Curran Associates, Inc. Lucas Resck, Isabelle Augenstein, and Anna Korhonen

  3. [3]

    DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

    Explainability and interpretability of multilin- gual large language models: A survey. InProceed- ings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 20465–20497, Suzhou, China. Association for Computational Lin- guistics. Alexis Ross, Ana Marasovi´c, and Matthew Peters. 2021. Explaining NLP models via minimal contrastiv...

  4. [4]

    science/technology

    andMMLU(Hendrycks et al., 2021) datasets. MMLUis a widely used benchmark for measuring multitask language understanding across 57 sub- jects. In contrast,MMLU-ProXextends the more challengingMMLU-Probenchmark, which fea- tures reasoning-focused questions and ten answer choices instead of four. Figures 11, 12 and 13 demonstrate that, in gen- eral, models a...

Pith tools

Reviewed June 30, 2026 · model on record in the stance chip above.