Pith. sign in

REVIEW 1 cited by

Interference Matrix: Quantifying Cross-Lingual Interference in Transformer Encoders

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.02256 v1 pith:NXAGK3IQ submitted 2025-08-04 cs.CL

Interference Matrix: Quantifying Cross-Lingual Interference in Transformer Encoders

classification cs.CL
keywords interferencelanguagematrixmodelsbettercross-linguallanguagesperformance
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In this paper, we present a comprehensive study of language interference in encoder-only Transformer models across 83 languages. We construct an interference matrix by training and evaluating small BERT-like models on all possible language pairs, providing a large-scale quantification of cross-lingual interference. Our analysis reveals that interference between languages is asymmetrical and that its patterns do not align with traditional linguistic characteristics, such as language family, nor with proxies like embedding similarity, but instead better relate to script. Finally, we demonstrate that the interference matrix effectively predicts performance on downstream tasks, serving as a tool to better design multilingual models to obtain optimal performance.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Cross-Lingual Transfer for Machine Translation in Turkic Languages

    cs.CL 2026-07 conditional novelty 6.0

    Among five Turkic languages, mT5 transfer is strongest for Turkish–Azerbaijani and Kazakh–Kyrgyz, direction and translation target matter, and Latinization helps surface metrics mainly in script-mismatched pairs.