REVIEW 2 cited by
The Secret is in the Spectra: Predicting Cross-lingual Task Performance with Spectral Similarity Measures
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Performance in cross-lingual NLP tasks is impacted by the (dis)similarity of languages at hand: e.g., previous work has suggested there is a connection between the expected success of bilingual lexicon induction (BLI) and the assumption of (approximate) isomorphism between monolingual embedding spaces. In this work we present a large-scale study focused on the correlations between monolingual embedding space similarity and task performance, covering thousands of language pairs and four different tasks: BLI, parsing, POS tagging and MT. We hypothesize that statistics of the spectrum of each monolingual embedding space indicate how well they can be aligned. We then introduce several isomorphism measures between two embedding spaces, based on the relevant statistics of their individual spectra. We empirically show that 1) language similarity scores derived from such spectral isomorphism measures are strongly associated with performance observed in different cross-lingual tasks, and 2) our spectral-based measures consistently outperform previous standard isomorphism measures, while being computationally more tractable and easier to interpret. Finally, our measures capture complementary information to typologically driven language distance measures, and the combination of measures from the two families yields even higher task performance correlations.
Forward citations
Cited by 2 Pith papers
-
Interpretable Syntactic Representations Enable Hierarchical Word Vectors
A linear projection of Word2Vec and GloVe embeddings onto eight part-of-speech axes yields compact interpretable vectors, and combining them with the original vectors gives small gains on a few downstream tasks.
-
Uncovering Cross-Linguistic Disparities in LLMs using Sparse Autoencoders
Sparse-autoencoder analysis of Gemma-2-2B shows that medium-to-low resource languages get up to 26% lower activations than English, and LoRA fine-tuning that explicitly minimizes the activation gap raises activations ...
Discussion (0). Continue with ORCID to comment.