Pith. sign in

REVIEW 1 cited by

Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.07474 v2 pith:ZPGTBRKF submitted 2024-11-12 cs.CL

classification cs.CL
keywords syntacticevaluationlanguagesmodelsmultilingualagreementbasquegeneralizations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models (LMs) are capable of acquiring elements of human-like syntactic knowledge. Targeted syntactic evaluation tests have been employed to measure how well they form generalizations about syntactic phenomena in high-resource languages such as English. However, we still lack a thorough understanding of LMs' capacity for syntactic generalizations in low-resource languages, which are responsible for much of the diversity of syntactic patterns worldwide. In this study, we develop targeted syntactic evaluation tests for three low-resource languages (Basque, Hindi, and Swahili) and use them to evaluate five families of open-access multilingual Transformer LMs. We find that some syntactic tasks prove relatively easy for LMs while others (agreement in sentences containing indirect objects in Basque, agreement across a prepositional phrase in Swahili) are challenging. We additionally uncover issues with publicly available Transformers, including a bias toward the habitual aspect in Hindi in multilingual BERT and underperformance compared to similar-sized models in XGLM-4.5B.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu

    cs.CL 2025-08 conditional novelty 7.0 of 10

    A new minimal-pair benchmark for Urdu grammar shows that multilingual models vary widely across syntactic phenomena, with LLaMA-3-70B best at 94.7% accuracy.

Pith tools