Pith. sign in

REVIEW 2 cited by

RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.19232 v3 pith:67I2EXV5 submitted 2024-06-27 cs.CL

classification cs.CL
keywords pairsminimallinguisticrublimplanguagemodelsrussianbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Minimal pairs are a well-established approach to evaluating the grammatical knowledge of language models. However, existing resources for minimal pairs address a limited number of languages and lack diversity of language-specific grammatical phenomena. This paper introduces the Russian Benchmark of Linguistic Minimal Pairs (RuBLiMP), which includes 45k pairs of sentences that differ in grammaticality and isolate a morphological, syntactic, or semantic phenomenon. In contrast to existing benchmarks of linguistic minimal pairs, RuBLiMP is created by applying linguistic perturbations to automatically annotated sentences from open text corpora and carefully curating test data. We describe the data collection protocol and present the results of evaluating 25 language models in various scenarios. We find that the widely used language models for Russian are sensitive to morphological and agreement-oriented contrasts but fall behind humans on phenomena requiring understanding of structural relations, negation, transitivity, and tense. RuBLiMP, the codebase, and other materials are publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Spontaneous Speech Variables for Evaluating LLMs Cognitive Plausibility

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Language models fine-tuned to predict speech reductions and prosodic prominences from text perform above random baselines, and models pretrained on conversational data outperform those pretrained on written data in En...

  2. Explain-then-Process: Using Grammar Prompting to Enhance Grammatical Acceptability Judgments

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Feeding an LLM-generated grammar explanation back to a model before a grammaticality judgment improves minimal-pair accuracy, with the largest gains for smaller models.

Pith tools