Pith. sign in

REVIEW 2 cited by

Linguini: A benchmark for language-agnostic linguistic reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.12126 v1 pith:44GVDLV3 submitted 2024-09-18 cs.CL

classification cs.CL
keywords linguisticbenchmarkmodelmodelsaccuracybest-performingknowledgelanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a new benchmark to measure a language model's linguistic reasoning skills without relying on pre-existing language-specific knowledge. The test covers 894 questions grouped in 160 problems across 75 (mostly) extremely low-resource languages, extracted from the International Linguistic Olympiad corpus. To attain high accuracy on this benchmark, models don't need previous knowledge of the tested language, as all the information needed to solve the linguistic puzzle is presented in the context. We find that, while all analyzed models rank below 25% accuracy, there is a significant gap between open and closed models, with the best-performing proprietary model at 24.05% and the best-performing open model at 8.84%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KoBALT: Korean Benchmark For Advanced Linguistic Tasks

    cs.CL 2025-05 conditional novelty 6.0 of 10

    KoBALT, an expert-crafted 700-question Korean linguistic benchmark, finds that even the best LLM answers only 61% correctly, with human preference ratings correlating moderately with benchmark accuracy.

  2. LingBench++: A Linguistically-Informed Benchmark and Reasoning Framework for Multi-Step and Cross-Cultural Inference with LLMs

    cs.CL 2025-07 conditional novelty 5.0 of 10

    LingBench++ adds stepwise reasoning evaluation, typological metadata, and a retrieval-augmented multi-agent framework to IOL-style linguistic puzzles, but the claimed gains rest on unreplicated single-run experiments.

Pith tools