Pith. sign in

REVIEW 1 cited by

Can language models handle recursively nested grammatical structures? A case study on comparing models and humans

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.15303 v3 pith:DHQT7XYB submitted 2022-10-27 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords humansmodelsstructureshumannestedcasegrammaticallanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

How should we compare the capabilities of language models (LMs) and humans? I draw inspiration from comparative psychology to highlight some challenges. In particular, I consider a case study: processing of recursively nested grammatical structures. Prior work suggests that LMs cannot handle these structures as reliably as humans can. However, the humans were provided with instructions and training, while the LMs were evaluated zero-shot. I therefore match the evaluation more closely. Providing large LMs with a simple prompt -- substantially less content than the human training -- allows the LMs to consistently outperform the human results, and even to extrapolate to more deeply nested conditions than were tested with humans. Further, reanalyzing the prior human data suggests that the humans may not perform above chance at the difficult structures initially. Thus, large LMs may indeed process recursively nested grammatical structures as reliably as humans. This case study highlights how discrepancies in the evaluation can confound comparisons of language models and humans. I therefore reflect on the broader challenge of comparing human and model capabilities, and highlight an important difference between evaluating cognitive models and foundation models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Human-Like Processing: Large Language Models Perform Equivalently on Forward and Backward Scientific Text

    cs.CL 2024-11 conditional novelty 5.0 of 10

    GPT-2 models trained on character-reversed neuroscience text perform as well on a neuroscience abstract-selection benchmark as models trained on normal text, despite higher perplexity.

Pith tools