Pith. sign in

REVIEW 3 cited by

Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.20850 v3 pith:GJEGMZBM submitted 2025-03-26 cs.CL

classification cs.CL
keywords preferencesdirectevidencedativeindirectlanguagelengthproperties
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models (LMs) tend to show human-like preferences on a number of syntactic phenomena, but the extent to which these are attributable to direct exposure to the phenomena or more general properties of language is unclear. We explore this with the English dative alternation (DO: "gave Y the X" vs. PO: "gave the X to Y"), using a controlled rearing paradigm wherein we iteratively train small LMs on systematically manipulated input. We focus on two properties that affect the choice of alternant: length and animacy. Both properties are directly present in datives but also reflect more global tendencies for shorter elements to precede longer ones and animates to precede inanimates. First, by manipulating and ablating datives for these biases in the input, we show that direct evidence of length and animacy matters, but easy-first preferences persist even without such evidence. Then, using LMs trained on systematically perturbed datasets to manipulate global length effects (re-linearizing sentences globally while preserving dependency structure), we find that dative preferences can emerge from indirect evidence. We conclude that LMs' emergent syntactic preferences come from a mix of direct and indirect sources.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the Interpretability of Whisper Encodings Using Sparse Autoencoders

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    Sparse autoencoders applied to Whisper ASR reveal monosemantic features across linguistic boundaries and demonstrate cross-lingual feature steering.

  2. The Holistic Storage of Verb+Up Phrases in Text-based and Audio-based Language Models

    cs.CL 2026-06 conditional novelty 6.0 of 10

    Frequency and predictability shift the internal representation of “up” inside V+up phrases away from standalone “up” in text and audio models, a pattern interpreted as holistic storage.

  3. semantic-features: A User-Friendly Tool for Studying Contextual Word Embeddings in Interpretable Semantic Spaces

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A library and demo reveal that masked language models encode the dative construction's person-like versus place-like reading of ambiguous recipients.

Pith tools