Pith. sign in

REVIEW 2 cited by

Implicit Representations of Meaning in Neural Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.00737 v1 pith:K2AXHRZV submitted 2021-06-01 cs.CL

classification cs.CL
keywords modelslanguageneuralrepresentationstheydatadynamicentity
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Does the effectiveness of neural language models derive entirely from accurate modeling of surface word co-occurrence statistics, or do these models represent and reason about the world they describe? In BART and T5 transformer language models, we identify contextual word representations that function as models of entities and situations as they evolve throughout a discourse. These neural representations have functional similarities to linguistic models of dynamic semantics: they support a linear readout of each entity's current properties and relations, and can be manipulated with predictable effects on language generation. Our results indicate that prediction in pretrained neural language models is supported, at least in part, by dynamic representations of meaning and implicit simulation of entity state, and that this behavior can be learned with only text as training data. Code and data are available at https://github.com/belindal/state-probes .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When Do Neural Networks Learn World Models?

    cs.LG 2025-02 conditional novelty 7.0 of 10

    With Boolean variables, a low-degree bias, and a task distribution weighted toward simple functions of the latents, multi-task training provably recovers the latent world model up to permutations and negations.

  2. Tracking World States with Language Models: State-Based Evaluation Using Chess

    cs.AI 2025-08 conditional novelty 6.0 of 10

    A model-agnostic chess evaluation metric measures state-tracking fidelity by comparing legal-move sets of predicted and true positions, showing GPT-4o's reconstruction quality degrades over longer games.

Pith tools