Pith. sign in

REVIEW 1 cited by

Language Model Behavior: A Comprehensive Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.11504 v2 pith:OKX5W4HF submitted 2023-03-20 cs.CL

classification cs.CL
keywords languagemodelstextcapabilitiesmodelbehaviorgeneratedrecent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformer language models have received widespread public attention, yet their generated text is often surprising even to NLP researchers. In this survey, we discuss over 250 recent studies of English language model behavior before task-specific fine-tuning. Language models possess basic capabilities in syntax, semantics, pragmatics, world knowledge, and reasoning, but these capabilities are sensitive to specific inputs and surface features. Despite dramatic increases in generated text quality as models scale to hundreds of billions of parameters, the models are still prone to unfactual responses, commonsense errors, memorized text, and social biases. Many of these weaknesses can be framed as over-generalizations or under-generalizations of learned patterns in text. We synthesize recent results to highlight what is currently known about large language model capabilities, thus providing a resource for applied work and for research in adjacent fields that use language models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Research Community Perspectives on "Intelligence" and Large Language Models

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A survey of 303 researchers finds that generalization, adaptability, and reasoning are the most agreed-upon criteria for intelligence, and that most researchers do not consider current LLM-based systems intelligent.

Pith tools