Pith. sign in

REVIEW 1 cited by

Linguistic Interpretability of Transformer-based Language Models: a systematic review

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.08001 v1 pith:VZODGZU5 submitted 2025-04-09 cs.CL

classification cs.CL
keywords modelsinterpretabilitylinguisticlanguagesurveyachieveanalysisarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models based on the Transformer architecture achieve excellent results in many language-related tasks, such as text classification or sentiment analysis. However, despite the architecture of these models being well-defined, little is known about how their internal computations help them achieve their results. This renders these models, as of today, a type of 'black box' systems. There is, however, a line of research -- 'interpretability' -- aiming to learn how information is encoded inside these models. More specifically, there is work dedicated to studying whether Transformer-based models possess knowledge of linguistic phenomena similar to human speakers -- an area we call 'linguistic interpretability' of these models. In this survey we present a comprehensive analysis of 160 research works, spread across multiple languages and models -- including multilingual ones -- that attempt to discover linguistic information from the perspective of several traditional Linguistics disciplines: Syntax, Morphology, Lexico-Semantics and Discourse. Our survey fills a gap in the existing interpretability literature, which either not focus on linguistic knowledge in these models or present some limitations -- e.g. only studying English-based models. Our survey also focuses on Pre-trained Language Models not further specialized for a downstream task, with an emphasis on works that use interpretability techniques that explore models' internal representations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models

    cs.CL 2026-01 conditional novelty 4.0 of 10

    A systematic review of 337 articles shows Transformers handle formal syntax well but perform worse and more variably at the syntax-semantics interface, with the field over-reliant on English and BERT.

Pith tools