Pith. sign in

A primer in BERTology: What we know about how BERT works.arXiv [cs.CL]

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it
abstract

Transformer-based models have pushed state of the art in many areas of NLP, but our understanding of what is behind their success is still limited. This paper is the first survey of over 150 studies of the popular BERT model. We review the current state of knowledge about how BERT works, what kind of information it learns and how it is represented, common modifications to its training objectives and architecture, the overparameterization issue and approaches to compression. We then outline directions for future research.

citation-role summary

background 1

citation-polarity summary

years

2026 5

roles

background 1

polarities

background 1

representative citing papers

Riemannian Geometry for Pre-trained Language Model Embeddings

cs.CL · 2026-07-08 · conditional · novelty 6.0

Aggregating per-token pullback metrics via the Fréchet mean on the SPD manifold outperforms Euclidean mean pooling for sentence classification, with most of the gain attributable to geometric aggregation rather than learned encoder structure.

citing papers explorer

Showing 5 of 5 citing papers.