Pith. sign in

REVIEW 2 cited by

Latin BERT: A Contextual Language Model for Classical Philology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.10053 v1 pith:GFIBYQ6W submitted 2020-09-21 cs.CL

classification cs.CL
keywords latinbertlanguagecontextualmodelclassicaltrainedused
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Latin BERT, a contextual language model for the Latin language, trained on 642.7 million words from a variety of sources spanning the Classical era to the 21st century. In a series of case studies, we illustrate the affordances of this language-specific model both for work in natural language processing for Latin and in using computational methods for traditional scholarship: we show that Latin BERT achieves a new state of the art for part-of-speech tagging on all three Universal Dependency datasets for Latin and can be used for predicting missing text (including critical emendations); we create a new dataset for assessing word sense disambiguation for Latin and demonstrate that Latin BERT outperforms static word embeddings; and we show that it can be used for semantically-informed search by querying contextual nearest neighbors. We publicly release trained models to help drive future work in this space.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From transcription to semantic corpus analysis: unsupervised learning of sentence representations for ancient languages

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Unsupervised TSDAE and CSE adaptation of specialized Latin/Greek LMs yields corpus-specific sentence encoders that outperform multilingual, distilled, and supervised baselines on biblical reuse detection and retrieval.

  2. Reviving Cultural Heritage: A Novel Approach for Comprehensive Historical Document Restoration

    cs.CV 2025-07 conditional novelty 6.0 of 10

    AutoHDR restores full pages of damaged historical documents by combining OCR damage detection, LLM-based text prediction, and patch-autoregressive image diffusion, improving OCR accuracy on severely damaged documents ...

Pith tools