Pith. sign in

REVIEW 2 cited by

Playing with Words at the National Library of Sweden -- Making a Swedish BERT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.01658 v1 pith:ROGZYKTO submitted 2020-07-03 cs.CL

classification cs.CL
keywords swedishbertmodelcreatekb-bertlanguageslibrarymodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper introduces the Swedish BERT ("KB-BERT") developed by the KBLab for data-driven research at the National Library of Sweden (KB). Building on recent efforts to create transformer-based BERT models for languages other than English, we explain how we used KB's collections to create and train a new language-specific BERT model for Swedish. We also present the results of our model in comparison with existing models - chiefly that produced by the Swedish Public Employment Service, Arbetsf\"ormedlingen, and Google's multilingual M-BERT - where we demonstrate that KB-BERT outperforms these in a range of NLP tasks from named entity recognition (NER) to part-of-speech tagging (POS). Our discussion highlights the difficulties that continue to exist given the lack of training data and testbeds for smaller languages like Swedish. We release our model for further exploration and research here: https://github.com/Kungbib/swedish-bert-models .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Matching and Linking Entries in Historical Swedish Encyclopedias

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Using embeddings and Wikidata linking, the authors find a small geographic shift in the Nordisk familjebok's entries between its first and second editions, away from Europe and toward the rest of the world.

  2. Prediction-powered estimators for finite population statistics in highly imbalanced textual data: Public hate crime estimation

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Using a BERT classifier's predicted hate-crime probabilities as an auxiliary sampling variable yields a Hansen-Hurwitz estimate of 6,051 hate crimes among 2022 Swedish police reports, with a design effect of 0.0068.

Pith tools