Pith. sign in

REVIEW 1 cited by

Large-Scale Contextualised Language Modelling for Norwegian

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.06546 v1 pith:KM4FGGRJ submitted 2021-04-13 cs.CL

classification cs.CL
keywords norwegianlanguagemodelscontextualiseddatalarge-scalenorlmsoftware
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the ongoing NorLM initiative to support the creation and use of very large contextualised language models for Norwegian (and in principle other Nordic languages), including a ready-to-use software environment, as well as an experience report for data preparation and training. This paper introduces the first large-scale monolingual language models for Norwegian, based on both the ELMo and BERT frameworks. In addition to detailing the training process, we present contrastive benchmark results on a suite of NLP tasks for Norwegian. For additional background and access to the data, models, and software, please see http://norlm.nlpl.eu

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Development of Pre-Trained Transformer-based Models for the Nepali Language

    cs.CL 2024-11 conditional novelty 5.0 of 10

    Nepali-language BERT, RoBERTa, and GPT-2 models trained on a new 27.5 GB corpus score 95.60 on Nep-gLUE, beating the prior best by about two points.

Pith tools