Pith. sign in

REVIEW 2 cited by

NLEBench+NorGLM: A Comprehensive Empirical Analysis and Benchmark Dataset for Generative Language Models in Norwegian

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.01314 v2 pith:RNAUN5GQ submitted 2023-12-03 cs.CL

classification cs.CL
keywords norwegianlanguagetasksdatasetmodelscomprehensivebenchmarkcapability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Norwegian, spoken by only 5 million population, is under-representative within the most impressive breakthroughs in NLP tasks. To the best of our knowledge, there has not yet been a comprehensive evaluation of the existing language models (LMs) on Norwegian generation tasks during the article writing process. To fill this gap, we 1) compiled the existing Norwegian dataset and pre-trained 4 Norwegian Open Language Models varied from parameter scales and architectures, collectively called NorGLM; 2) introduced a comprehensive benchmark, NLEBench, for evaluating natural language generation capabilities in Norwegian, encompassing translation and human annotation. Based on the investigation, we find that: 1) the mainstream, English-dominated LM GPT-3.5 has limited capability in understanding the Norwegian context; 2) the increase in model parameter scales demonstrates limited impact on the performance of downstream tasks when the pre-training dataset is constrained in size; 3) smaller models also demonstrate the reasoning capability through Chain-of-Thought; 4) a multi-task dataset that includes synergy tasks can be used to verify the generalizability of LLMs on natural language understanding and, meanwhile, test the interconnectedness of these NLP tasks. We share our resources and code for reproducibility under a CC BY-NC 4.0 license.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A new Norwegian summarization benchmark, NorSumm, contains 63 news articles each with three human-authored summaries in Bokmål and Nynorsk, and open LLMs perform poorly on it.

  2. Small Languages, Big Models: A Study of Continual Training on Languages of Norway

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A three-stage continual training recipe (tokenizer change, embedding alignment, full retraining) produces NorMistral-11B, an open Norwegian and Northern Sámi language model that improves on most Norwegian benchmarks a...

Pith tools