A new Norwegian summarization benchmark, NorSumm, contains 63 news articles each with three human-authored summaries in Bokmål and Nynorsk, and open LLMs perform poorly on it.
NorNE: Annotating Named Entities for Norwegian
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
This paper presents NorNE, a manually annotated corpus of named entities which extends the annotation of the existing Norwegian Dependency Treebank. Comprising both of the official standards of written Norwegian (Bokm{\aa}l and Nynorsk), the corpus contains around 600,000 tokens and annotates a rich set of entity types including persons, organizations, locations, geo-political entities, products, and events, in addition to a class corresponding to nominals derived from names. We here present details on the annotation effort, guidelines, inter-annotator agreement and an experimental analysis of the corpus using a neural sequence labeling architecture.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Benchmarking Abstractive Summarisation: A Dataset of Human-authored Summaries of Norwegian News Articles
A new Norwegian summarization benchmark, NorSumm, contains 63 news articles each with three human-authored summaries in Bokmål and Nynorsk, and open LLMs perform poorly on it.