Pith. sign in

REVIEW 1 cited by

Language Resources for Dutch Large Language Modelling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.12852 v1 pith:CWJQ34NW submitted 2023-12-20 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelsdutchlanguagemodeldatadatasetsfine-tunedfirst
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the rapid expansion of types of large language models, there remains a notable gap in models specifically designed for the Dutch language. This gap is not only a shortage in terms of pretrained Dutch models but also in terms of data, and benchmarks and leaderboards. This work provides a small step to improve the situation. First, we introduce two fine-tuned variants of the Llama 2 13B model. We first fine-tuned Llama 2 using Dutch-specific web-crawled data and subsequently refined this model further on multiple synthetic instruction and chat datasets. These datasets as well as the model weights are made available. In addition, we provide a leaderboard to keep track of the performance of (Dutch) models on a number of generation tasks, and we include results of a number of state-of-the-art models, including our own. Finally we provide a critical conclusion on what we believe is needed to push forward Dutch language models and the whole eco-system around the models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Interpretable phenotyping of Heart Failure patients with Dutch discharge letters

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Discharge letters alone let BERT and Aug-Linear models classify HFrEF vs HFpEF with external-validation AUCs of 0.84 and 0.81, and Aug-Linear explanations matched clinicians better than SHAP/LIME.

Pith tools