Pith. sign in

REVIEW 1 cited by

DUMB: A Benchmark for Smart Evaluation of Dutch Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.13026 v2 pith:ZXJKOOT7 submitted 2023-05-22 cs.CL

classification cs.CL
keywords dutchmodelstasksbenchmarkdumblanguageperformanceavailable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce the Dutch Model Benchmark: DUMB. The benchmark includes a diverse set of datasets for low-, medium- and high-resource tasks. The total set of nine tasks includes four tasks that were previously not available in Dutch. Instead of relying on a mean score across tasks, we propose Relative Error Reduction (RER), which compares the DUMB performance of language models to a strong baseline which can be referred to in the future even when assessing different sets of language models. Through a comparison of 14 pre-trained language models (mono- and multi-lingual, of varying sizes), we assess the internal consistency of the benchmark tasks, as well as the factors that likely enable high performance. Our results indicate that current Dutch monolingual models under-perform and suggest training larger Dutch models with other architectures and pre-training objectives. At present, the highest performance is achieved by DeBERTaV3 (large), XLM-R (large) and mDeBERTaV3 (base). In addition to highlighting best strategies for training larger Dutch models, DUMB will foster further research on Dutch. A public leaderboard is available at https://dumbench.nl.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. skLEP: A Slovak General Language Understanding Benchmark

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A nine-task Slovak-language understanding benchmark with translated and newly curated datasets, plus the first broad fine-tuned model comparison for Slovak.

Pith tools