Pith. sign in

REVIEW 3 cited by

German's Next Language Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.10906 v4 pith:KCMZZ64O submitted 2020-10-21 cs.CL cs.LG

classification cs.CLcs.LG
keywords modelsgermanmodeldatalanguageperformancesizetraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work we present the experiments which lead to the creation of our BERT and ELECTRA based German language models, GBERT and GELECTRA. By varying the input training data, model size, and the presence of Whole Word Masking (WWM) we were able to attain SoTA performance across a set of document classification and named entity recognition (NER) tasks for both models of base and large size. We adopt an evaluation driven approach in training these models and our results indicate that both adding more data and utilizing WWM improve model performance. By benchmarking against existing German models, we show that these models are the best German models to date. Our trained models will be made publicly available to the research community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Llama-GENBA-10B: A Trilingual Large Language Model for German, English and Bavarian

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Llama-GENBA-10B is a 10B-parameter trilingual model that reports top Bavarian scores among sub-10B models on a machine-translated benchmark the authors built.

  2. Digital Guardians: Can GPT-4, Perspective API, and Moderation API reliably detect hate speech in reader comments of German online newspapers?

    cs.CL 2025-01 conditional novelty 4.0 of 10

    On the German HOCON34k hate speech test set, GPT-4o with one-shot prompting scored highest on the combined F2/MCC metric, about 5 points above the fine-tuned BERT baseline.

  3. Efficient Multi-Task Inferencing with a Shared Backbone and Lightweight Task-Specific Adapters for Automatic Scoring

    cs.CL 2024-12 conditional novelty 3.0 of 10

    A frozen BERT-style backbone with per-task LoRA adapters scores 27 PISA items at 60% lower GPU memory and 40% lower latency, with a 4.5% QWK drop.

Pith tools