Pith. sign in

REVIEW 2 cited by

G-SciEdBERT: A Contextualized LLM for Science Assessment Tasks in German

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.06584 v2 pith:QJA27V4C submitted 2024-02-09 cs.CL cs.AI

classification cs.CLcs.AI
keywords scoringg-sciedbertgermanscienceg-bertresponsesaccuracycontextualized
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The advancement of natural language processing has paved the way for automated scoring systems in various languages, such as German (e.g., German BERT [G-BERT]). Automatically scoring written responses to science questions in German is a complex task and challenging for standard G-BERT as they lack contextual knowledge in the science domain and may be unaligned with student writing styles. This paper presents a contextualized German Science Education BERT (G-SciEdBERT), an innovative large language model tailored for scoring German-written responses to science tasks and beyond. Using G-BERT, we pre-trained G-SciEdBERT on a corpus of 30K German written science responses with 3M tokens on the Programme for International Student Assessment (PISA) 2018. We fine-tuned G-SciEdBERT on an additional 20K student-written responses with 2M tokens and examined the scoring accuracy. We then compared its scoring performance with G-BERT. Our findings revealed a substantial improvement in scoring accuracy with G-SciEdBERT, demonstrating a 10.2% increase of quadratic weighted Kappa compared to G-BERT (mean difference = 0.1026, SD = 0.069). These insights underline the significance of specialized language models like G-SciEdBERT, which is trained to enhance the accuracy of contextualized automated scoring, offering a substantial contribution to the field of AI in education.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 3 citations worldwide. Full citation record

  1. AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics

    physics.ed-ph 2026-07 conditional novelty 6.0 of 10

    Every tested AI scorer systematically underestimates conceptual understanding in linguistically weaker secondary-school physics explanations relative to trained experts.

  2. Efficient Multi-Task Inferencing with a Shared Backbone and Lightweight Task-Specific Adapters for Automatic Scoring

    cs.CL 2024-12 conditional novelty 3.0 of 10

    A frozen BERT-style backbone with per-task LoRA adapters scores 27 PISA items at 60% lower GPU memory and 40% lower latency, with a 4.5% QWK drop.

Pith tools