Pith. sign in

REVIEW 1 cited by

AustroTox: A Dataset for Target-Based Austrian German Offensive Language Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08080 v1 pith:IY2J77DO submitted 2024-06-12 cs.CL cs.AIcs.CY

classification cs.CLcs.AIcs.CY
keywords languagemodelsdetectionoffensiveannotationsaustrianaustrotoxdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model interpretability in toxicity detection greatly profits from token-level annotations. However, currently such annotations are only available in English. We introduce a dataset annotated for offensive language detection sourced from a news forum, notable for its incorporation of the Austrian German dialect, comprising 4,562 user comments. In addition to binary offensiveness classification, we identify spans within each comment constituting vulgar language or representing targets of offensive statements. We evaluate fine-tuned language models as well as large language models in a zero- and few-shot fashion. The results indicate that while fine-tuned models excel in detecting linguistic peculiarities such as vulgar dialect, large language models demonstrate superior performance in detecting offensiveness in AustroTox. We publish the data and code.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Context-Aware Content Moderation for German Newspaper Comments

    cs.CL 2025-05 conditional novelty 4.0 of 10

    LSTM and CNN models for German newspaper comment moderation improve when given article title and user history, while ChatGPT-3.5 zero-shot classification does not.

Pith tools