Pith. sign in

REVIEW 1 cited by

Characteristics of Harmful Text: Towards Rigorous Benchmarking of Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.08325 v2 pith:EZTJQGKN submitted 2022-06-16 cs.CL cs.AIcs.CY

classification cs.CLcs.AIcs.CY
keywords benchmarkslanguagecharacteristicsharmfulmodelstextforesightharms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models produce human-like text that drive a growing number of applications. However, recent literature and, increasingly, real world observations, have demonstrated that these models can generate language that is toxic, biased, untruthful or otherwise harmful. Though work to evaluate language model harms is under way, translating foresight about which harms may arise into rigorous benchmarks is not straightforward. To facilitate this translation, we outline six ways of characterizing harmful text which merit explicit consideration when designing new benchmarks. We then use these characteristics as a lens to identify trends and gaps in existing benchmarks. Finally, we apply them in a case study of the Perspective API, a toxicity classifier that is widely used in harm benchmarks. Our characteristics provide one piece of the bridge that translates between foresight and effective evaluation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 12 citations worldwide. Full citation record

  1. The Societal Impact of Foundation Models: Advancing Evidence-based AI Policy

    cs.AI 2025-06 conditional novelty 2.0 of 10

    A dissertation that synthesizes prior work on foundation models into a three-part framework: conceptual framing, empirical measurement (HELM, FMTI), and evidence-based AI policy.

Pith tools