Pith. sign in

REVIEW 1 cited by

Assessing agreement on classification tasks: the kappa statistic

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv cmp-lg/9602004 v1 pith:KD2PNKJF submitted 1996-02-27 cmp-lg cs.CL

classification cmp-lgcs.CL
keywords analysisarguecognitivecomputationalcontentcurrentlydialoguediscourse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Currently, computational linguists and cognitive scientists working in the area of discourse and dialogue argue that their subjective judgments are reliable using several different statistics, none of which are easily interpretable or comparable to each other. Meanwhile, researchers in content analysis have already experienced the same difficulties and come up with a solution in the kappa statistic. We discuss what is wrong with reliability measures as they are currently used for discourse and dialogue work in computational linguistics and cognitive science, and argue that we would be better off as a field adopting techniques from content analysis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Automating Credit Card Limit Adjustments Using Machine Learning

    cs.LG 2025-01 conditional novelty 4.0 of 10

    An XGBoost model with cost-sensitive learning reports Cohen's kappa of 0.81 against a bank committee's credit card limit decisions and is proposed to automate the process.

Pith tools