Pith. sign in

REVIEW 3 cited by

Analysis and Comparison of Classification Metrics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.05355 v4 pith:QYN6A6SC submitted 2022-09-12 cs.LG cs.AI

classification cs.LGcs.AI
keywords metricserrorqualitybalancedcalibrationclassificationexpectedlearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

A variety of different performance metrics are commonly used in the machine learning literature for the evaluation of classification systems. Some of the most common ones for measuring quality of hard decisions are standard and balanced accuracy, standard and balanced error rate, F-beta score, and Matthews correlation coefficient (MCC). In this document, we review the definition of these and other metrics and compare them with the expected cost (EC), a metric introduced in every statistical learning course but rarely used in the machine learning literature. We show that both the standard and balanced error rates are special cases of the EC. Further, we show its relation with F-beta score and MCC and argue that EC is superior to these traditional metrics for being based on first principles from statistics, and for being more general, interpretable, and adaptable to any application scenario. The metrics mentioned above measure the quality of hard decisions. Yet, most modern classification systems output continuous scores for the classes which we may want to evaluate directly. Metrics for measuring the quality of system scores include the area under the ROC curve, equal error rate, cross-entropy, Brier score, and Bayes EC or Bayes risk, among others. The last three metrics are special cases of a family of metrics given by the expected value of proper scoring rules (PSRs). We review the theory behind these metrics, showing that they are a principled way to measure the quality of the posterior probabilities produced by a system. Finally, we show how to use these metrics to compute a system's calibration loss and compare this metric with the widely-used expected calibration error (ECE), arguing that calibration loss based on PSRs is superior to the ECE for being more interpretable, more general, and directly applicable to the multi-class case, among other reasons.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 18 citations worldwide. Full citation record

  1. Time-dependent density estimation using binary classifiers

    stat.ML 2025-06 conditional novelty 6.0 of 10

    A time-dependent classifier whose pre-activation approximates the partial time derivative of the log-density makes fully data-driven, path-independent evaluation of evolving probability densities possible.

  2. Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Supervised contrastive alignment of frozen WavLM layers modestly lifts Mandarin depression F1 under LOSO while quantifying that speaker leakage inflated prior Mandarin F1 by ~0.23.

  3. Evaluating AI capabilities in detecting conspiracy theories on YouTube

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Zero-shot text LLMs detect conspiracy YouTube videos with high recall but low precision, a fine-tuned RoBERTa remains competitive, and thumbnails add little value.

Pith tools