Pith. sign in

REVIEW 1 cited by

Bad practices in evaluation methodology relevant to class-imbalanced problems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1812.01388 v1 pith:JMV566SP submitted 2018-12-04 cs.LG stat.ML

classification cs.LGstat.ML
keywords evaluationalgorithmsclassfieldmetricpracticesproblemsable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

For research to go in the right direction, it is essential to be able to compare and quantify performance of different algorithms focused on the same problem. Choosing a suitable evaluation metric requires deep understanding of the pursued task along with all of its characteristics. We argue that in the case of applied machine learning, proper evaluation metric is the basic building block that should be in the spotlight and put under thorough examination. Here, we address tasks with class imbalance, in which the class of interest is the one with much lower number of samples. We encountered non-insignificant amount of recent papers, in which improper evaluation methods are used, borrowed mainly from the field of balanced problems. Such bad practices may heavily bias the results in favour of inappropriate algorithms and give false expectations of the state of the field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Temporal cross-validation impacts multivariate time series subsequence anomaly detection evaluation

    stat.ML 2025-06 conditional novelty 4.0 of 10

    Sliding-window cross-validation yields higher and more stable AUC-PR than walk-forward for subsequence anomaly detection on two small CAN fault datasets.

Pith tools