Pith. sign in

REVIEW 2 cited by

How Complex is your classification problem? A survey on measuring classification complexity

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1808.03591 v3 pith:ABTD2VSH submitted 2018-08-10 cs.LG stat.ML

classification cs.LGstat.ML
keywords classificationcomplexitymeasuresproblemscharacteristicsdatadatasetsextracted
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Characteristics extracted from the training datasets of classification problems have proven to be effective predictors in a number of meta-analyses. Among them, measures of classification complexity can be used to estimate the difficulty in separating the data points into their expected classes. Descriptors of the spatial distribution of the data and estimates of the shape and size of the decision boundary are among the known measures for this characterization. This information can support the formulation of new data-driven pre-processing and pattern recognition techniques, which can in turn be focused on challenges highlighted by such characteristics of the problems. This paper surveys and analyzes measures which can be extracted from the training datasets in order to characterize the complexity of the respective classification problems. Their use in recent literature is also reviewed and discussed, allowing to prospect opportunities for future work in the area. Finally, descriptions are given on an R package named Extended Complexity Library (ECoL) that implements a set of complexity measures and is made publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Is data-efficient learning feasible with quantum models?

    quant-ph 2025-08 conditional novelty 5.0 of 10

    Quantum kernels can beat an untuned classical kernel on datasets whose labels the authors deliberately construct from the quantum kernel's own spectrum, showing data efficiency by construction.

  2. On the Effect of Ruleset Tuning and Data Imbalance on Explainable Network Security Alert Classifications: a Case-Study on DeepCASE

    cs.CR 2025-07 conditional novelty 5.0 of 10

    Ruleset tuning that reduces label imbalance improves DeepCASE's alert classification and the correctness of its explanations in a real SOC dataset.

Pith tools