Pith. sign in

REVIEW 10 cited by

Metrics for Multi-Class Classification: an Overview

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.05756 v1 pith:BIIRNOPI submitted 2020-08-13 stat.ML cs.LG

Metrics for Multi-Class Classification: an Overview

classification stat.ML cs.LG
keywords classificationdifferentmetricsmulti-classdevelopmentlearningmachinemodel
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Classification tasks in machine learning involving more than two classes are known by the name of "multi-class classification". Performance indicators are very useful when the aim is to evaluate and compare different classification models or machine learning techniques. Many metrics come in handy to test the ability of a multi-class classifier. Those metrics turn out to be useful at different stage of the development process, e.g. comparing the performance of two different models or analysing the behaviour of the same model by tuning different parameters. In this white paper we review a list of the most promising multi-class metrics, we highlight their advantages and disadvantages and show their possible usages during the development of a classification model.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning

    cs.CL 2026-05 unverdicted novelty 7.0

    UA-Legal-Bench is a new five-task benchmark for Ukrainian legal reasoning that demonstrates task-dependent few-shot prompting effects and the need for macro-F1 over accuracy on imbalanced classes.

  2. Online Set Learning from Precision and Recall Feedback

    cs.LG 2026-05 unverdicted novelty 7.0

    A hypothesis class is learnable in this online precision-recall feedback model if and only if it has finite VC dimension, with algorithms achieving regret bounds in realizable and agnostic settings despite ERM failing.

  3. Revisiting Privacy Leakage in Machine Unlearning: Membership Inference Beyond the Forgotten Set

    cs.CR 2026-05 unverdicted novelty 7.0

    TC-UMIA is a population-level attack using pre- and post-unlearning predictions to infer membership across forget, retain, and unseen sets, revealing added privacy leakage to retained data.

  4. Revisiting Privacy Leakage in Machine Unlearning: Membership Inference Beyond the Forgotten Set

    cs.CR 2026-05 unverdicted novelty 7.0

    Unlearning increases privacy leakage for the retain set, and a new tri-class membership inference attack distinguishes forget, retain, and unseen data using pre- and post-unlearning model outputs.

  5. HADS-Net:A Hybrid Attention-Augmented Dual-Stream Network with Physics-Informed Augmentation for Breast Ultrasound Image Classification

    cs.CV 2026-05 unverdicted novelty 5.0

    HADS-Net fuses physics-informed texture features from EfficientNet-B3 with Sobel boundary features via cross-attention to reach 96.58% accuracy and 0.9978 macro ROC-AUC on the BUSI breast ultrasound dataset.

  6. Solving Constrained Affine Heaviside Composite Optimization Problems by a Progressive IP Approach

    math.OC 2026-05 unverdicted novelty 5.0

    A progressive IP method with successive decomposition and approximation solves constrained affine Heaviside composite optimization problems, with proven convergence to local optima and numerical support from classific...

  7. Annotation Quality in Aspect-Based Sentiment Analysis: A Case Study Comparing Experts, Students, Crowdworkers, and Large Language Model

    cs.CL 2026-05 unverdicted novelty 5.0

    Expert re-annotations of a German ABSA dataset serve as ground truth to evaluate how students, crowdworkers, and LLMs affect inter-annotator agreement and downstream performance on ACSA and TASD tasks using BERT, T5, ...

  8. A Comprehensive Survey of Agents for Computer Use: Foundations, Challenges, and Future Directions

    cs.AI 2025-01 unverdicted novelty 5.0

    A survey of 87 agents for computer use and 33 datasets that introduces a three-dimensional taxonomy across domain, interaction, and agent perspectives and identifies six research gaps.

  9. Towards Continuous-variable Quantum Neural Networks for Biomedical Imaging

    quant-ph 2025-11 conditional novelty 4.0

    A 4-qumode Gaussian CV-QNN classifies MedMNIST images with accuracy statistically indistinguishable from a 42-parameter classical linear model and a DV-QNN.

  10. Classification of Disease from Lungs X-ray Images using VGG16, VGG19 and ResNet50 Models

    cs.CV 2026-07 reject novelty 2.0

    Fine-tuned VGG16, VGG19 and ResNet-50v2 classify a public Kaggle chest X-ray set (COVID-19, normal, viral pneumonia) at 85–96% accuracy; the abstract claims tuberculosis and lung-cancer coverage the experiments never include.