Pith. sign in

REVIEW 29 cited by

Metrics for Multi-Class Classification: an Overview

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2008.05756 v1 pith:BIIRNOPI submitted 2020-08-13 stat.ML cs.LG

classification stat.MLcs.LG
keywords classificationdifferentmetricsmulti-classdevelopmentlearningmachinemodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Classification tasks in machine learning involving more than two classes are known by the name of "multi-class classification". Performance indicators are very useful when the aim is to evaluate and compare different classification models or machine learning techniques. Many metrics come in handy to test the ability of a multi-class classifier. Those metrics turn out to be useful at different stage of the development process, e.g. comparing the performance of two different models or analysing the behaviour of the same model by tuning different parameters. In this white paper we review a list of the most promising multi-class metrics, we highlight their advantages and disadvantages and show their possible usages during the development of a classification model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 29 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 66 citations worldwide. Full citation record

  1. Apnea Burden-Guided Framework: Enhancing Out-of-Distribution Generalization in PPG-Based Sleep Apnea Characterization

    eess.SP 2026-08 conditional novelty 6.0 of 10

    A PPG-plus-SpO2 deep learning framework that predicts apnea burden and converts it to AHI improved out-of-distribution severity classification versus SpO2 alone in a 30-subject external cohort.

  2. TransformEEG: Towards Improving Model Generalizability in Deep Learning-based EEG Parkinson's Disease Detection

    cs.LG 2025-07 conditional novelty 6.0 of 10

    TransformEEG, a convolutional-transformer with a depthwise tokenizer, reports the highest median balanced accuracy (78.4-80.1%) and lowest variability across splits among eight EEG models for Parkinson's detection on ...

  3. Rhythm Features for Speaker Identification

    eess.AS 2025-06 conditional novelty 6.0 of 10

    Using character-duration sequences from WhisperX, a transformer identifies speakers with balanced accuracy of 0.39 on LibriSpeech but only 0.03 on VoxCeleb1, and fusion with x-vectors does not improve accuracy.

  4. SneakPeek: Data-Aware Model Selection and Scheduling for Inference Serving on the Edge

    cs.DC 2025-05 conditional novelty 6.0 of 10

    SneakPeek estimates each model's accuracy from live data with a fast kNN and Bayesian update, then uses those estimates in grouped scheduling and short-circuit inference to raise utility on a single GPU.

  5. From Newswire to Nexus: Using text-based actor embeddings and transformer networks to forecast conflict dynamics

    cs.CY 2025-01 reject novelty 6.0 of 10

    Transformer models fine-tuned on dyad-level newswire digests forecast conflict escalation better than a fatality-history baseline at one-month horizons, but lose skill at six months.

  6. Multi-label Classification using Deep Multi-order Context-aware Kernel Networks

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Proposes DMCKN, a deep kernel network that aggregates multi-order spatial context via attention and random walks, showing modest gains on two multi-label benchmarks.

  7. Transferring self-supervised pre-trained models for SHM data anomaly detection with scarce labeled data

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Autoencoder pre-training on unlabeled bridge monitoring data, followed by fine-tuning on a few hundred labels, raises anomaly detection F1 by 3 to 9 points over supervised training in two real bridge datasets.

  8. Exploring LLM Capabilities in Extracting DCAT-Compatible Metadata for Data Cataloging

    cs.IR 2025-07 conditional novelty 5.0 of 10

    Large language models can generate DCAT-compatible metadata for data catalogs at quality close to human annotations, though the strongest evidence is for simple extraction tasks.

  9. Comparing Credit Risk Estimates in the Gen-AI Era

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Few-shot GPT-4o underperforms logistic regression and KNN on German Credit Data across all tested prompt and example-selection configurations.

  10. Runtime Analysis of Evolutionary NAS for Multiclass Classification

    cs.NE 2025-06 conditional novelty 5.0 of 10

    On a hand-built multiclass benchmark, (1+1)-ENAS with one-bit or bit-wise mutation finds an optimal architecture in O(rM ln(rM)) expected generations, with lower bound Omega(rM ln M), so the two mutations have nearly ...

  11. Evaluation in EEG Emotion Recognition: State-of-the-Art Review and Unified Framework

    eess.SP 2025-05 conditional novelty 5.0 of 10

    EEG emotion recognition research is built on inconsistent evaluation protocols, and the new EEGain framework provides a unified benchmark across six datasets and four models.

  12. LLM-Assisted Automated Deductive Coding of Dialogue Data: Leveraging Dialogue-Specific Characteristics to Enhance Contextual Understanding

    cs.CL 2025-04 conditional novelty 5.0 of 10

    An LLM-assisted pipeline that separates communicative acts from events, uses multi-model voting, and applies a consistency check reaches Cohen's kappa above 0.80 with human coders on a small student-dialogue corpus.

  13. Rule-Based Modeling of Low-Dimensional Data with PCA and Binary Particle Swarm Optimization (BPSO) in ANFIS

    cs.CV 2025-02 conditional novelty 5.0 of 10

    ANFIS-PCA-BPSO applies PCA to normalized firing strengths and selects components with BPSO, reducing ANFIS rule count and training time at a small accuracy cost.

  14. A Transferable Physics-Informed Framework for Battery Degradation Diagnosis, Knee-Onset Detection and Knee Prediction

    eess.SY 2025-01 conditional novelty 5.0 of 10

    A transferable framework using histogram features, a physics-inspired neural network, and XGBoost estimates battery degradation modes, detects degradation phases, and aims to predict capacity knees online.

  15. Adapting OpenAI's CLIP Model for Few-Shot Image Inspection in Manufacturing Quality Control: An Expository Case Study with Multiple Application Examples

    cs.CV 2025-01 conditional novelty 5.0 of 10

    Frozen CLIP embeddings plus a nearest-example cosine rule give a strong few-shot baseline for simple and textured manufacturing inspections, but not for complex multi-component scenes.

  16. Rare Event Detection in Imbalanced Multi-Class Datasets Using an Optimal MIP-Based Ensemble Weighting Approach

    cs.LG 2024-12 conditional novelty 5.0 of 10

    A MIP-based ensemble weighting scheme jointly selects K classifiers and assigns per-class weights; it reports average balanced accuracy gains of 4.53% over six baselines on four imbalanced CPS datasets.

  17. Machine learning-based classification for Single Photon Space Debris Light Curves

    astro-ph.IM 2024-11 conditional novelty 5.0 of 10

    Machine learning, especially feature-based Random Forest and XGBoost, can classify single-photon space debris light curves with accuracies up to about 90.7 percent.

  18. Towards Continuous-variable Quantum Neural Networks for Biomedical Imaging

    quant-ph 2025-11 conditional novelty 4.0 of 10

    A 4-qumode Gaussian CV-QNN classifies MedMNIST images with accuracy statistically indistinguishable from a 42-parameter classical linear model and a DV-QNN.

  19. Proactive HIV Care: AI-Based Comorbidity Prediction from Routine EHR Data

    cs.CY 2025-08 conditional novelty 4.0 of 10

    Across six models and 2,200 HIV outpatients, including demographic features always improved multi-label comorbidity prediction, with XGBoost best at 45.8% macro F1; gender was recoverable from labs at 92.8%.

  20. A Systematic Literature Review on Multi-label Data Stream Classification

    cs.LG 2025-08 conditional novelty 4.0 of 10

    A systematic review of 58 multi-label data stream classifiers finds that concept drift is addressed by about half the papers, while label latency and concept evolution are each handled by only a small minority, and re...

  21. Ensemble BERT for Medication Event Classification on Electronic Health Records (EHRs)

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Majority voting over multiple pretrained BERT models improves medication event classification on the n2c2 CMED dataset, but the paper lacks error bars and code.

  22. An Explorative Analysis of SVM Classifier and ResNet50 Architecture on African Food Classification

    cs.CV 2025-05 conditional novelty 4.0 of 10

    On 1,658 images of six African foods, a fine-tuned ResNet50 and an SVM using raw pixels both reach about 81 percent accuracy, with SVM slightly ahead on macro F1.

  23. Dusty stellar sources classification by implementing machine learning methods based on spectroscopic observations in the Magellanic Clouds

    astro-ph.GA 2025-04 conditional novelty 4.0 of 10

    A probabilistic random forest trained on 618 spectroscopically confirmed dusty stars achieves 89% accuracy and relabels more than 23,000 sources through a consensus of four models.

  24. Image Classification with Deep Reinforcement Active Learning

    cs.CV 2024-12 conditional novelty 4.0 of 10

    An active learning method that uses deep reinforcement learning (DDPG) to decide which unlabeled images to query, after pre-ranking images by margin uncertainty, reports modest accuracy improvements on CIFAR-10, SVHN,...

  25. Performance evaluation of predictive AI models to support medical decisions: Overview and guidance

    cs.LG 2024-12 accept novelty 3.0 of 10

    The paper recommends validating binary medical prediction models with AUROC, a calibration plot, net benefit with decision curves, and risk distribution plots, and warns against improper classification metrics.

  26. Enhancing Supply Chain Visibility with Generative AI: An Exploratory Case Study on Relationship Prediction in Knowledge Graphs

    cs.CE 2024-12 reject novelty 3.0 of 10

    Pretrained language model embeddings improve supply chain link prediction in knowledge graphs, but the reported near-perfect accuracy likely reflects entity memorization rather than true generalization.

  27. LIMBA: An Open-Source Framework for the Preservation and Valorization of Low-Resource Languages using Generative Models

    cs.CL 2024-11 conditional novelty 3.0 of 10

    LIMBA is a proposed pipeline that combines collection, grammatical tagging, translation, speech, and generative modules to build language models for low-resource languages, with preliminary Sardinian experiments.

  28. Classification of Disease from Lungs X-ray Images using VGG16, VGG19 and ResNet50 Models

    cs.CV 2026-07 reject novelty 2.0 of 10

    Fine-tuned VGG16, VGG19 and ResNet-50v2 classify a public Kaggle chest X-ray set (COVID-19, normal, viral pneumonia) at 85–96% accuracy; the abstract claims tuberculosis and lung-cancer coverage the experiments never include.

  29. Multidimensional classification of posts for online course discussion forum curation

    cs.CL 2025-08 reject novelty 2.0 of 10

    Bayesian fusion of a generic LLM and a local classifier ties the best individual classifier on MOOC forum labels and lags fine-tuning, undermining the paper's headline claim.

Pith tools