Pith. sign in

REVIEW 28 cited by

Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.16061 v1 pith:ESXXDISC submitted 2020-10-11 cs.LG stat.MEstat.ML

Evaluation: from precision, recall and F-measure to ROC, informedness, markedness and correlation

classification cs.LG stat.MEstat.ML
keywords informednessmeasurescasechancemarkednessprecisionrecallused
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Commonly used evaluation measures including Recall, Precision, F-Measure and Rand Accuracy are biased and should not be used without clear understanding of the biases, and corresponding identification of chance or base case levels of the statistic. Using these measures a system that performs worse in the objective sense of Informedness, can appear to perform better under any of these commonly used measures. We discuss several concepts and measures that reflect the probability that prediction is informed versus chance. Informedness and introduce Markedness as a dual measure for the probability that prediction is marked versus chance. Finally we demonstrate elegant connections between the concepts of Informedness, Markedness, Correlation and Significance as well as their intuitive relationships with Recall and Precision, and outline the extension from the dichotomous case to the general multi-class case.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Tail-Aware Adaptive-k: Query-Adaptive Context Selection for Retrieval-Augmented Generation

    cs.IR 2026-06 unverdicted novelty 7.0

    TAA-k finds query-adaptive retrieval cutoffs by first using knee detection to isolate a candidate window around the relevance-to-noise transition, then applying EVT goodness-of-fit tests inside that window.

  2. Satisfiability Solving with LLMs: A Matched-Pair Evaluation of Reasoning Capability

    cs.AI 2026-05 unverdicted novelty 7.0

    A matched-pair protocol and Accurate Differentiation Rate metric reveal that conventional LLM accuracy on SAT problems is often inflated by over-predicting satisfiability, while cross-representation agreement exceeds ...

  3. Everywhere Valid Bounds on False Discovery Proportions in Conformal Inference

    stat.ME 2026-05 unverdicted novelty 7.0

    Derives simultaneous high-probability upper bounds on realized FDP for conformal p-values that hold for arbitrary post-hoc thresholds via an envelope on the null empirical distribution function.

  4. Time series causal discovery with variable lags

    cs.LG 2026-04 unverdicted novelty 7.0

    A Tabu-based algorithm learns time-ordered causal graphs from time series by optimizing per-edge lags with a decomposable BIC score and explicit lag penalty.

  5. Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization

    cs.CR 2026-04 unverdicted novelty 7.0

    HyPE detects harmful prompts as outliers in hyperbolic space and HyPS sanitizes them using explainable attribution, outperforming prior defenses in accuracy and robustness across datasets and adversarial scenarios.

  6. Direction for Detection: A Survey of Automated Vulnerability Detection and all of its Pain Points

    cs.SE 2024-12 conditional novelty 7.0

    ML4AVD research remains locked into binary function-level classification of C/C++ vulnerabilities because twelve pain points in the pipeline reinforce each other through feedback loops.

  7. PERSONAJUDGE: Simulating Individual Human Preference Judgments with Evaluator-Specific Demonstration Data

    cs.HC 2026-07 conditional novelty 6.5

    Evaluator-specific demonstrations with retrospective reasoning improve LLM simulation of individual preference judges by up to 9.9 points over a non-personalized base judge, while interface telemetry often degrades accuracy.

  8. Rethinking Clinical Relevance in Chest X-ray Machine Learning: How Evaluation References Define Performance

    eess.IV 2026-07 conditional novelty 6.0

    Chest X-ray AI model rankings and image-quality metric rankings change substantially with the choice of evaluation reference, so benchmark scores are not neutral.

  9. Everywhere Valid Bounds on False Discovery Proportions in Conformal Inference

    stat.ME 2026-05 unverdicted novelty 6.0

    Derives simultaneous finite-sample distribution-free upper bounds on false discovery proportions for conformal p-values that hold for every possible rejection threshold.

  10. AMAR: Lightweight Attention-Based Multi-User Activity Recognition from Wi-Fi CSI

    eess.SP 2026-05 unverdicted novelty 6.0

    AMAR uses a transformer with learnable query embeddings for set-based prediction of concurrent activities from composite Wi-Fi CSI, combined with edge feature extraction and vector quantization for bandwidth-efficient...

  11. Retrieval with Multiple Query Vectors through Anomalous Pattern Detection

    cs.LG 2026-05 unverdicted novelty 6.0

    A retrieval approach identifies anomalous dimensions in a set of query vectors and retrieves database vectors that are anomalous across those dimensions, with performance improving as query set size grows to around 8.

  12. HFS-TriNet: A Three-Branch Collaborative Feature Learning Network for Prostate Cancer Classification from TRUS Videos

    cs.CV 2026-04 unverdicted novelty 6.0

    HFS-TriNet applies heuristic frame selection and a three-branch network (ResNet50, SAM-based with temporal attention, WTCR) to classify prostate cancer from TRUS videos.

  13. Beyond Statistical Co-occurrence: Unlocking Intrinsic Semantics for Tabular Data Clustering

    cs.AI 2026-04 unverdicted novelty 6.0

    TagCC anchors statistical tabular representations to LLM-derived textual semantic concepts via contrastive learning jointly optimized with a clustering objective, outperforming prior methods on benchmarks.

  14. Event Detection in Videos: A Framework for the Development of New Methods

    cs.CV 2026-07 conditional novelty 5.5

    A framework of tagged multi-environment datasets (including new FSD and SUC), probabilistic Tile-based ranking, and explicit application scenarios for fair video event detection.

  15. From Unsupervised Subgroups to Hypothetical State-Intervention Policies: An Evaluation of Selected Subgrouping Methods in Observational Health Data

    cs.LG 2026-07 conditional novelty 5.0

    Phenotype-first unsupervised subgroups yield comparable held-out policy utilities across clustering methods, with no statistically significant differences, while the individuals prioritized differ substantially.

  16. Lung-R1: A Knowledge Graph-Guided LLM for Pulmonary Diagnostic Reasoning

    cs.AI 2026-06 unverdicted novelty 5.0

    Introduces the first structured pulmonary knowledge graph LungKG and uses it to train Lung-R1, which reaches SOTA on EMR-based pulmonary diagnosis tasks.

  17. Retrieval with Multiple Query Vectors through Anomalous Pattern Detection

    cs.LG 2026-05 unverdicted novelty 5.0

    Multi-query vector retrieval via anomalous pattern detection on shared dimensions improves retrieval as query-set size grows, especially from 1 to 8 queries.

  18. Interpretable facial dynamics as behavioral and perceptual traces of deepfakes

    cs.CV 2026-04 unverdicted novelty 5.0

    Face-swapped deepfakes exhibit measurable higher-order temporal irregularities in facial dynamics that are most pronounced during emotional expressions, enabling above-chance ML detection and partial alignment with hu...

  19. Time series causal discovery with variable lags

    cs.LG 2026-04 unverdicted novelty 5.0

    A Tabu-based algorithm learns time-ordered causal graphs from time series with variable per-edge lags using a decomposable BIC score and explicit lag penalty.

  20. From Large Language Model Predicates to Logic Tensor Networks: Neurosymbolic Offer Validation in Regulated Procurement

    cs.AI 2026-04 conditional novelty 5.0

    LLM-scored offer predicates aggregated by a Logic Tensor Network classify procurement documents about as accurately as BERT or LLM baselines while exposing auditable predicate and rule truth values.

  21. From Large Language Model Predicates to Logic Tensor Networks: Neurosymbolic Offer Validation in Regulated Procurement

    cs.AI 2026-04 unverdicted novelty 5.0

    A neurosymbolic pipeline extracts predicates from offer texts with an LLM and validates them via Logic Tensor Networks, delivering performance comparable to standard models plus built-in interpretability on a real corpus.

  22. MAC: Masked Agent Collaboration Boosts Large Language Model Medical Decision-Making

    cs.AI 2025-07 unverdicted novelty 5.0

    MAC framework selects Pareto-optimal LLM agents and masks low cross-consistency outputs for adaptive collaboration in medical decision-making.

  23. A Deep Multiscale Neural Network for Accurate Neurological Disorder Detection from MRI Scans and Real-Time Web Deployment

    cs.CV 2026-06 conditional novelty 4.5

    End-Net, a multiscale inception-based CNN, reaches 0.9761 accuracy on a balanced multi-class MRI dataset of Alzheimer, tumors, MS and controls and is deployed as a public web service.

  24. Which Anatomy Matters Under Limited Labels? A Data-Efficient Anatomy-Aware Benchmark for Cardiac Pathology Prediction

    eess.IV 2026-05 unverdicted novelty 4.0

    Anatomy-aware descriptors from RV, myocardium and LV outperform model complexity in low-label 5-class cardiac pathology prediction on the ACDC MRI dataset.

  25. Automated Detection of Urological Events in Bladder Pressure Signals with a Two-Stage Machine Learning Framework Validated on External Datasets

    eess.SP 2026-05 conditional novelty 4.0

    A two-stage MLP detects urological events in vesical pressure signals with 84% accuracy for voiding versus non-voiding and 90% for abdominal versus detrusor overactivity on external validation data.

  26. CPEMH: An Agentic Framework for Prompt-Driven Behavior Evaluation and Assurance in Foundation-Model Systems for Mental Health Screening

    cs.AI 2026-05 unverdicted novelty 4.0

    CPEMH is a new agentic framework that orchestrates AI agents to design, evaluate, and select prompts for stable and traceable behavior in foundation models applied to mental health screening from transcripts.

  27. Beyond Semantics: An Evidential Reasoning-Aware Multi-View Learning Framework for Trustworthy Mental Health Prediction

    cs.CL 2026-05 unverdicted novelty 4.0

    A multi-view evidential framework combines semantic and reasoning information to improve accuracy and provide trustworthy uncertainty estimates for mental health prediction on text data.

  28. A Deep Multiscale Neural Network for Accurate Neurological Disorder Detection from MRI Scans and Real-Time Web Deployment

    cs.CV 2026-06 unverdicted novelty 3.0

    End-Net, a multiscale CNN with inception modules, claims superior accuracy on four-class neurological disorder MRI classification and includes online deployment.