Pith. sign in

REVIEW 2 cited by

Macro F1 and Macro F1

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.03347 v3 pith:VJEL7DZZ submitted 2019-11-08 cs.LG stat.ML

classification cs.LGstat.ML
keywords computationsmacrodifferentonlybinarycalculatecircumstancesclassification
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The 'macro F1' metric is frequently used to evaluate binary, multi-class and multi-label classification problems. Yet, we find that there exist two different formulas to calculate this quantity. In this note, we show that only under rare circumstances the two computations can be considered equivalent. More specifically, one formula well 'rewards' classifiers which produce a skewed error type distribution. In fact, the difference in outcome of the two computations can be as high as 0.5. The two computations may not only diverge in their scalar result but can also lead to different classifier rankings.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 24 citations worldwide. Full citation record

  1. Dependency Triad: A Metric to Quantify the Dependencies Between Attributes for Local Differential Privacy

    cs.CR 2026-08 conditional novelty 6.0 of 10

    The Dependency Triad summarizes pairwise attribute dependence with three parameters and delivers a constant-time upper-bound estimate of correlation-induced privacy leakage.

  2. Joint Modeling of Entities and Discourse Relations for Coherence Assessment

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Jointly modeling entities and discourse relations improves coherence assessment accuracy over text-only and single-feature models on GCDC, CoheSentia, and TOEFL.

Pith tools