Pith. sign in

REVIEW 7 cited by

Squared Earth Mover's Distance-based Loss for Training Deep Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1611.05916 v4 pith:YGR4BVVZ submitted 2016-11-17 cs.CV

classification cs.CV
keywords classesdistanceinter-classlossrelationshipssquaredclassificationdeep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the context of single-label classification, despite the huge success of deep learning, the commonly used cross-entropy loss function ignores the intricate inter-class relationships that often exist in real-life tasks such as age classification. In this work, we propose to leverage these relationships between classes by training deep nets with the exact squared Earth Mover's Distance (also known as Wasserstein distance) for single-label classification. The squared EMD loss uses the predicted probabilities of all classes and penalizes the miss-predictions according to a ground distance matrix that quantifies the dissimilarities between classes. We demonstrate that on datasets with strong inter-class relationships such as an ordering between classes, our exact squared EMD losses yield new state-of-the-art results. Furthermore, we propose a method to automatically learn this matrix using the CNN's own features during training. We show that our method can learn a ground distance matrix efficiently with no inter-class relationship priors and yield the same performance gain. Finally, we show that our method can be generalized to applications that lack strong inter-class relationships and still maintain state-of-the-art performance. Therefore, with limited computational overhead, one can always deploy the proposed loss function on any dataset over the conventional cross-entropy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating and Pricing Advertisements in AI-Generated Responses

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A persona-agent simulation generates click-intent labels for ads inside AI answers, a distilled evaluator reproduces those labels and beats zero-shot LLM judges on directional tests, and the same score drives a truthf...

  2. Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A Schwartz-aware energy decoder improves theory-coherent label sets on 19 refined values at no F1 cost, while training-time geometry and LLM prompting do not match it.

  3. Reliable Conformal Prediction for Ordinal Classification Using the Ranked Probability Score

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    RPS-based conformal prediction for ordinal classification yields median-centered contiguous sets with a favorable width-miscoverage tradeoff compared to prior methods.

  4. Differentiable Halo Mass Prediction and the Cosmology-Dependence of Halo Mass Functions

    astro-ph.CO 2025-07 conditional novelty 6.0 of 10

    A differentiable U-Net predicts halo mass functions and their cosmology derivatives from initial density fields, matching finite-difference gradients of simulations and emulators to within model scatter.

  5. Conveyance: A Versatile Framework for Learning in Structured Class Spaces

    cs.LG 2026-05 unverdicted novelty 5.0 of 10

    Conveyance is a margin-based loss for structured class spaces that encodes graph relations without joint distributions and matches specialized baselines on hierarchical, ordinal, and multiple-instance tasks.

  6. Aleatoric and Epistemic Uncertainty Measures for Ordinal Classification through Binary Reduction

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Order-consistent binary reduction, summing entropy or variance uncertainties over all ordered splits, provides competitive aleatoric and epistemic uncertainty measures for ordinal classification.

  7. Explaining Automatic Image Assessment

    cs.CV 2025-02 reject novelty 4.0 of 10

    Training separate NIMA-style models on depth, saliency, and blur versions of AVA shows saliency carries the most signal among non-RGB modalities, while the standard 5.0 threshold inflates baselines to above 70 percent.

Pith tools