REVIEW 7 cited by
Squared Earth Mover's Distance-based Loss for Training Deep Neural Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In the context of single-label classification, despite the huge success of deep learning, the commonly used cross-entropy loss function ignores the intricate inter-class relationships that often exist in real-life tasks such as age classification. In this work, we propose to leverage these relationships between classes by training deep nets with the exact squared Earth Mover's Distance (also known as Wasserstein distance) for single-label classification. The squared EMD loss uses the predicted probabilities of all classes and penalizes the miss-predictions according to a ground distance matrix that quantifies the dissimilarities between classes. We demonstrate that on datasets with strong inter-class relationships such as an ordering between classes, our exact squared EMD losses yield new state-of-the-art results. Furthermore, we propose a method to automatically learn this matrix using the CNN's own features during training. We show that our method can learn a ground distance matrix efficiently with no inter-class relationship priors and yield the same performance gain. Finally, we show that our method can be generalized to applications that lack strong inter-class relationships and still maintain state-of-the-art performance. Therefore, with limited computational overhead, one can always deploy the proposed loss function on any dataset over the conventional cross-entropy.
Forward citations
Cited by 7 Pith papers
-
Evaluating and Pricing Advertisements in AI-Generated Responses
A persona-agent simulation generates click-intent labels for ads inside AI answers, a distilled evaluator reproduces those labels and beats zero-shot LLM judges on directional tests, and the same score drives a truthf...
-
Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection
A Schwartz-aware energy decoder improves theory-coherent label sets on 19 refined values at no F1 cost, while training-time geometry and LLM prompting do not match it.
-
Reliable Conformal Prediction for Ordinal Classification Using the Ranked Probability Score
RPS-based conformal prediction for ordinal classification yields median-centered contiguous sets with a favorable width-miscoverage tradeoff compared to prior methods.
-
Differentiable Halo Mass Prediction and the Cosmology-Dependence of Halo Mass Functions
A differentiable U-Net predicts halo mass functions and their cosmology derivatives from initial density fields, matching finite-difference gradients of simulations and emulators to within model scatter.
-
Conveyance: A Versatile Framework for Learning in Structured Class Spaces
Conveyance is a margin-based loss for structured class spaces that encodes graph relations without joint distributions and matches specialized baselines on hierarchical, ordinal, and multiple-instance tasks.
-
Aleatoric and Epistemic Uncertainty Measures for Ordinal Classification through Binary Reduction
Order-consistent binary reduction, summing entropy or variance uncertainties over all ordered splits, provides competitive aleatoric and epistemic uncertainty measures for ordinal classification.
-
Explaining Automatic Image Assessment
Training separate NIMA-style models on depth, saliency, and blur versions of AVA shows saliency carries the most signal among non-RGB modalities, while the standard 5.0 threshold inflates baselines to above 70 percent.
Discussion (0). Continue with ORCID to comment.