Pith. sign in

REVIEW 3 cited by

Evaluation of Similarity-based Explanations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.04528 v2 pith:Q4XWL6WI submitted 2020-06-08 cs.LG stat.ML

classification cs.LGstat.ML
keywords metricsrelevanceexplanationssimilarity-basedexplanationpredictionstestsusers
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Explaining the predictions made by complex machine learning models helps users to understand and accept the predicted outputs with confidence. One promising way is to use similarity-based explanation that provides similar instances as evidence to support model predictions. Several relevance metrics are used for this purpose. In this study, we investigated relevance metrics that can provide reasonable explanations to users. Specifically, we adopted three tests to evaluate whether the relevance metrics satisfy the minimal requirements for similarity-based explanation. Our experiments revealed that the cosine similarity of the gradients of the loss performs best, which would be a recommended choice in practice. In addition, we showed that some metrics perform poorly in our tests and analyzed the reasons of their failure. We expect our insights to help practitioners in selecting appropriate relevance metrics and also aid further researches for designing better relevance metrics for explanations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ProDS: Preference-oriented Data Selection for Instruction Tuning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    ProDS picks instruction-tuning data by matching training-sample gradients to preference gradients from DPO, achieving slight gains over prior selection methods on MMLU, TYDIQA, BBH, and Alpaca-style tests.

  2. Position: The Most Expensive Part of an LLM should be its Training Data

    cs.CL 2025-04 conditional novelty 6.0 of 10

    Even at conservative wages, recreating LLM training data from scratch would cost 10 to 1000 times more than the compute and energy used to train the models.

  3. Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities

    cs.CL 2025-01 conditional novelty 6.0 of 10

    BIDS, a balanced influence-based data selection algorithm using per-task normalization and iterative greedy selection, improves balanced multi-capability instruction tuning and can outperform full-dataset training on ...

Pith tools