Pith. sign in

Improving LIME Robustness with Smarter Locality Sampling

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Explainability algorithms such as LIME have enabled machine learning systems to adopt transparency and fairness, which are important qualities in commercial use cases. However, recent work has shown that LIME's naive sampling strategy can be exploited by an adversary to conceal biased, harmful behavior. We propose to make LIME more robust by training a generative adversarial network to sample more realistic synthetic data which the explainer uses to generate explanations. Our experiments demonstrate that our proposed method demonstrates an increase in accuracy across three real-world datasets in detecting biased, adversarial behavior compared to vanilla LIME. This is achieved while maintaining comparable explanation quality, with up to 99.94\% in top-1 accuracy in some cases.

citation-role summary

background 1

citation-polarity summary

fields

math.OC 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Coherent Local Explanations for Mathematical Optimization

math.OC · 2025-02-07 · conditional · novelty 6.0

CLEMO fits local linear explanations for optimization models, adding a regularizer so predicted objective values match the objective of predicted decisions and predicted decisions stay feasible.

citing papers explorer

Showing 1 of 1 citing paper.

  • Coherent Local Explanations for Mathematical Optimization math.OC · 2025-02-07 · conditional · none · ref 31 · internal anchor

    CLEMO fits local linear explanations for optimization models, adding a regularizer so predicted objective values match the objective of predicted decisions and predicted decisions stay feasible.