Pith. sign in

REVIEW 10 cited by

True to the Model or True to the Data?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.16234 v1 pith:777735IN submitted 2020-06-29 cs.LG stat.ML

classification cs.LGstat.ML
keywords choicedatatrueconditionalexpectationmodelobservationalshapley
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A variety of recent papers discuss the application of Shapley values, a concept for explaining coalitional games, for feature attribution in machine learning. However, the correct way to connect a machine learning model to a coalitional game has been a source of controversy. The two main approaches that have been proposed differ in the way that they condition on known features, using either (1) an interventional or (2) an observational conditional expectation. While previous work has argued that one of the two approaches is preferable in general, we argue that the choice is application dependent. Furthermore, we argue that the choice comes down to whether it is desirable to be true to the model or true to the data. We use linear models to investigate this choice. After deriving an efficient method for calculating observational conditional expectation Shapley values for linear models, we investigate how correlation in simulated data impacts the convergence of observational conditional expectation Shapley values. Finally, we present two real data examples that we consider to be representative of possible use cases for feature attribution -- (1) credit risk modeling and (2) biological discovery. We show how a different choice of value function performs better in each scenario, and how possible attributions are impacted by modeling choices.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. QuadraSHAP: Stable and Scalable Shapley Values for Product Games via Gauss-Legendre Quadrature

    cs.LG 2026-05 conditional novelty 7.0 of 10

    Shapley values in product games equal the integral of a degree-(d-1) polynomial over [0,1], allowing provably exact or near-exact computation via Gauss-Legendre quadrature with O(d m_q) work.

  2. QuadraSHAP: Stable and Scalable Shapley Values for Product Games via Gauss-Legendre Quadrature

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    Shapley values in product games equal an exact one-dimensional integral of a polynomial, computable via Gauss-Legendre quadrature with linear cost in the number of features.

  3. Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    In high-stakes settings, Shapley explanations increase analyst confidence but do not improve decision accuracy, and standard metrics fail to predict human utility.

  4. From Decision Trees to Boolean Logic: A Fast and Unified SHAP Algorithm

    cs.LG 2025-11 unverdicted novelty 6.0 of 10

    WOODELF computes Background SHAP for tree ensembles in linear time via pseudo-Boolean formulas that encode trees, features, and background data, with reported speedups of 16x on CPU and 165x on GPU for million-row datasets.

  5. Identifying the post-pandemic determinants of low performing students in Latin America through Interpretable Machine Learning methods

    econ.GN 2025-09 reject novelty 6.0 of 10

    Using stacked ML models and SHAP values on PISA 2022, the paper ranks the correlates of bottom and low performance in 10 Latin American countries, with grade repetition, family SES, and ICT access as the most consiste...

  6. On Spectral Properties of Gradient-based Explanation Methods

    cs.LG 2025-08 conditional novelty 6.0 of 10

    Gradient-based explanations behave like frequency-band selectors: the gradient acts as a high-pass filter, perturbation as a low-pass filter, and their combination creates explanations that shift with the perturbation scale.

  7. ELATE: Evolutionary Language model for Automated Time-series Engineering

    cs.LG 2025-08 conditional novelty 5.0 of 10

    An LLM-guided evolutionary feature engineering method for time-series forecasting reduces RMSE by 8.4% on average across seven datasets.

  8. A case for data valuation transparency via DValCards

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Data valuation is unstable across imputation methods and can penalize minority groups; the paper proposes DValCards to document and constrain such valuation use.

  9. Analysing drivers and interdependencies in European electricity markets using XAI

    cs.AI 2026-06 unverdicted novelty 4.0 of 10

    DNNs plus SHAP/SSHAP applied to 39 European bidding zones identify solar and gas as key price drivers and simulate a single-price EU market.

  10. On the Complexity-Faithfulness Trade-off of Gradient-Based Explanations

    cs.LG 2025-08 reject novelty 4.0 of 10

    The paper introduces EF and ΔEF as spectral metrics, but ΔEF is derived from EF, making the complexity-faithfulness trade-off partly tautological.

Pith tools