Pith. sign in

REVIEW 2 major objections 4 minor 49 references

Diagnostic accuracy alone cannot reveal whether a medical LLM used case evidence appropriately; a role-aware behavioral audit exposes failures concentrated in negated and local findings.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 09:08 UTC pith:CTZASW2C

load-bearing objection A careful, well-scoped behavioral audit of evidence use in medical LLMs; the core method is standard interaction mining, but the diagnosis-relative interpretation layer and the clinical review are genuinely useful, with one load-bearing weakness: the role annotations that drive the headline numbers are never validated. the 2 major comments →

arxiv 2607.20848 v1 pith:CTZASW2C submitted 2026-07-23 cs.AI

Auditing Evidence Use in Medical LLM Diagnosis

classification cs.AI
keywords medical LLM evaluationevidence-use auditdiagnostic accuracyinteraction miningdiagnosis-relative rolesnegated evidenceclinical validationcounterfactual analysis
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that evaluating medical LLMs by final diagnostic accuracy is insufficient, because a model can select the right diagnosis while weighting evidence in clinically incoherent ways. To test this, it introduces a behavioral audit that decomposes each case into evidence units, scores candidate diagnoses under controlled subsets, and mines low-order interactions in diagnostic margins. The audit separates interaction discovery from failure assignment, using diagnosis-relative evidence roles, stability checks, and blinded clinical review. Across five LLMs and three datasets, faithful support and differential conflict/cancellation account for most interaction strength; however, all adjudicated invalid or shortcut-like mechanisms in the DDXPlus sample concentrate in negated or absent evidence and clinically local fields. The paper concludes that accuracy hides a measurable class of evidence-use failures and that role-aware audits should accompany accuracy reporting.

Core claim

The central claim is that accuracy-based evaluation cannot show whether a medical LLM used the case evidence appropriately, and the paper establishes a role-aware audit protocol that can. The audit's key empirical result is that across models and datasets, most mined evidence interactions are clinically plausible: faithful support plus differential conflict/cancellation accounts for 86.6–89.6% of DDXPlus interaction strength. But after stability filtering and blinded clinical review, the adjudicated invalid or shortcut-like cases are not distributed evenly; they all fall in the negated/absent evidence bucket. The paper interprets this as evidence that accuracy masks a specific class of failu

What carries the argument

The central mechanism is the low-order interaction effect I_d(S) computed from diagnostic margins under controlled evidence subsets, using the Möbius inversion formula over subsets. This interaction score quantifies how a combination of evidence units changes the margin for a diagnosis beyond their individual effects. The audit interprets these interactions through diagnosis-relative evidence roles (supporting, competitor-supporting, negated/absent, background, local), then applies prompt-, order-, and perturbation-stability filtering and blinded clinical review to separate plausible differential-diagnosis effects from candidate failures. The role-aware interpretation carries the argument.

Load-bearing premise

The audit's conclusions rest on the correctness of the diagnosis-relative evidence role annotations, which are derived from DDXPlus's structured per-diagnosis lists and were not sensitivity-tested against role misannotation; if these roles are wrong, the mechanism shares and the concentration of invalid cases in the negated/absent bucket could be artifacts.

What would settle it

Re-run the audit with independently re-annotated diagnosis-relative roles and with a review rubric that does not mention negation or local fields; if the adjudicated invalid cases no longer concentrate in the negated/absent bucket, the paper's central finding is a product of its annotation and rubric choices. A complementary automated check: flip the polarity of the negated evidence in the 11 invalid cases and verify that the model's margin for the target diagnosis moves in the expected direction; if it does not, the 'misuse' label is not behaviorally grounded.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If the paper is right, medical LLM evaluations should report evidence-use structure (e.g., mechanism-strength shares) alongside accuracy, since the two are not interchangeable.
  • The audit protocol can be applied to other diagnostic datasets and models; the expectation is that most interaction strength will reflect faithful support and conflict/cancellation, while stable failures will concentrate in negated/absent and local evidence.
  • Stability filtering (prompt wording, candidate order, neutralization) is a necessary precondition for failure claims; unfiltered mined interactions over-flag, and filtering raises adjudicated precision from 0.55 to 0.80 on the enriched queue.
  • The concentration of invalid mechanisms in negation polarity suggests targeted countermeasures: models may benefit from training or prompting that is robust to absent findings, and from de-biasing local fields with no clinical bridge.
  • Order-3 interaction auditing surfaces adjudicated invalid mechanisms that singleton and pairwise baselines miss (11/11 vs 0/11 and 4/11), motivating interaction-order coverage in evidence-use audits.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the pattern generalizes, error-analysis for medical LLMs should stratify by evidence polarity and clinical specificity, not just by final-answer correctness; this could reveal systematic negation-handling weaknesses across many clinical tasks.
  • A direct testable extension: construct a synthetic set of cases where negated findings are flipped to positive and vice versa; if the audit's claim is right, the model's margin changes should be disproportionately large in the originally-negated interactions.
  • The audit's separation of local incoherence from global reliance offers a template for interpreting attribution results in other high-stakes domains: a feature can be 'irrelevant' in a local interaction without being the cause of the final prediction.
  • Since the failure taxonomy depends on diagnosis-relative role annotations, datasets with structured roles are valuable; for narrative datasets, an LLM-based role annotator could be used, but its error rate would need to be checked because role misannotation could shift the failure buckets.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper proposes a behavioral audit of evidence use in medical LLM diagnosis. For each clinical case, patient information is decomposed into evidence units, candidate diagnoses are scored under controlled subsets, and low-order margin interactions are mined using a Möbius-style interaction index. Because medical evidence is diagnosis-relative, the audit separates discovery of interactions from assignment of failures: large or negative interactions may reflect plausible differential diagnosis, while suspicious interactions are subjected to stability checks (prompt wording, option order, deletion vs. neutralization) and blinded clinical review. The main empirical claims are that (1) across five open-weight LLMs and three datasets, faithful support plus differential conflict/cancellation accounts for 86.6–89.6% of DDXPlus interaction strength, so most evidence interactions are clinically plausible; and (2) in an enriched 130-item blinded five-reviewer DDXPlus review, all 11 adjudicated invalid or shortcut-like cases fall in the negated/absent misuse bucket, with zero invalid cases in the faithful, conflict, or background buckets (Table 5). The paper also reports stability filtering that raises adjudicated precision from 0.55 to 0.80, order-ablation results, and targeted counterfactual checks.

Significance. If the results hold, the paper makes a useful methodological contribution: it provides a concrete, role-aware framework for auditing evidence use rather than only final-answer accuracy, and it shows that accuracy can hide a measurable class of evidence-use failures concentrated in negation polarity and clinically local fields. The strengths are real: the five-model/three-dataset evaluation, the blinded five-reviewer adjudication with high inter-rater agreement (Fleiss' κ=0.82), the deletion-vs-neutralization controls, the stability filtering, and an unusually candid Limitations section. The central quantitative claims, however, rest on an unvalidated automated role-annotation layer and on a review rubric that substantially reproduces the automatic taxonomy. These are load-bearing gaps that prevent the empirical conclusions from being fully established.

major comments (2)
  1. [§2.2, Appendix B.2; Tables 4, 5, 14, 16] The diagnosis-relative role labels are the linchpin of the two central quantitative claims: the mechanism shares in Table 4/Table 14 (from which the 86.6–89.6% 'faithful+conflict' dominance is computed) and the bucket assignments behind the 11/11 invalid-case concentration in Table 5. The automatic role assignment is derived from DDXPlus structured symptom/antecedent lists, but the paper provides no inter-annotator agreement, gold-standard comparison, or perturbation analysis for these labels. The selection-sensitivity analysis in Table 16 varies which evidence units are audited but holds the role definitions fixed. The Limitations section itself states that results depend on role annotations; given that admission, a perturbation of role labels (e.g., reclassifying a share of competitor-supporting cues as background, or negated cues as low-specificity) is needed to show that the 86.6–89.
  2. [§3.3, Appendix F.2; Tables 5, 13, 20] The clinical review rubric defines 'invalid or shortcut-like' as 'an absent finding increasing support in the wrong direction or a clinically local field affecting an unrelated diagnosis' — almost verbatim the same conditions used by the automatic pre-flagging rules in §2.4 and Table 13. Reviewers were blinded to the provisional bucket, but the rubric itself encodes the taxonomy, so the concentration of invalid cases in the negated/absent bucket in Table 5 is partly circular. The independent counterfactual check is only a partial corrective: for no-travel, strong-change rates are 11.6% vs. 5.6% (Table 20), and for skin/rash the 'plausible skin' control is also highly sensitive (45.0%). To support the claim that the concentration reflects model behavior rather than the label scheme, the authors should either (a) re-run the review with a rubric that makes no reference to polarity or specif
minor comments (4)
  1. [Table 5 caption and abstract] The 130-item review sample is enriched for high-strength interactions and candidate-failure buckets, as stated in §3.3 and Appendix F.1. However, Table 5's 'Total' row and the abstract's phrasing ('invalid or shortcut-like cases concentrate in negated or absent findings') omit this enrichment caveat in the main text. Add the qualifier directly in the table caption and in Finding 3.
  2. [Table 3] The caption should more prominently state that the order-3 row is a reference set, not an independent sufficiency result; the text does say this, but the table alone can mislead readers into interpreting the 0/11 and 4/11 rows as independent retrieval failures.
  3. [Table 4] The columns 'Count' and 'Strength' are both described as percentages, but their sum in the table does not reach 100% for Count while Strength sums to approximately 99%. Clarify how the percentages are normalized and why Count and Strength differ.
  4. [Appendix F.2] The review rubric examples use the same terms as the automatic labels ('absent finding', 'clinically local field'). Consider rewording the rubric to use neutral, descriptive language so that reviewers are less likely to reproduce the taxonomy verbatim.

Circularity Check

1 steps flagged

The invalid-mechanism concentration is substantially self-definitional: the review rubric's 'invalid or shortcut-like' restates the automatic negated/local misuse labels, and the sample enriches those buckets.

specific steps
  1. self definitional [Section 2.4; Table 13; Section 3.3/Table 5; Appendix F.2]
    "The pipeline first checks candidate-misuse conditions: (i) positive interactions containing explicitly negated or excluding evidence, then (ii) interactions containing clinically local evidence with no role-supported path to the audited diagnosis. ... An invalid or shortcut-like mechanism reflects an implausible evidence use, such as an absent finding increasing support in the wrong direction or a clinically local field affecting an unrelated diagnosis."

    Table 13 defines 'negated or absent evidence misuse' as 'an explicitly absent or excluding finding increases support in an unexpected direction,' and 'irrelevant evidence misuse' as a clinically local field strongly affecting the audited diagnosis. Appendix F.2 defines the adjudicated outcome 'invalid or shortcut-like' with the same two clauses. Because Section 2.4's priority rules assign any positive interaction containing negated or local evidence to the misuse bucket before faithful/conflict labels, the buckets are mutually exclusive by construction. The review sample is enriched for the negated bucket (20/130 reviewed vs ~2.2% of interaction strength), and the rubric tells reviewers to call exactly that pattern invalid. Table 5's 11 invalid cases all in the negated bucket and zero in c

full rationale

The paper's main derivation chain is not a fit-to-input tautology: interactions are computed from option-scored margins (Eqs. 1-2), role labels come from DDXPlus structured finding lists, and the dominance of faithful+conflict/cancellation (86.6-89.6%) is an empirical distribution that is tested across alternative evidence-unit selectors (Table 16) and holds there. I do not flag that result as circular. The circularity is concentrated in the clinical-validation finding. The review rubric's definition of 'invalid or shortcut-like' is nearly verbatim the automatic candidate-failure labels in Table 13, and the priority rules in Section 2.4 guarantee that such interactions are bucketed as negated/absent or irrelevant/local before any other label. Combined with the explicit enrichment of the negated bucket in the 130-item sample (Table 18), the headline result that all 11 adjudicated invalid cases fall in the negated/absent bucket is partly by construction. The blinded reviewers and the concrete case examples provide real evidence that some patterns are clinically incoherent, so this is partial circularity rather than a fully empty result. There are no load-bearing self-citations or imported uniqueness claims. The counterfactual and robustness checks are independent and appropriately cautious, but they do not rescue the definitional overlap in the invalid-concentration finding.

Axiom & Free-Parameter Ledger

5 free parameters · 5 axioms · 1 invented entities

The audit contributes a measurement protocol, not a derivation, so free parameters are the hand-chosen thresholds, order, unit counts, and sampling enrichments that shape every reported percentage; the axioms are the role-annotation and adjudication premises the failure claims rest on. The absence of role-misannotation tests and the rubric-taxonomy overlap are the largest un-audited dependencies.

free parameters (5)
  • Salience threshold |I_y(S)| >= 0.5 = 0.5
    Hand-chosen cutoff (Appendix G.1) for which interactions count as salient; all mechanism-share percentages and review sampling are computed over interactions passing this threshold.
  • Maximum interaction order |S| <= 3 = 3
    Hand-chosen audit scope (Table 3). The paper's own ablation shows singleton (0/11) and pairwise (4/11) orders miss most adjudicated invalid items, so the failure findings depend on mining third-order interactions.
  • Audited unit count per case = 8 (DDXPlus), 6 (CupCase/MedCase)
    Tractable subset-intervention space (Table 1); choice determines which interactions exist and the 256-subset budget for DDXPlus.
  • Clinical review enrichment proportions = 30/40/20/20/20 = 130 items
    Table 18. Enriched for high-strength and candidate-failure buckets; all 11 invalid cases come from the 20-item negated bucket, so invalid rates within buckets are sampling artifacts, not population estimates.
  • Balanced evidence-unit selector = main setting
    Design choice for which 8 units are audited (Appendix B.2); the paper's own sensitivity shows top-interaction overlap only ~0.5 across selectors, so exact interaction findings are selector-dependent even though mechanism-level structure is stable.
axioms (5)
  • domain assumption DDXPlus structured annotations (per-diagnosis symptom vs antecedent lists) correctly encode diagnosis-relative clinical relevance.
    Invoked in Appendix B.2 for role assignment; the entire mechanism distribution (Finding 2) is computed by bucketing interactions under these roles, and role misannotation is not sensitivity-tested.
  • domain assumption Blinded clinical coherence judgment by five PhD-trained medical reviewers is a valid ground truth for evidence-use failure.
    Invoked in §3.3; the reviewers' rubric defines invalidity partly in terms of the same categories the pipeline pre-flags, so judgment independence is partial.
  • domain assumption The margin value function v_d(T) over option-letter log probabilities is a stable scalar that supports interaction decomposition.
    §2.1; token-prior and position effects are acknowledged and mitigated by candidate-order permutations, but margins are never calibrated.
  • standard math Interaction index (Eq. 2), the discrete Möbius inversion, is a valid and meaningful decomposition of joint evidence effects on the margin.
    §2.3; standard cooperative-game interaction theory (Grabisch-Roubens, Shapley-Taylor), assumed to apply to a non-additive, non-monotone LLM margin function.
  • domain assumption Hidden deleted evidence and 'not specified' neutralization correctly model evidence absence without asserting presence.
    §2.1, Appendix G.3; the deletion-vs-neutralization comparison shows signal reduction but not elimination, so some residual interaction strength is attributed to genuine evidence semantics rather than deletion artifacts.
invented entities (1)
  • Diagnosis-relative evidence role taxonomy and mechanism label set (faithful support, conflict/cancellation, redundancy, background/history, negated/absent misuse, irrelevant misuse) no independent evidence
    purpose: Maps mined interaction sign+strength+roles to clinical interpretations and filters candidate failures (§2.2, §2.4, Table 13).
    An analytic construct, not a physical postulate. Its validity is supported only inside the paper (blinded review, counterfactuals, selector sensitivity); no external dataset or prediction outside the paper falsifies or confirms it, and the review rubric that partially validates it was written to mirror it.

pith-pipeline@v1.3.0-alltime-deepseek · 20194 in / 22867 out tokens · 226829 ms · 2026-08-01T09:08:59.314147+00:00 · methodology

0 comments
read the original abstract

Medical LLMs are often evaluated by whether they select the correct diagnosis, but diagnostic accuracy alone does not show whether the model used the case evidence appropriately. We present a behavioral audit of evidence use in medical diagnosis. For each case, we decompose patient information into evidence units, score candidate diagnoses under controlled evidence subsets, and mine low-order interactions in diagnostic margins. Because medical evidence is diagnosis-relative, the audit separates interaction discovery from failure assignment: large or negative interactions can reflect plausible differential diagnosis, while suspicious interactions require robustness checks and clinical review. We evaluate five open-weight LLMs on DDXPlus, CupCase, and MedCase. Across datasets, faithful support and differential conflict or cancellation account for most interaction strength, showing that many evidence interactions are clinically plausible rather than failures. In a DDXPlus-focused blinded five-reviewer 130-item enriched review sample, invalid or shortcut-like cases concentrate in negated or absent findings and clinically local evidence. These results show that accuracy can hide candidate evidence-use failures and motivate role-aware audits for medical LLM evaluation.

Figures

Figures reproduced from arXiv: 2607.20848 by Fuji Ren, Jiawen Deng, Junchi Liao.

Figure 1
Figure 1. Figure 1: Overview of the evidence-use audit. Cases are decomposed into evidence units, scored under controlled [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 3
Figure 3. Figure 3: Clinical validation. (a) Adjudicated review [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Robustness and filtering. (a) Stability filter [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Case-study summary. Each row contrasts the evidence pattern, mined interaction, targeted counterfactual [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

49 extracted references · 10 linked inside Pith

  1. [1]

    Nature , volume=

    Large language models encode clinical knowledge , author=. Nature , volume=. 2023 , publisher=

  2. [2]

    Nature medicine , volume=

    Toward expert-level medical question answering with large language models , author=. Nature medicine , volume=. 2025 , publisher=

  3. [3]

    PLoS digital health , volume=

    Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models , author=. PLoS digital health , volume=. 2023 , publisher=

  4. [4]

    Applied Sciences , volume=

    What disease does this patient have? a large-scale open domain question answering dataset from medical exams , author=. Applied Sciences , volume=. 2021 , publisher=

  5. [5]

    Conference on health, inference, and learning , pages=

    Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering , author=. Conference on health, inference, and learning , pages=. 2022 , organization=

  6. [6]

    Advances in neural information processing systems , volume=

    Ddxplus: A new dataset for automatic medical diagnosis , author=. Advances in neural information processing systems , volume=

  7. [7]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume=

    Cupcase: Clinically uncommon patient cases and diagnoses dataset , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=

  8. [8]

    arXiv preprint arXiv:2505.11733 , year=

    Medcasereasoning: Evaluating and learning diagnostic reasoning from clinical case reports , author=. arXiv preprint arXiv:2505.11733 , year=

  9. [9]

    5-coder technical report , author=

    Qwen2. 5-coder technical report , author=. arXiv preprint arXiv:2409.12186 , year=

  10. [10]

    Hugging Face repository , howpublished =

    Ankit Pal, Malaikannan Sankarasubbu , title =. Hugging Face repository , howpublished =. 2024 , publisher =

  11. [11]

    arXiv preprint arXiv:2604.05081 , year=

    MedGemma 1.5 Technical Report , author=. arXiv preprint arXiv:2604.05081 , year=

  12. [12]

    Findings of the association for computational linguistics: acl 2024 , pages=

    Biomistral: A collection of open-source pretrained large language models for medical domains , author=. Findings of the association for computational linguistics: acl 2024 , pages=

  13. [13]

    arXiv preprint arXiv:2408.06142 , year=

    Med42-v2: A suite of clinical llms , author=. arXiv preprint arXiv:2408.06142 , year=

  14. [14]

    Why should i trust you?

    " Why should i trust you?" Explaining the predictions of any classifier , author=. Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining , pages=

  15. [15]

    International conference on machine learning , pages=

    Axiomatic attribution for deep networks , author=. International conference on machine learning , pages=. 2017 , organization=

  16. [16]

    Advances in neural information processing systems , volume=

    A unified approach to interpreting model predictions , author=. Advances in neural information processing systems , volume=

  17. [17]

    International conference on machine learning , pages=

    The shapley taylor interaction index , author=. International conference on machine learning , pages=. 2020 , organization=

  18. [18]

    Proceedings of the 58th annual meeting of the association for computational linguistics , pages=

    ERASER: A benchmark to evaluate rationalized NLP models , author=. Proceedings of the 58th annual meeting of the association for computational linguistics , pages=

  19. [19]

    Proceedings of the 58th annual meeting of the association for computational linguistics , pages=

    Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness? , author=. Proceedings of the 58th annual meeting of the association for computational linguistics , pages=

  20. [20]

    arXiv preprint arXiv:1909.12434 , year=

    Learning the difference that makes a difference with counterfactually-augmented data , author=. arXiv preprint arXiv:1909.12434 , year=

  21. [21]

    Findings of the Association for Computational Linguistics: EMNLP 2020 , pages=

    Evaluating models’ local decision boundaries via contrast sets , author=. Findings of the Association for Computational Linguistics: EMNLP 2020 , pages=

  22. [22]

    Proceedings of the 58th annual meeting of the association for computational linguistics , pages=

    Beyond accuracy: Behavioral testing of NLP models with CheckList , author=. Proceedings of the 58th annual meeting of the association for computational linguistics , pages=

  23. [23]

    Proceedings of the conference on fairness, accountability, and transparency , pages=

    Model cards for model reporting , author=. Proceedings of the conference on fairness, accountability, and transparency , pages=

  24. [24]

    Nature , volume=

    Foundation models for generalist medical artificial intelligence , author=. Nature , volume=. 2023 , publisher=

  25. [25]

    Nature medicine , volume=

    Large language models in medicine , author=. Nature medicine , volume=. 2023 , publisher=

  26. [26]

    Communications medicine , volume=

    The future landscape of large language models in medicine , author=. Communications medicine , volume=. 2023 , publisher=

  27. [27]

    arXiv preprint arXiv:2303.13375 , year=

    Capabilities of gpt-4 on medical challenge problems , author=. arXiv preprint arXiv:2303.13375 , year=

  28. [28]

    arXiv preprint arXiv:2009.03300 , year=

    Measuring massive multitask language understanding , author=. arXiv preprint arXiv:2009.03300 , year=

  29. [29]

    Transactions on machine learning research , year=

    Beyond the imitation game: Quantifying and extrapolating the capabilities of language models , author=. Transactions on machine learning research , year=

  30. [30]

    arXiv preprint arXiv:2211.09110 , year=

    Holistic evaluation of language models , author=. arXiv preprint arXiv:2211.09110 , year=

  31. [31]

    Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=

    Truthfulqa: Measuring how models mimic human falsehoods , author=. Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers) , pages=

  32. [32]

    Pubmedqa: A dataset for biomedical research question answering , author=. Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP) , pages=

  33. [33]

    Scientific data , volume=

    MIMIC-III, a freely accessible critical care database , author=. Scientific data , volume=. 2016 , publisher=

  34. [34]

    Proceedings of the 2018 conference on empirical methods in natural language processing , pages=

    emrqa: A large corpus for question answering on electronic medical records , author=. Proceedings of the 2018 conference on empirical methods in natural language processing , pages=

  35. [35]

    Advances in Neural Information Processing Systems , volume=

    Language models don't always say what they think: Unfaithful explanations in chain-of-thought prompting , author=. Advances in Neural Information Processing Systems , volume=

  36. [36]

    arXiv preprint arXiv:2307.13702 , year=

    Measuring faithfulness in chain-of-thought reasoning , author=. arXiv preprint arXiv:2307.13702 , year=

  37. [37]

    International Journal of game theory , volume=

    An axiomatic approach to the concept of interaction among players in cooperative games , author=. International Journal of game theory , volume=. 1999 , publisher=

  38. [38]

    Journal of Machine Learning Research , volume=

    Explaining explanations: Axiomatic feature interactions for deep networks , author=. Journal of Machine Learning Research , volume=

  39. [39]

    arXiv preprint arXiv:1705.04977 , year=

    Detecting statistical interactions from neural network weights , author=. arXiv preprint arXiv:1705.04977 , year=

  40. [40]

    Transactions of the Association for Computational Linguistics , volume=

    Data statements for natural language processing: Toward mitigating system bias and enabling better science , author=. Transactions of the Association for Computational Linguistics , volume=. 2018 , publisher=

  41. [41]

    Communications of the ACM , volume=

    Datasheets for datasets , author=. Communications of the ACM , volume=. 2021 , publisher=

  42. [42]

    Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=

    On the dangers of stochastic parrots: Can language models be too big? , author=. Proceedings of the 2021 ACM conference on fairness, accountability, and transparency , pages=

  43. [43]

    Science , volume=

    Dissecting racial bias in an algorithm used to manage the health of populations , author=. Science , volume=. 2019 , publisher=

  44. [44]

    Proceedings of the AAAI conference on artificial intelligence , volume=

    Anchors: High-precision model-agnostic explanations , author=. Proceedings of the AAAI conference on artificial intelligence , volume=

  45. [45]

    International conference on machine learning , pages=

    Understanding black-box predictions via influence functions , author=. International conference on machine learning , pages=. 2017 , organization=

  46. [46]

    International conference on machine learning , pages=

    Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav) , author=. International conference on machine learning , pages=. 2018 , organization=

  47. [47]

    Attention is not explanation , author=. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , pages=

  48. [48]

    Proceedings of the 57th annual meeting of the association for computational linguistics , pages=

    Is attention interpretable? , author=. Proceedings of the 57th annual meeting of the association for computational linguistics , pages=

  49. [49]

    Advances in neural information processing systems , volume=

    Sanity checks for saliency maps , author=. Advances in neural information processing systems , volume=