Pith. sign in

REVIEW 3 major objections 2 minor 1 cited by

The Case for Model Science: Verify, Explore, Steer, Refine

T0 review · 3 major / 2 minor · reviewed 2026-06-28 · grok-4.3

Pith's one-line read AI research should consolidate scattered model analysis into a new discipline called Model Science built on four perspectives: Verify, Explore, Steer, and Refine.

desk verdict Position paper proposes 'Model Science' with four perspectives but rests on untested analogies without adaptation details or supporting evidence. read the letter →

arxiv 2606.01189 v1 pith:LRQKUYV4 submitted 2026-05-31 cs.AI

classification cs.AI
keywords ModelScienceAIanalysisbenchmarklimitationsverificationexplorationsinglecasestudiesinfrastructureexplainable
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that benchmarks have driven progress but cannot explain why models succeed or fail and routinely miss critical issues such as hallucinations and shortcuts. It proposes moving to Model Science by adapting lessons from cognitive science on multiple levels of analysis, neuroscience on single-case depth, medicine on paired training and research, and agriculture on shared infrastructure. The resulting discipline rests on three foundations: the four functional perspectives that address complementary questions about model behaviour, catalogues of datasets models and findings for cumulative knowledge, and detailed study of individual model instances. A sympathetic reader would care because current methods leave deployed systems that serve billions of users poorly understood. The proposal treats these elements as ready to be assembled into systematic practice.

What carries the argument

The four functional perspectives Verify, Explore, Steer, and Refine that together address complementary questions about model behaviour and form one of the three foundations for Model Science.

What would settle it

A sustained effort to build the proposed catalogues and apply the four perspectives that produces no new explanations of model failures beyond existing benchmarks would undermine the case for Model Science.

Watch

Extended reading notes

Core claim

We argue that the AI community is now ready to move beyond benchmarking and consolidate scattered efforts in model analysis into a systematic discipline, a direction we term Model Science. Precedents from cognitive science, neuroscience, medicine, and agriculture show that complex systems require complementary levels of analysis, single-case depth, specialised training alongside research, and shared infrastructure. These lessons support three foundations: consolidation around the four perspectives Verify, Explore, Steer, and Refine; catalogues of datasets, models, and findings; and deep analysis of individual model instances rather than only model families.

Load-bearing premise

That precedents and practices from cognitive science, neuroscience, medicine, and agriculture can be transferred directly to create a viable new discipline for AI models.

Editorial extensions

If this is right

  • Benchmarks will be supplemented by methods that identify why models succeed or fail rather than only measuring performance.
  • Shared catalogues of datasets, models, and findings will enable cumulative progress instead of repeated isolated studies.
  • Deep analysis of single model instances will reveal patterns that population-level studies across model families miss.
  • Specialised training in model analysis will develop in parallel with research practice.
  • Complementary levels of analysis will become standard for understanding complex model behaviours.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The four perspectives could provide a common language for integrating existing scattered tools for model inspection and control.
  • Infrastructure for Model Science might extend to regulatory requirements that demand evidence from Verify and Steer activities before large-scale deployment.
  • Single-instance analysis could become routine for high-stakes applications where aggregated metrics are known to overlook rare but severe failure modes.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The paper argues that the AI community should move beyond benchmark-driven research to establish a new systematic discipline termed 'Model Science.' It draws on precedents from cognitive science (complementary levels of analysis), neuroscience (value of single-case studies), medicine (specialized training alongside research), and agriculture (shared infrastructure for cumulative progress) to propose three foundations: (1) four functional perspectives—Verify, Explore, Steer, and Refine—for analyzing model behavior; (2) shared infrastructure including catalogues of datasets, models, and findings; and (3) deep analysis of individual model instances rather than only model families.

Significance. If the proposed framework gains traction, it could help organize scattered model analysis efforts and address benchmark limitations such as failure to explain why models succeed or fail on tasks like hallucination detection. The paper correctly identifies that current leaderboards track performance gains but provide limited insight into internal mechanisms. However, the significance is limited by the absence of any concrete mappings, pilot implementations, or falsifiable predictions demonstrating that the cited field analogies can be adapted to computational models without substantial modification.

major comments (3)
  1. [Abstract / Foundations section] Abstract and the section outlining the three foundations: the central readiness claim—that precedents 'point the way forward' and that the community is 'now ready' to consolidate into Model Science—rests on an untested transferability assumption. No specific argument is given showing why single-case neuroscience methods would reveal LLM internals differently from population benchmarks, why medicine-style training would scale to model analysis, or how agriculture-style catalogues would overcome rapid model obsolescence.
  2. [Precedents from neuroscience / single-instance analysis paragraph] Discussion of the neuroscience precedent for single-instance analysis: the manuscript states that 'single cases can reveal what population studies miss' but supplies no mapping or example demonstrating how this would apply to trained neural networks, where population-level benchmarks are the dominant evaluation paradigm due to the statistical nature of learned parameters.
  3. [Infrastructure discussion] Infrastructure foundation (catalogues of datasets, models, and findings): the proposal assumes such shared resources would enable cumulative progress, yet the text does not address or provide evidence against the risk that fast iteration cycles in AI would render catalogues obsolete faster than in agriculture, undermining the cumulative-knowledge goal.
minor comments (2)
  1. [Four perspectives section] The four perspectives (Verify, Explore, Steer, Refine) are introduced at a high level; concrete operational definitions or example workflows for each would improve clarity.
  2. [Precedents paragraphs] The manuscript would benefit from additional citations to specific methodological papers in the referenced fields (e.g., single-case studies in neuroscience) to ground the analogies.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for the constructive feedback on our position paper. We address each major comment below, clarifying the manuscript's scope as an argument for establishing Model Science rather than an empirical validation of the proposed analogies.

read point-by-point responses
  1. Referee: [Abstract / Foundations section] Abstract and the section outlining the three foundations: the central readiness claim—that precedents 'point the way forward' and that the community is 'now ready' to consolidate into Model Science—rests on an untested transferability assumption. No specific argument is given showing why single-case neuroscience methods would reveal LLM internals differently from population benchmarks, why medicine-style training would scale to model analysis, or how agriculture-style catalogues would overcome rapid model obsolescence.

    Authors: We agree that the paper presents the precedents as suggestive rather than demonstrating specific transferability arguments or mappings. As a position paper, the intent is to outline why the community should pursue such consolidation, with the task of adapting and testing these ideas left to future work in the proposed discipline. We will revise the abstract and foundations section to explicitly frame the analogies as hypotheses to be investigated rather than established transfers. revision: yes

  2. Referee: [Precedents from neuroscience / single-instance analysis paragraph] Discussion of the neuroscience precedent for single-instance analysis: the manuscript states that 'single cases can reveal what population studies miss' but supplies no mapping or example demonstrating how this would apply to trained neural networks, where population-level benchmarks are the dominant evaluation paradigm due to the statistical nature of learned parameters.

    Authors: The neuroscience reference is used to illustrate the potential value of single-instance analysis alongside population methods. We acknowledge the absence of a concrete mapping to neural networks. In revision, we will expand this paragraph with a short note on how methods such as circuit analysis on individual models could serve an analogous role to single-case studies, while recognizing that population benchmarks remain central due to the statistical nature of training. revision: partial

  3. Referee: [Infrastructure discussion] Infrastructure foundation (catalogues of datasets, models, and findings): the proposal assumes such shared resources would enable cumulative progress, yet the text does not address or provide evidence against the risk that fast iteration cycles in AI would render catalogues obsolete faster than in agriculture, undermining the cumulative-knowledge goal.

    Authors: The manuscript does not discuss the risk of rapid obsolescence in AI relative to slower-moving fields like agriculture. This is a substantive concern that merits direct engagement. We will add a dedicated paragraph to the infrastructure section acknowledging this challenge and outlining potential mitigations, such as maintaining versioned catalogues focused on general principles and failure modes rather than transient model specifics. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity; proposal is a conceptual argument resting on external analogies without self-referential reductions.

full rationale

The paper advances a disciplinary proposal by invoking precedents from cognitive science, neuroscience, medicine, and agriculture to motivate four perspectives, shared infrastructure, and single-instance analysis. No equations, fitted parameters, or 'predictions' appear that reduce to inputs by construction. No self-citation chains or uniqueness theorems are invoked to justify the framework. The load-bearing step is the transferability assumption itself, which is external and falsifiable rather than tautological. This is a normal non-finding for a position paper whose derivation chain is self-contained against external benchmarks.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No free parameters, mathematical axioms, or invented entities are introduced; the paper is a high-level conceptual argument relying on domain assumptions about the limitations of benchmarks and transferability of practices from other sciences.

assumptions (2)
  • domain assumption Benchmarks reveal whether models perform but not why they succeed or fail, and miss critical failure modes such as hallucinations or shortcuts.
    Stated in the abstract as the core motivation for moving beyond benchmarking.
  • domain assumption Lessons from cognitive science, neuroscience, medicine, and agriculture can inform the foundations of a new discipline for AI model analysis.
    Invoked to justify the three proposed foundations without further justification in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Case for Model Science: Verify, Explore, Steer, Refine." pith.science (2026). https://pith.science/paper/LRQKUYV4

@misc{pith2026260601189,
  author       = {Pith},
  title        = {Pith review of: The Case for Model Science: Verify, Explore, Steer, Refine},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LRQKUYV4}},
  note         = {Machine review of arXiv:2606.01189}
}
read the original abstract

We argue that the AI community is now ready to move beyond benchmarking and consolidate scattered efforts in model analysis into a systematic discipline, a direction we term Model Science. Complex AI models now serve billions of users, yet our understanding of how they work lags far behind our ability to deploy them. Decades of benchmark-driven research have delivered remarkable progress: extensive leaderboards, a wide range of performance metrics, tracking capability gains across diverse tasks; yet this success has also revealed the limits of benchmarks as they tell us whether models perform but not why they succeed or fail, they miss critical failure modes, such as hallucinations or shortcuts. Precedents from established sciences point the way forward: cognitive science shows that understanding complex systems requires complementary levels of analysis; neuroscience demonstrates that deep study of single cases reveals what population studies miss; medicine teaches that specialised training must develop alongside research practice; and agriculture models how shared infrastructure and principles enable cumulative progress. These lessons inform three foundations for Model Science. First, we propose to consolidate research around four functional perspectives: Verify, Explore, Steer, and Refine that address complementary questions about model behaviour. Second, we discuss the required infrastructure for cumulative knowledge: catalogues of datasets, models and findings. Third, we highlight the need for deep analysis of individual model instances, not just model families, because single cases can reveal what population studies miss.

Figures

Figures reproduced from arXiv: 2606.01189 by the authors.

Figure 1
Figure 1. An illustration of the foundations of the Model Science. The base is the infrastructure: data [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Interpretability-Guided Soft Pruning of Attention Heads in Vision Transformers

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A differentiable soft-pruning framework (SAPER) and a Laplacian-based head visualization reduce ViT attention-head counts up to 24-fold with 4-11 point accuracy drops on ImageNet-1K.

Reference graph

Works this paper leans on

283 extracted references · 69 canonical work pages · cited by 1 Pith paper

  1. [1]

    Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES) , pages =

    Slack, Dylan and Hilgard, Sophie and Jia, Emily and Singh, Sameer and Lakkaraju, Himabindu , title =. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society (AIES) , pages =. 2020 , doi =

  2. [2]

    and Applebaum, Andy and Miller, Doug P

    Strom, Blake E. and Applebaum, Andy and Miller, Doug P. and Nickels, Kathryn C. and Pennington, Adam G. and Thomas, Cody B. , title =. 2018 , url =

  3. [3]

    Information Fusion , volume =

    Woźnica, Katarzyna and Wilczyński, Piotr and Biecek, Przemysław , title =. Information Fusion , volume =. 2026 , doi =

  4. [4]

    Lindsey, Jack and Gurnee, Wes and Ameisen, Emmanuel and Chen, Brian and Pearce, Adam and Turner, Nicholas L. and Citro, Craig and Abrahams, David and Carter, Shan and Hosmer, Basil and Marcus, Jonathan and Sklar, Michael and Templeton, Adly and Bricken, Trenton and McDougall, Callum and Cunningham, Hoagy and Henighan, Thomas and Jermyn, Adam and Jones, An...

  5. [5]

    Ecosystem graphs: The social footprint of foundation models.arXiv preprint arXiv:2303.15772,

    Bommasani, Rishi and Soylu, Dilara and Liao, Thomas I. and Creel, Kathleen A. and Liang, Percy , title =. arXiv preprint arXiv:2303.15772 , year =

  6. [6]

    Nature , volume =

    Rahwan, Iyad and Cebrian, Manuel and Obradovich, Nick and others , title =. Nature , volume =. 2019 , doi =

  7. [7]

    Ullman and Fernando Martinez-Plumed and Joshua B

    Ryan Burnell and Wout Schellaert and John Burden and Tomer D. Ullman and Fernando Martinez-Plumed and Joshua B. Tenenbaum and Danaja Rutar and Lucy G. Cheke and Jascha Sohl-Dickstein and Melanie Mitchell and Douwe Kiela and Murray Shanahan and Ellen M. Voorhees and Anthony G. Cohn and Joel Z. Leibo and Jose Hernandez-Orallo , title =. Science , volume =. ...

  8. [8]

    Beyond the Leaderboard: A Survey of the Science of Evaluation, Benchmarking, and Methodologies for Large Language Models , rights=

    Sheikhi, Saeid and Loven, Lauri and Kostakos, Panos , year=. Beyond the Leaderboard: A Survey of the Science of Evaluation, Benchmarking, and Methodologies for Large Language Models , rights=. doi:10.1109/ACCESS.2026.3686088 , journal=

Show all 283 references
  1. [9]

    and Hanna, Alex and Paullada, Amandalynne , booktitle =

    Raji, Deborah and Denton, Emily and Bender, Emily M. and Hanna, Alex and Paullada, Amandalynne , booktitle =

  2. [10]

    2017 , institution =

    Kelly, Markelle and Longjohn, Rachel and Nottingham, Kolby , title =. 2017 , institution =

  3. [11]

    , title =

    Fisher, Ronald A. , title =. Annals of Eugenics , volume =. 1936 , doi =

  4. [12]

    Gradient-Based Learning Applied to Document Recognition , journal =

    LeCun, Yann and Bottou, L. Gradient-Based Learning Applied to Document Recognition , journal =. 1998 , doi =

  5. [13]

    and Santorini, Beatrice and Marcinkiewicz, Mary Ann , title =

    Marcus, Mitchell P. and Santorini, Beatrice and Marcinkiewicz, Mary Ann , title =. Computational Linguistics , volume =

  6. [14]

    Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages =

    Rajpurkar, Pranav and Zhang, Jian and Lopyrev, Konstantin and Liang, Percy , title =. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP) , pages =. 2016 , doi =

  7. [15]

    , title =

    Wang, Alex and Singh, Amanpreet and Michael, Julian and Hill, Felix and Levy, Omer and Bowman, Samuel R. , title =. Proceedings of the 2018 EMNLP Workshop BlackboxNLP , pages =. 2018 , doi =

  8. [16]

    , title =

    Wang, Alex and Pruksachatkun, Yada and Nangia, Nikita and Singh, Amanpreet and Michael, Julian and Hill, Felix and Levy, Omer and Bowman, Samuel R. , title =. Advances in Neural Information Processing Systems (NeurIPS) , volume =

  9. [17]

    and Harman, Donna K

    Voorhees, Ellen M. and Harman, Donna K. , title =. TREC: Experiment and Evaluation in Information Retrieval , publisher =

  10. [18]

    Proceedings of KDD Cup and Workshop , year =

    Bennett, James and Lanning, Stan , title =. Proceedings of KDD Cup and Workshop , year =

  11. [19]

    IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages =

    Deng, Jia and Dong, Wei and Socher, Richard and Li, Li-Jia and Li, Kai and Fei-Fei, Li , title =. IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , pages =. 2009 , doi =

  12. [20]

    and Fei-Fei, Li , title =

    Russakovsky, Olga and Deng, Jia and Su, Hao and Krause, Jonathan and Satheesh, Sanjeev and Ma, Sean and Huang, Zhiheng and Karpathy, Andrej and Khosla, Aditya and Bernstein, Michael and Berg, Alexander C. and Fei-Fei, Li , title =. International Journal of Computer Vision , vo...

  13. [21]

    , title =

    Krizhevsky, Alex and Sutskever, Ilya and Hinton, Geoffrey E. , title =. Advances in Neural Information Processing Systems (NeurIPS) , volume =

  14. [22]

    Big Data & Society , volume =

    Denton, Emily and Hanna, Alex and Amironesei, Razvan and Smart, Andrew and Nicole, Hilary , title =. Big Data & Society , volume =. 2021 , doi =

  15. [23]

    and Dahl, George , title =

    Bowman, Samuel R. and Dahl, George , title =. arXiv preprint arXiv:2104.02145 , year =

  16. [24]

    arXiv preprint arXiv:2310.18018 , year =

    Sainz, Oscar and Campos, Jon Ander and Garc. arXiv preprint arXiv:2310.18018 , year =

  17. [25]

    2021 , note =

    Ruder, Sebastian , title =. 2021 , note =

  18. [26]

    Shortcut Learning in Deep Neural Networks , booktitle =

    Geirhos, Robert and Jacobsen, J. Shortcut Learning in Deep Neural Networks , booktitle =. 2020 , doi =

  19. [27]

    and others , title =

    Wilkinson, Mark D. and others , title =. Scientific Data , volume =. 2016 , doi =

  20. [28]

    , title =

    Seyhan, Attila A. , title =. Translational Medicine Communications , volume =. 2019 , doi =

  21. [29]

    , title =

    Fuster, Valentin and Sweeny, Joseph M. , title =. Circulation , volume =. 2011 , doi =

  22. [30]

    2025 , note =

    Bourdois, Loïck , title =. 2025 , note =

  23. [31]

    and Adeli, Ehsan and others , title =

    Bommasani, Rishi and Hudson, Drew A. and Adeli, Ehsan and others , title =. arXiv preprint arXiv:2108.07258 , year =

  24. [32]

    2026 , url =

    Artificial Intelligence Index Report 2026 , institution =. 2026 , url =

  25. [33]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =

    Erasing Concepts from Diffusion Models , author =. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =

  26. [34]

    Li and Ann-Kathrin Dombrowski and Shashwat Goel and Long Phan and Gabriel Mukobi and Nathan Helm-Burger and Rassin Lababidi and Lennart Justen and Andrew B

    Nathaniel Li and Alexander Pan and Anjali Gopal and Summer Yue and Daniel Berrios and Alice Gatti and Justin D. Li and Ann-Kathrin Dombrowski and Shashwat Goel and Long Phan and Gabriel Mukobi and Nathan Helm-Burger and Rassin Lababidi and Lennart Justen and Andrew B. Liu and ...

  27. [35]

    arXiv preprint arXiv:2502.18969 , year =

    (Mis)Fitting: A Survey of Scaling Laws , author =. arXiv preprint arXiv:2502.18969 , year =

  28. [36]

    arXiv preprint arXiv:2001.08361 , year =

    Scaling Laws for Neural Language Models , author =. arXiv preprint arXiv:2001.08361 , year =

  29. [37]

    Bo and Tianyu Xu and Ishan Chatterjee and Katrina Passarella-Ward and Achin Kulshrestha and D Shin , journal =

    Jessica Y. Bo and Tianyu Xu and Ishan Chatterjee and Katrina Passarella-Ward and Achin Kulshrestha and D Shin , journal =. Steerable Chatbots: Personalizing

  30. [38]

    Alireza Salemi and Sheshera Mysore and Michael Bendersky and Hamed Zamani , booktitle =

  31. [39]

    Emergent Misalignment: Narrow finetuning can produce broadly misaligned

    Jan Betley and Daniel Chee Hian Tan and Niels Warncke and Anna Sztyber-Betley and Xuchan Bao and Mart. Emergent Misalignment: Narrow finetuning can produce broadly misaligned. Forty-second International Conference on Machine Learning , year=

  32. [40]

    Training language models to follow instructions with human feedback , volume =

    Ouyang, Long and Wu, Jeffrey and Jiang, Xu and Almeida, Diogo and Wainwright, Carroll and Mishkin, Pamela and Zhang, Chong and Agarwal, Sandhini and Slama, Katarina and Ray, Alex and Schulman, John and Hilton, Jacob and Kelton, Fraser and Miller, Luke and Simens, Maddie and As...

  33. [41]

    International Conference on Machine Learning , pages =

    Scaling Laws for Reward Model Overoptimization , author =. International Conference on Machine Learning , pages =. 2023 , journal =

  34. [42]

    Advances in Neural Information Processing Systems , volume =

    Training Language Models to Follow Instructions with Human Feedback , author =. Advances in Neural Information Processing Systems , volume =

  35. [43]

    Will We Run Out of Data? Limits of

    Villalobos, Pablo and Ho, Anson and Hallawa, Jaime and Atkinson, Tamay and Sevilla, Jaime , journal =. Will We Run Out of Data? Limits of. 2024 , note =

  36. [44]

    Nature , year =

    The Curse of Recursion: Training on Generated Data Makes Models Forget , author =. Nature , year =

  37. [45]

    arXiv preprint arXiv:2306.11644 , year =

    Textbooks Are All You Need , author =. arXiv preprint arXiv:2306.11644 , year =

  38. [46]

    arXiv preprint , year =

    Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance , author =. arXiv preprint , year =

  39. [47]

    Proceedings of the National Academy of Sciences , volume =

    Overcoming Catastrophic Forgetting in Neural Networks , author =. Proceedings of the National Academy of Sciences , volume =

  40. [48]

    Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , pages =

    The Fallacy of AI Functionality , author =. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , pages =. 2022 , doi =

  41. [49]

    arXiv preprint arXiv:2209.07858 , year =

    Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned , author=. arXiv preprint arXiv:2209.07858 , year =

  42. [50]

    Auditing

    Sobieski, Bart. Auditing. arXiv preprint arXiv:2602.02560 , year =

  43. [51]

    2024 , booktitle =

    Groves, Lara and Metcalf, Jacob and Kennedy, Alayna and Vecchione, Briana and Strait, Andrew , title =. 2024 , booktitle =

  44. [52]

    Casper, Stephen and Ezell, Carson and Siegmann, Charlotte and Kolt, Noam and et. al. , booktitle =. Black-Box Access is Insufficient for Rigorous. 2024 , doi =

  45. [53]

    2024 , isbn =

    Lam, Khoa and Lange, Benjamin and Blili-Hamelin, Borhane and Davidovic, Jovana and Brown, Shea and Hasan, Ali , title =. 2024 , isbn =. doi:10.1145/3630106.3658957 , booktitle =

  46. [54]

    Elena, Mihaela and Valentin, Mihai and Rafaela, Coman and Codrut, Turcan , journal =. An

  47. [55]

    Proceedings of the 41st International Conference on Machine Learning , pages =

    Position: Explain to Question not to Justify , author =. Proceedings of the 41st International Conference on Machine Learning , pages =. 2024 , volume =

  48. [56]

    Model Science: Getting Serious About Verification, Explanation and Control of AI Systems , ISBN=

    Biecek, Przemyslaw and Samek, Wojciech , year=. Model Science: Getting Serious About Verification, Explanation and Control of AI Systems , ISBN=. doi:10.3233/FAIA250784 , booktitle=

  49. [57]

    Explainable

    Holzinger, Andreas and Saranti, Anna and Molnar, Christoph and Biecek, Przemyslaw and Samek, Wojciech , booktitle =. Explainable. 2022 , publisher =

  50. [58]

    ACM Computing Surveys , volume =

    A Survey of Methods for Explaining Black Box Models , author =. ACM Computing Surveys , volume =. 2019 , doi =

  51. [59]

    Proceedings of the 34th International Conference on Machine Learning , pages =

    Axiomatic Attribution for Deep Networks , author =. Proceedings of the 34th International Conference on Machine Learning , pages =. 2017 , volume =

  52. [60]

    and Cogswell, Michael and Das, Abhishek and Vedantam, Ramakrishna and Parikh, Devi and Batra, Dhruv , booktitle =

    Selvaraju, Ramprasaath R. and Cogswell, Michael and Das, Abhishek and Vedantam, Ramakrishna and Parikh, Devi and Batra, Dhruv , booktitle =. Grad-

  53. [61]

    PLOS ONE , volume =

    On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation , author =. PLOS ONE , volume =. 2015 , doi =

  54. [62]

    Explainable AI: Interpreting, Explaining and Visualizing Deep Learning , pages =

    Layer-Wise Relevance Propagation: An Overview , author =. Explainable AI: Interpreting, Explaining and Visualizing Deep Learning , pages =. 2019 , publisher =

  55. [63]

    Ribeiro, Marco Tulio and Singh, Sameer and Guestrin, Carlos , booktitle =. ``. 2016 , doi =

  56. [64]

    Advances in Neural Information Processing Systems , volume =

    A Unified Approach to Interpreting Model Predictions , author =. Advances in Neural Information Processing Systems , volume =

  57. [65]

    Unmasking

    Lapuschkin, Sebastian and W. Unmasking. Nature Communications , volume =. 2019 , doi =

  58. [66]

    Advances in Neural Information Processing Systems , volume =

    Sanity Checks for Saliency Maps , author =. Advances in Neural Information Processing Systems , volume =

  59. [67]

    Transformer Circuits Thread , year =

    In-context Learning and Induction Heads , author =. Transformer Circuits Thread , year =

  60. [68]

    Transformer Circuits Thread , year =

    A Mathematical Framework for Transformer Circuits , author =. Transformer Circuits Thread , year =

  61. [69]

    Transformer Circuits Thread , year =

    Towards Monosemanticity: Decomposing Language Models with Dictionary Learning , author =. Transformer Circuits Thread , year =

  62. [70]

    The Twelfth International Conference on Learning Representations , year=

    Sparse Autoencoders Find Highly Interpretable Features in Language Models , author=. The Twelfth International Conference on Learning Representations , year=

  63. [71]

    and McDougall, Callum and MacDiarmid, Monte and Tamkin, Alex and Durmus, Esin and Hume, Tristan and Mosconi, Francesco and Freeman, C

    Templeton, Adly and Conerly, Tom and Marcus, Jonathan and Lindsey, Jack and Bricken, Trenton and Chen, Brian and Pearce, Adam and Citro, Craig and Ameisen, Emmanuel and Jones, Andy and Cunningham, Hoagy and Turner, Nicholas L. and McDougall, Callum and MacDiarmid, Monte and Ta...

  64. [72]

    Nauta, Meike and Seifert, Christin , booktitle =. The. 2023 , publisher =

  65. [73]

    Information Fusion , volume =

    Adversarial Attacks and Defenses in Explainable Artificial Intelligence: A Survey , author =. Information Fusion , volume =. 2024 , doi =

  66. [74]

    Statistics Surveys , volume =

    Interpretable Machine Learning: Fundamental Principles and 10 Grand Challenges , author =. Statistics Surveys , volume =. 2022 , doi =

  67. [75]

    and Bischl, Bernd and Torgo, Luis , journal =

    Vanschoren, Joaquin and van Rijn, Jan N. and Bischl, Bernd and Torgo, Luis , journal =

  68. [76]

    Position: Science of

    Jiang, Han and Zhang, Susu and Yi, Xiaoyuan and Xie, Xing and Xiao, Ziang , journal =. Position: Science of

  69. [77]

    Advances in Neural Information Processing Systems , volume =

    Benchmark Data Repositories for Better Benchmarking , author =. Advances in Neural Information Processing Systems , volume =

  70. [78]

    Mantovani and Jan N

    Bernd Bischl and Giuseppe Casalicchio and Matthias Feurer and Pieter Gijsbers and Frank Hutter and Michel Lang and Rafael G. Mantovani and Jan N. van Rijn and Joaquin Vanschoren , journal =

  71. [79]

    Communications of the ACM , volume =

    Datasheets for Datasets , author =. Communications of the ACM , volume =. 2021 , doi =

  72. [80]

    and Le, Trang T

    Romano, Joseph D. and Le, Trang T. and La Cava, William and Greber, John T. and Goldber, Daniel E. and Moore, Jason H. , journal =. 2022 , doi =

  73. [81]

    Annals of the New York Academy of Sciences , volume =

    Holistic Evaluation of Language Models , author =. Annals of the New York Academy of Sciences , volume =. 2023 , doi =

  74. [82]

    Advances in Neural Information Processing Systems , volume =

    We Should Chart an Atlas of All the World's Models , author =. Advances in Neural Information Processing Systems , volume =. 2025 , url =

  75. [83]

    2024 , howpublished =

    Neuronpedia: Open Platform for Mechanistic Interpretability Research , author =. 2024 , howpublished =

  76. [84]

    Proceedings of the Conference on Fairness, Accountability, and Transparency , pages =

    Model Cards for Model Reporting , author =. Proceedings of the Conference on Fairness, Accountability, and Transparency , pages =. 2019 , doi =

  77. [85]

    Journal of Legal Analysis , volume =

    Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models , author =. Journal of Legal Analysis , volume =. 2024 , doi =

  78. [86]

    and Celi, Leo Anthony and Gichoya, Judy and Jurafsky, Dan and Szolovits, Peter and Bates, David W

    Zack, Travis and Lehman, Eric and Suzgun, Mirac and Rodriguez, Jorge A. and Celi, Leo Anthony and Gichoya, Judy and Jurafsky, Dan and Szolovits, Peter and Bates, David W. and Abdulnour, Raja-Elie E. and Butte, Atul J. and Alsentzer, Emily , journal =. Assessing the potential o...

  79. [87]

    2024 , issn =

    Explainable Artificial Intelligence (XAI) 2.0: A manifesto of open challenges and interdisciplinary research directions , journal =. 2024 , issn =. doi:https://doi.org/10.1016/j.inffus.2024.102301 , url =

  80. [88]

    Advances in Neural Information Processing Systems (NeurIPS) , volume =

    Stable Bias: Analyzing Societal Representations in Diffusion Models , author =. Advances in Neural Information Processing Systems (NeurIPS) , volume =

  81. [89]

    2023 , howpublished =

    Generative. 2023 , howpublished =

  82. [90]

    and Janizek, Joseph D

    DeGrave, Alex J. and Janizek, Joseph D. and Lee, Su-In , journal =. 2021 , doi =

  83. [91]

    and Liles, W

    Trivedi, Anusua and Robinson, Caleb and Blazes, Marian and Ortiz, Anthony and Desbiens, Jocelyn and Gupta, Sunil and Dodhia, Rahul and Bhatraju, Pavan K. and Liles, W. Conrad and Lee, Aaron Y. and Kalpathy-Cramer, Jayashree and Lavista Ferres, Juan M. , journal =. Deep learnin...

  84. [92]

    and Rasheed, Shirin and Moradidakhel, Arghavan and Tahir, Amjed and Khomh, Foutse , booktitle =

    Majdinasab, Vahid and Bishop, Michael J. and Rasheed, Shirin and Moradidakhel, Arghavan and Tahir, Amjed and Khomh, Foutse , booktitle =. Assessing the Security of. 2024 , doi =

  85. [93]

    Do Users Write More Insecure Code with

    Perry, Neil and Srivastava, Megha and Kumar, Deepak and Boneh, Dan , booktitle =. Do Users Write More Insecure Code with. 2023 , doi =

  86. [94]

    Overview of

    Wang, Lei and Wen, Zehua and Liu, Shi-Wei and Zhang, Lihong and Finley, Cierra and Lee, Ho-Jin and Fan, Hua-Jun Shawn , journal =. Overview of. 2024 , doi =

  87. [95]

    Proceedings of the AAAI Conference on Artificial Intelligence , year =

    Stable Diffusion Exposed: Gender Bias from Prompt to Image , author =. Proceedings of the AAAI Conference on Artificial Intelligence , year =

  88. [96]

    NeurIPS , year=

    Stable Bias: Analyzing Societal Representations in Diffusion Models , author=. NeurIPS , year=

  89. [97]

    Assessing the Security of

    Majdinasab, Vahid and others , booktitle=. Assessing the Security of

  90. [98]

    2026 , howpublished =

    Alphabet. 2026 , howpublished =

  91. [99]

    2025 , howpublished =

    The Future of. 2025 , howpublished =

  92. [100]

    2026 , howpublished =

    Claude. 2026 , howpublished =

  93. [101]

    2024 , howpublished =

    With 10x Growth Since 2023,. 2024 , howpublished =

  94. [102]

    2026 , url =

    Whisper. 2026 , url =

  95. [103]

    2024 , howpublished =

  96. [104]

    Journal of Clinical Oncology , volume =

    Sybil: A Validated Deep Learning Model to Predict Future Lung Cancer Risk From a Single Low-Dose Chest Computed Tomography , author =. Journal of Clinical Oncology , volume =. 2023 , doi =

  97. [105]

    Validation of

    Pasquinelli, Mary and others , booktitle =. Validation of. 2025 , note =

  98. [106]

    and Schellmann, Hilke and Sloane, Mona , title =

    Koenecke, Allison and Choi, Anna Seo Gyeong and Mei, Katelyn X. and Schellmann, Hilke and Sloane, Mona , title =. 2024 , booktitle =

  99. [107]

    Ribeiro, Marco Tulio and Singh, Sameer and Guestrin, Carlos , booktitle =. ``

  100. [108]

    Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages =

    Annotation Artifacts in Natural Language Inference Data , author =. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , pages =

  101. [109]

    European Review , volume =

    Strathern, Marilyn , title =. European Review , volume =

  102. [110]

    Statistics in Medicine , volume =

    Biomarkers and Surrogate Endpoints in Clinical Trials , author =. Statistics in Medicine , volume =

  103. [111]

    Science , volume =

    Estimating the Reproducibility of Psychological Science , author =. Science , volume =

  104. [112]

    2017 , publisher =

    The Testing Charade: Pretending to Make Schools Better , author =. 2017 , publisher =

  105. [113]

    Journal of Neurology, Neurosurgery & Psychiatry , volume =

    Loss of Recent Memory After Bilateral Hippocampal Lesions , author =. Journal of Neurology, Neurosurgery & Psychiatry , volume =

  106. [114]

    and Damasio, Antonio R

    Damasio, Hanna and Grabowski, Thomas and Frank, Randall and Galaburda, Albert M. and Damasio, Antonio R. , journal =. The Return of

  107. [115]

    Bulletin de la Soci

    Remarks on the Seat of the Faculty of Articulated Language , author =. Bulletin de la Soci

  108. [116]

    1988 , publisher =

    From Neuropsychology to Mental Structure , author =. 1988 , publisher =

  109. [117]

    Interpretability in the Wild: a Circuit for Indirect Object Identification in

    Kevin Ro Wang and Alexandre Variengien and Arthur Conmy and Buck Shlegeris and Jacob Steinhardt , booktitle=. Interpretability in the Wild: a Circuit for Indirect Object Identification in. 2023 , url=

  110. [118]

    Transformer Circuits Thread , year =

    Toy Models of Superposition , author =. Transformer Circuits Thread , year =

  111. [119]

    Brain and Cognition , volume =

    On Drawing Inferences About the Structure of Normal Cognitive Systems from the Analysis of Patterns of Impaired Performance: The Case for Single-Patient Studies , author =. Brain and Cognition , volume =

  112. [120]

    1982 , publisher =

    Vision: A Computational Investigation into the Human Representation and Processing of Visual Information , author =. 1982 , publisher =

  113. [121]

    arXiv preprint arXiv:2503.13401 , year =

    Levels of Analysis for Large Language Models , author =. arXiv preprint arXiv:2503.13401 , year =. 2503.13401 , archiveprefix =

  114. [122]

    1957 , publisher =

    Verbal Behavior , author =. 1957 , publisher =

  115. [123]

    Chomsky, Noam , title =

  116. [124]

    Psychological Review , volume =

    The Magical Number Seven, Plus or Minus Two: Some Limits on Our Capacity for Processing Information , author =. Psychological Review , volume =

  117. [125]

    IRE Transactions on Information Theory , volume =

    The Logic Theory Machine: A Complex Information Processing System , author =. IRE Transactions on Information Theory , volume =

  118. [126]

    Bell System Technical Journal , volume =

    A Mathematical Theory of Communication , author =. Bell System Technical Journal , volume =

  119. [127]

    Locating and Editing Factual Associations in

    Meng, Kevin and Bau, David and Andonian, Alex and Belinkov, Yonatan , journal =. Locating and Editing Factual Associations in

  120. [128]

    arXiv preprint arXiv:2210.07229 , year =

    Mass-Editing Memory in a Transformer , author =. arXiv preprint arXiv:2210.07229 , year =

  121. [129]

    arXiv preprint arXiv:2308.10248 , year =

    Activation Addition: Steering Language Models Without Optimization , author =. arXiv preprint arXiv:2308.10248 , year =

  122. [130]

    Proceedings of the 38th International Conference on Machine Learning , pages =

    Learning Transferable Visual Models From Natural Language Supervision , author =. Proceedings of the 38th International Conference on Machine Learning , pages =. 2021 , volume =

  123. [131]

    NEJM AI , author =

    A. NEJM AI , author =. doi:10.1056/AIoa2400640 , language =

  124. [132]

    2023 , publisher =

    Gao, Bowen and Qiang, Bo and Tan, Haichuan and Jia, Yinjun and Ren, Minsi and Lu, Minsi and Liu, Jingjing and Ma, Wei-Ying and Lan, Yanyan , booktitle =. 2023 , publisher =

  125. [133]

    2025 , eprint =

    Agrawal, Kumar Krishna and Liu, Longchao and Lian, Long and Nercessian, Michael and Harguindeguy, Natalia and Wu, Yufu and Mikhael, Peter and Lin, Gigin and Sequist, Lecia V and Fintelmann, Florian and Darrell, Trevor and Bai, Yutong and Chung, Maggie and Yala, Adam , journal ...

  126. [134]

    Nature , volume =

    Merlin: A Computed Tomography Vision--Language Foundation Model and Dataset , author =. Nature , volume =. 2026 , publisher =

  127. [135]

    Merlin: A Vision Language Foundation Model for

    Blankemeier, Louis and Cohen, Joseph Paul and Kumar, Ashwin and Van Veen, Dave and Gardezi, Syed Jamal Safdar and Paschali, Magdalini and Chen, Zhihong and Delbrouck, Jean-Benoit and Reis, Eduardo and Truyts, Cesar and Bluethgen, Christian and Jensen, Malte Engmann Kjeldskov a...

  128. [136]

    Nature Machine Intelligence , volume =

    Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead , author =. Nature Machine Intelligence , volume =

  129. [137]

    Distill , year =

    Zoom In: An Introduction to Circuits , author =. Distill , year =

  130. [138]

    Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (

    Kim, Been and others , booktitle =. Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (

  131. [139]

    ICML , year =

    Understanding Black-box Predictions via Influence Functions , author =. ICML , year =

  132. [140]

    NeurIPS , year =

    Sanity Checks for Saliency Maps , author =. NeurIPS , year =

  133. [141]

    NAACL , year =

    Attention is not Explanation , author =. NAACL , year =

  134. [142]

    Information Fusion , volume =

    Adversarial Attacks and Defenses in Explainable Artificial Intelligence: A Survey , author =. Information Fusion , volume =

  135. [143]

    Neither Valid nor Reliable? Investigating the Use of

    Chehbouni, Khaoula and Haddou, Mohammed and Cheung, Jackie Chi Kit and Farnadi, Golnoosh , booktitle =. Neither Valid nor Reliable? Investigating the Use of. 2025 , institution =

  136. [144]

    Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =

    Measurement and Fairness , author =. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =. 2021 , publisher =

  137. [145]

    Raina, Vyas and Liusie, Adian and Gales, Mark , booktitle =. Is. 2024 , publisher =

  138. [146]

    Know Thy Judge: On the Robustness Meta-Evaluation of

    Eiras, Francisco and Zemour, Eliott and Lin, Eric and Mugunthan, Vaikkunth , journal =. Know Thy Judge: On the Robustness Meta-Evaluation of

  139. [147]

    The Thirteenth International Conference on Learning Representations , year =

    Unsupervised Model Tree Heritage Recovery , author =. The Thirteenth International Conference on Learning Representations , year =

  140. [148]

    Advances in Neural Information Processing Systems , volume =

    Hyper-Representations as Generative Models: Sampling Unseen Neural Network Weights , author =. Advances in Neural Information Processing Systems , volume =

  141. [149]

    The Thirteenth International Conference on Learning Representations , year =

    Deep Linear Probe Generators for Weight Space Learning , author =. The Thirteenth International Conference on Learning Representations , year =

  142. [150]

    The Thirteenth International Conference on Learning Representations , year =

    PhyloLM: Inferring the Phylogeny of Large Language Models and Predicting Their Performances in Benchmarks , author =. The Thirteenth International Conference on Learning Representations , year =

  143. [151]

    Pattern Recognition , author =

    Checklist for responsible deep learning modeling of medical images based on. Pattern Recognition , author =. 2021 , pages =. doi:10.1016/j.patcog.2021.108035 , language =

  144. [152]

    Journal of the Operational Research Society , author =

    Transparency, auditability, and explainability of machine learning models in credit scoring , volume =. Journal of the Operational Research Society , author =. 2022 , pages =. doi:10.1080/01605682.2021.1922098 , number =

  145. [153]

    doi:10.1007/s10676-024-09795-1 , number =

    Ethics and Information Technology , author =. doi:10.1007/s10676-024-09795-1 , number =

  146. [154]

    auditor: an

    Gosiewska, Alicja and Biecek, Przemyslaw , year =. auditor: an. doi:10.48550/arXiv.1809.07763 , publisher =

  147. [155]

    Nature Human Behaviour , author =

    Protect our environment from information overload , volume =. Nature Human Behaviour , author =. 2024 , note =. doi:10.1038/s41562-024-01833-8 , number =

  148. [156]

    Ethics and Information Technology , author =

    Generative. Ethics and Information Technology , author =. doi:10.1007/s10676-023-09728-4 , number =

  149. [157]

    Birds look like cars:

    Baniecki, Hubert and Biecek, Przemyslaw , year =. Birds look like cars:

  150. [158]

    Mieleszczenko-Kowszewicz, Wiktoria and Płudowski, Dawid and Kołodziejczyk, Filip and Świstak, Jakub and Sienkiewicz, Julian and Biecek, Przemysław , year =. The

  151. [159]

    Adversarial attacks and defenses in explainable artificial intelligence:

    Baniecki, Hubert and Biecek, Przemyslaw , year =. Adversarial attacks and defenses in explainable artificial intelligence:. doi:10.1016/j.inffus.2024.102303 , journal =

  152. [160]

    and Nalepa, Jakub and Longépé, Nicolas and Biecek, Przemyslaw , year =

    Zaigrajew, Vladimir and Baniecki, Hubert and Tulczyjew, Lukasz and Wijata, Agata M. and Nalepa, Jakub and Longépé, Nicolas and Biecek, Przemyslaw , year =. Red

  153. [161]

    Baniecki, Hubert and Kretowicz, Wojciech and Biecek, Przemyslaw , year =. Fooling. Lecture. doi:10.1007/978-3-031-26409-2_8 , pages =

  154. [162]

    Komorowski, Piotr and Baniecki, Hubert and Biecek, Przemysław , year =. Towards. doi:10.1109/cvprw59228.2023.00383 , booktitle =

  155. [163]

    Płudowski, Dawid and Spinnato, Francesco and Wilczyński, Piotr and Kotowski, Krzysztof and Ntagiou, Evridiki Vasileia and Guidotti, Riccardo and Biecek, Przemysław , year =

  156. [164]

    The Thirteenth International Conference on Learning Representations , year=

    Rethinking Visual Counterfactual Explanations Through Region Constraint , author=. The Thirteenth International Conference on Learning Representations , year=

  157. [165]

    ACM Computing Surveys , author =

    A. ACM Computing Surveys , author =. 2019 , pages =. doi:10.1145/3236009 , number =

  158. [166]

    Explainable AI Methods - A Brief Overview

    Holzinger, Andreas and Saranti, Anna and Molnar, Christoph and Biecek, Przemyslaw and Samek, Wojciech. Explainable AI Methods - A Brief Overview. xxAI - Beyond Explainable AI: International Workshop, Held in Conjunction with ICML 2020, July 18, 2020, Vienna, Austria, Revised a...

  159. [167]

    Resistance

    Wilczyński, Piotr and Mieleszczenko-Kowszewicz, Wiktoria and Biecek, Przemysław , year =. Resistance. Frontiers in

  160. [168]

    Sobieski, Bartlomiej and Biecek, Przemyslaw , year =. Global. Lecture. doi:10.1007/978-3-031-73036-8_5 , pages =

  161. [169]

    doi:10.1007/s10994-022-06204-w , journal =

    Hryniewska, Weronika and Grudzień, Adrianna and Biecek, Przemysław , year =. doi:10.1007/s10994-022-06204-w , journal =

  162. [170]

    The grammar of interactive explanatory model analysis , issn =

    Baniecki, Hubert and Parzych, Dariusz and Biecek, Przemyslaw , year =. The grammar of interactive explanatory model analysis , issn =. doi:10.1007/s10618-023-00924-w , journal =

  163. [171]

    Computational Statistics , author =

    Exploring local explanations of nonlinear models using animated linear projections , volume =. Computational Statistics , author =. 2025 , pages =. doi:10.1007/s00180-023-01453-2 , number =

  164. [172]

    2019 , pages =

    Journal of Open Source Software , author =. 2019 , pages =. doi:10.21105/joss.01798 , number =

  165. [173]

    Explanatory

    Biecek, Przemyslaw and Burzykowski, Tomasz , year =. Explanatory

  166. [174]

    Kuzba, Michal and Biecek, Przemyslaw , year =. What. Communications in. doi:10.1007/978-3-030-65965-3_30 , pages =

  167. [175]

    Ouyang, Long and Wu, Jeff and Jiang, Xu and Almeida, Diogo and Wainwright, Carroll L. and Mishkin, Pamela and Zhang, Chong and Agarwal, Sandhini and Slama, Katarina and Ray, Alex and Schulman, John and Hilton, Jacob and Kelton, Fraser and Miller, Luke and Simens, Maddie and As...

  168. [176]

    2022 , author =

    Constitutional. 2022 , author =

  169. [177]

    and Finn, Chelsea , year =

    Rafailov, Rafael and Sharma, Archit and Mitchell, Eric and Ermon, Stefano and Manning, Christopher D. and Finn, Chelsea , year =. Direct

  170. [178]

    Lee, Harrison and Phatale, Samrat and Mansoor, Hassan and Mesnard, Thomas and Ferret, Johan and Lu, Kellie and Bishop, Colton and Hall, Ethan and Carbune, Victor and Rastogi, Abhinav and Prakash, Sushant , year =

  171. [179]

    Zhou, Ziqi and Hu, Shengshan and Li, Minghui and Zhang, Hangtao and Zhang, Yechao and Jin, Hai , year =

  172. [180]

    2022 , author =

    On the. 2022 , author =

  173. [181]

    Artificial Intelligence , author =

    Explanation in artificial intelligence:. Artificial Intelligence , author =. 2019 , pages =. doi:10.1016/j.artint.2018.07.007 , language =

  174. [182]

    Principles of

    Kulesza, Todd and Burnett, Margaret and Wong, Weng-Keen and Stumpf, Simone , year =. Principles of. doi:10.1145/2678025.2701399 , booktitle =

  175. [183]

    Choo, Jaegul and Liu, Shixia , year =. Visual

  176. [184]

    Hohman, Fred and Kahng, Minsuk and Pienta, Robert and Chau, Duen Horng , year =. Visual

  177. [185]

    , year =

    Strobelt, Hendrik and Webson, Albert and Sanh, Victor and Hoover, Benjamin and Beyer, Johanna and Pfister, Hanspeter and Rush, Alexander M. , year =. Interactive and. doi:10.1109/tvcg.2022.3209479 , journal =

  178. [186]

    doi:10.1145/3581641.3584059 , author =

    2023 , pages =. doi:10.1145/3581641.3584059 , author =

  179. [187]

    2023 , pages =

    IEEE Transactions on Visualization and Computer Graphics , author =. 2023 , pages =. doi:10.1109/tvcg.2023.3327168 , urldate =

  180. [188]

    doi:10.1145/3491101.3519729 , urldate =

    Wu, Tongshuang and Jiang, Ellen and Donsbach, Aaron and Gray, Jeff and Molina, Alejandra and Terry, Michael and Cai, Carrie J , year =. doi:10.1145/3491101.3519729 , urldate =

  181. [189]

    and Wang, Jiayao and Berger, Matthew , year =

    DeRose, Joseph F. and Wang, Jiayao and Berger, Matthew , year =. Attention

  182. [190]

    IEEE Transactions on Visualization and Computer Graphics , author =

    How. IEEE Transactions on Visualization and Computer Graphics , author =. 2023 , pages =. doi:10.1109/tvcg.2023.3261935 , number =

  183. [191]

    Yeh, Catherine and Chen, Yida and Wu, Aoyu and Chen, Cynthia and Viégas, Fernanda and Wattenberg, Martin , year =

  184. [192]

    and Cogswell, Michael and Das, Abhishek and Vedantam, Ramakrishna and Parikh, Devi and Batra, Dhruv , year =

    Selvaraju, Ramprasaath R. and Cogswell, Michael and Das, Abhishek and Vedantam, Ramakrishna and Parikh, Devi and Batra, Dhruv , year =. Grad-. doi:10.1109/iccv.2017.74 , booktitle =

  185. [193]

    Interpreting

    Kaur, Harmanpreet and Nori, Harsha and Jenkins, Samuel and Caruana, Rich and Wallach, Hanna and Wortman Vaughan, Jennifer , year =. Interpreting. doi:10.1145/3313831.3376219 , booktitle =

  186. [194]

    Does the

    Bansal, Gagan and Wu, Tongshuang and Zhou, Joyce and Fok, Raymond and Nushi, Besmira and Kamar, Ece and Ribeiro, Marco Tulio and Weld, Daniel , year =. Does the. doi:10.1145/3411764.3445717 , booktitle =

  187. [195]

    and Steinhardt, Jacob , year =

    Gandelsman, Yossi and Efros, Alexei A. and Steinhardt, Jacob , year =. Interpreting

  188. [196]

    2021 , isbn =

    Przemyslaw Biecek and Tomasz Burzykowski , title =. 2021 , isbn =

  189. [197]

    What does

    Shtedritski, Aleksandar and Rupprecht, Christian and Vedaldi, Andrea , year =. What does. doi:10.1109/iccv51070.2023.01101 , booktitle =

  190. [198]

    and Lakkaraju, Himabindu , year =

    Bhalla, Usha and Oesterling, Alex and Srinivas, Suraj and Calmon, Flavio P. and Lakkaraju, Himabindu , year =. Interpreting

  191. [199]

    Interpreting

    Zaigrajew, Vladimir and Baniecki, Hubert and Biecek, Przemyslaw , year =. Interpreting

  192. [200]

    Successor

    Gould, Rhys and Ong, Euan and Ogden, George and Conmy, Arthur , year =. Successor

  193. [201]

    Nam, Andrew and Conklin, Henry and Yang, Yukang and Griffiths, Thomas and Cohen, Jonathan and Leslie, Sarah-Jane , year =. Causal

  194. [202]

    Nature Communications , author =

    Protein language models trained on multiple sequence alignments learn phylogenetic relationships , volume =. Nature Communications , author =. doi:10.1038/s41467-022-34032-y , number =

  195. [203]

    Proceedings of the National Academy of Sciences , author =

    Acquisition of chess knowledge in. Proceedings of the National Academy of Sciences , author =. doi:10.1073/pnas.2206625119 , number =

  196. [204]

    Unveiling

    Pálsson, Aðalsteinn and Björnsson, Yngvi , year =. Unveiling. doi:10.24963/ijcai.2023/541 , booktitle =

  197. [205]

    Dreyer, Maximilian and Hufe, Lorenz and Berend, Jim and Wiegand, Thomas and Lapuschkin, Sebastian and Samek, Wojciech , year =. From

  198. [206]

    Locating and

    Meng, Kevin and Bau, David and Andonian, Alex and Belinkov, Yonatan , year =. Locating and

  199. [207]

    Meng, Kevin and Sharma, Arnab Sen and Andonian, Alex and Belinkov, Yonatan and Bau, David , year =. Mass-

  200. [208]

    Interpretable and

    Delfosse, Quentin and Shindo, Hikaru and Dhami, Devendra and Kersting, Kristian , year =. Interpretable and

  201. [209]

    Tan, Juntao and Zhang, Yongfeng , year =

  202. [210]

    doi:10.48550/arXiv.2404.03118 , publisher =

    Stan, Gabriela Ben Melech and Aflalo, Estelle and Rohekar, Raanan Yehezkel and Bhiwandiwalla, Anahita and Tseng, Shao-Yen and Olson, Matthew Lyle and Gurwicz, Yaniv and Wu, Chenfei and Duan, Nan and Lal, Vasudev , year =. doi:10.48550/arXiv.2404.03118 , publisher =

  203. [211]

    2022 , author =

    Red. 2022 , author =

  204. [212]

    Lovisotto, Giulio and Finnie, Nicole and Munoz, Mauricio and Mummadi, Chaithanya Kumar and Metzen, Jan Hendrik , year =. Give

  205. [213]

    Computers in Biology and Medicine , author =

    Overview of. Computers in Biology and Medicine , author =. 2024 , pages =. doi:10.1016/j.compbiomed.2024.108620 , urldate =

  206. [214]

    Nature Communications , author =

    Unmasking. Nature Communications , author =. doi:10.1038/s41467-019-08987-4 , number =

  207. [215]

    Ribeiro, Marco Tulio and Singh, Sameer and Guestrin, Carlos , year =. ". doi:10.1145/2939672.2939778 , booktitle =

  208. [216]

    Nature Machine Intelligence , author =

    Explainable. Nature Machine Intelligence , author =. 2025 , pages =. doi:10.1038/s42256-025-01000-2 , number =

  209. [217]

    and Schellmann, Hilke and Sloane, Mona , year =

    Koenecke, Allison and Choi, Anna Seo Gyeong and Mei, Katelyn X. and Schellmann, Hilke and Sloane, Mona , year =. Careless. doi:10.1145/3630106.3658996 , booktitle =

  210. [218]

    Journal of Legal Analysis , volume =

    Dahl, Matthew and Magesh, Varun and Suzgun, Mirac and Ho, Daniel E , title =. Journal of Legal Analysis , volume =. 2024 , issn =. doi:10.1093/jla/laae003 , url =

  211. [219]

    Assessing the potential of

    Zack, Travis and. Assessing the potential of. The Lancet Digital Health , year =. doi:10.1016/s2589-7500(23)00225-x , language =

  212. [220]

    Assessing the

    Majdinasab, Vahid and Bishop, Michael Joshua and Rasheed, Shawn and Moradidakhel, Arghavan and Tahir, Amjed and Khomh, Foutse , year =. Assessing the. doi:10.1109/saner60148.2024.00051 , booktitle =

  213. [221]

    2024 , pages =

    Academic Pathology , author =. 2024 , pages =. doi:10.1016/j.acpath.2023.100099 , number =

  214. [222]

    Adultification

    Castleman, Jane and Korolova, Aleksandra , year =. Adultification. Proceedings of the 2025. doi:10.1145/3715275.3732178 , urldate =

  215. [223]

    Statistical modeling: the two cultures , url =

    Breiman, Leo , doi =. Statistical modeling: the two cultures , url =. Statistical Science , issn =

  216. [224]

    , title =

    Lipton, Zachary C. , title =. 2018 , publisher =. doi:10.1145/3236386.3241340 , journal =

  217. [225]

    Adversarial machine learning: a taxonomy and terminology of attacks and mitigations , url =

    Vassilev, Apostol and Oprea, Alina and Fordyce, Alie and Anderson, Hyrum , year =. Adversarial machine learning: a taxonomy and terminology of attacks and mitigations , url =

  218. [226]

    Patterns , author =

    Extensions of the. Patterns , author =. 2020 , pages =. doi:10.1016/j.patter.2020.100129 , language =

  219. [227]

    doi:10.35629/8028-1002031924 , journal =

    Mitoi Elena and Moldoveanu Valentin and Cazazian Rafaela and Turlea Codrut , year =. doi:10.35629/8028-1002031924 , journal =

  220. [228]

    arXiv , note=

    LLaMA 2: Open Foundation and Fine-Tuned Chat Models , author=. arXiv , note=

  221. [229]

    and Wright, Marvin N

    Hiabu, Munir and Meyer, Joseph T. and Wright, Marvin N. , year =. Unifying local and global model explanations by functional decomposition of low dimensional structures , url =

  222. [230]

    Decision Support Systems , author =

    Simpler is better:. Decision Support Systems , author =. 2021 , pages =. doi:10.1016/j.dss.2021.113556 , language =

  223. [231]

    doi:10.2139/ssrn.4389797 , language =

    SSRN Electronic Journal , author =. doi:10.2139/ssrn.4389797 , language =

  224. [232]

    Hendrycks, Dan , year =. Natural

  225. [233]

    and Jati, Arindam and Mukherjee, Sumanta and Aggarwal, Nupur and Sarpatwar, Kanthi and Ganapavarapu, Giridhar and Vaculin, Roman , year =

    Raykar, Vikas C. and Jati, Arindam and Mukherjee, Sumanta and Aggarwal, Nupur and Sarpatwar, Kanthi and Ganapavarapu, Giridhar and Vaculin, Roman , year =

  226. [234]

    and Lundberg, Scott M

    Chen, Hugh and Covert, Ian C. and Lundberg, Scott M. and Lee, Su-In , year =. Algorithms to estimate

  227. [235]

    Explainable

    Holzinger, Andreas and Saranti, Anna and Molnar, Christoph and Biecek, Przemyslaw and Samek, Wojciech , editor =. Explainable. 2022 , doi =

  228. [236]

    Wang, Jiaxuan and Wiens, Jenna and Lundberg, Scott , year =. Shapley

  229. [237]

    Asymmetric

    Frye, Christopher and Rowat, Colin and Feige, Ilya , year =. Asymmetric

  230. [238]

    Jain, Aditya and Ravula, Manish and Ghosh, Joydeep , year =. Biased

  231. [239]

    Graph-guided random forest for gene set selection , url =

    Pfeifer, Bastian and Baniecki, Hubert and Saranti, Anna and Biecek, Przemyslaw and Holzinger, Andreas , year =. Graph-guided random forest for gene set selection , url =

  232. [240]

    Explainable expected goal models for performance analysis in football analytics , url =

    Cavus, Mustafa and Biecek, Przemysław , year =. Explainable expected goal models for performance analysis in football analytics , url =. doi:10.1109/DSAA54385.2022.10032440 , booktitle =

  233. [241]

    Proceedings of the AAAI Conference on Artificial Intelligence , author =

    Manipulating. Proceedings of the AAAI Conference on Artificial Intelligence , author =. 2022 , pages =. doi:10.1609/aaai.v36i11.21590 , number =

  234. [242]

    Exploring

    Spyrison, Nicholas and Cook, Dianne , year =. Exploring

  235. [243]

    Slack, Dylan and Hilgard, Sophie and Jia, Emily and Singh, Sameer and Lakkaraju, Himabindu , year =. Fooling

  236. [244]

    2023 , pages =

    Knowledge-Based Systems , author =. 2023 , pages =. doi:10.1016/j.knosys.2022.110234 , urldate =

  237. [245]

    NeurIPS Workshops , year =

    Is Attention Interpretation? A Quantitative Assessment On Sets , author =. NeurIPS Workshops , year =

  238. [246]

    ACL , year =

    Bibal, Adrien and Cardon, R. ACL , year =

  239. [247]

    Bordt, Sebastian and Raidl, Eric and Luxburg, Ulrike , booktitle =

  240. [248]

    2022 , journal=

    Formal Algorithms for Transformers , author=. 2022 , journal=

  241. [249]

    2017 , booktitle=

    Attention Is All You Need , author=. 2017 , booktitle=

  242. [250]

    and Pfister, Tomas , booktitle=

    Arik, Sercan Ö. and Pfister, Tomas , booktitle=

  243. [251]

    ICLR , year=

    Noah Hollmann and Samuel M. ICLR , year=

  244. [252]

    ICLR , year=

    Samuel M. ICLR , year=

  245. [253]

    Bayan Bruss and Tom Goldstein , title =

    Gowthami Somepalli and Micah Goldblum and Avi Schwarzschild and C. Bayan Bruss and Tom Goldstein , title =. arXiv preprint arXiv:2106.01342 , year =

  246. [254]

    Chirag Agarwal and Satyapriya Krishna and Eshika Saxena and Martin Pawelczyk and Nari Johnson and Isha Puri and Marinka Zitnik and Himabindu Lakkaraju , booktitle=

  247. [255]

    KDD , year =

    Chen, Tianqi and Guestrin, Carlos , title =. KDD , year =

  248. [256]

    ArXiv , year=

    True to the Model or True to the Data? , author=. ArXiv , year=

  249. [257]

    2021 , issn =

    Explaining individual predictions when features are dependent: More accurate approximations to Shapley values , journal =. 2021 , issn =

  250. [258]

    International Conference on Artificial Intelligence and Statistics , year=

    Feature relevance quantification in explainable AI: A causality problem , author=. International Conference on Artificial Intelligence and Statistics , year=

  251. [259]

    Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =

    Unifying local and global model explanations by functional decomposition of low dimensional structures , author =. Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =. 2023 , editor =

  252. [260]

    Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =

    From Shapley Values to Generalized Additive Models and back , author =. Proceedings of The 26th International Conference on Artificial Intelligence and Statistics , pages =. 2023 , editor =

  253. [261]

    Journal of Machine Learning Research , year =

    Che-Ping Tsai and Chih-Kuan Yeh and Pradeep Ravikumar , title =. Journal of Machine Learning Research , year =

  254. [262]

    Proceedings of the 37th International Conference on Machine Learning , pages =

    The Shapley Taylor Interaction Index , author =. Proceedings of the 37th International Conference on Machine Learning , pages =. 2020 , editor =

  255. [263]

    Advances in Neural Information Processing Systems 30 , publisher =

    A unified approach to interpreting model predictions , author =. Advances in Neural Information Processing Systems 30 , publisher =

  256. [264]

    Nature Machine Intelligence , author =

    From local explanations to global understanding with explainable. Nature Machine Intelligence , author =. 2020 , pages =. doi:10.1038/s42256-019-0138-9 , language =

  257. [265]

    Shapley, Lloyd S , year = 1953, booktitle =

  258. [266]

    Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence,

    The Shapley Value in Machine Learning , author =. Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence,. 2022 , month =

  259. [267]

    Plos one , volume=

    Machine learning predicts and provides insights into milk acidification rates of Lactococcus lactis , author=. Plos one , volume=. 2021 , publisher=

  260. [268]

    Olsen, Lars Henry Berge and Glad, Ingrid Kristine and Jullum, Martin and Aas, Kjersti , year =. A

  261. [269]

    2024 , volume =

    Position: Explain to Question not to Justify , author =. 2024 , volume =

  262. [270]

    , title =

    Tukey, John W. , title =. The Annals of Mathematical Statistics , volume =. 1962 , publisher =

  263. [271]

    , title =

    Cleveland, William S. , title =. International Statistical Review , volume =. 2001 , publisher =

  264. [272]

    Journal of Computational and Graphical Statistics , volume =

    Donoho, David , title =. Journal of Computational and Graphical Statistics , volume =. 2017 , publisher =

  265. [273]

    , title =

    Tukey, John W. , title =

  266. [274]

    Statistical Science , volume =

    Breiman, Leo , title =. Statistical Science , volume =. 2001 , publisher =

  267. [275]

    , title =

    Turing, Alan M. , title =. Mind , volume =

  268. [276]

    Nature , volume =

    Learning Representations by Back-Propagating Errors , author =. Nature , volume =

  269. [277]

    Proceedings of the 10th European Conference on Artificial Intelligence (ECAI) , pages =

    Planning as Satisfiability , author =. Proceedings of the 10th European Conference on Artificial Intelligence (ECAI) , pages =

  270. [278]

    Artificial Intelligence , volume =

    Collaborative Plans for Complex Group Action , author =. Artificial Intelligence , volume =

  271. [279]

    The Entropy Formula for the

    Grisha Perelman , howpublished =. The Entropy Formula for the

  272. [280]

    Causality , author =

  273. [281]

    2026 , author =

    What can artificial intelligence do for soil health in agriculture? , journal =. 2026 , author =

  274. [282]

    2026 , author =

    Agricultural Systems , volume =. 2026 , author =

  275. [283]

    2024 , author =

    Human-Centered AI in smart farming: Toward Agriculture 5.0 , journal =. 2024 , author =

Pith tools

Reviewed June 28, 2026 · model on record in the stance chip above.