Pith. sign in

REVIEW 3 major objections 5 minor 30 references

Perspectives on Explanation Formats From Two Stakeholder Groups in Germany: Software Providers and Dairy Farmers

T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Software providers misjudge what farmers want from AI explanations.

desk verdict A small, honest exploratory study whose central claim about provider misperceptions is plausible but statistically thin—worth a serious referee, not worth citing as evidence yet. read the letter →

arxiv 2506.11665 v1 pith:PUBBM6KT submitted 2025-06-13 cs.HC

classification cs.HC
keywords explainableAIdairyfarmingdecisionsupportsystemsstakeholderperspectivescomprehensibilitytrustexplanationformatsuserrequirements
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that software providers in the German dairy industry make assumptions about farmers' explanation preferences that do not match what farmers themselves report. It compares 13 software providers' guesses with ratings from 14 farmers in a prior study, using four explanation formats for mastitis warnings in a hypothetical herd management system. Providers rated the rule-based format as best suited for farmers, while farmers most favored the time series. The paper argues this mismatch may be one reason behind the cautious adoption of digital decision support systems, and that a thorough user requirements analysis could close the gap. The authors present the finding as a tendency rather than a representative result, given the small sample sizes.

What carries the argument

The instrument is a hypothetical herd management system that assesses mastitis risk in dairy cows, offering four explanation formats: textual natural-language statements, rule-based if-then checks, a herd comparison across cows, and a per-cow time series of health parameters. In the earlier farmer study, 14 farmers rated each format for comprehensibility and trust on 5-point Likert scales; in this study, 13 software providers rated the same formats as they believed farmers would, and then picked the single best format. Comparing medians across the two groups turns the four formats into a measurement device for perception gaps between providers and end users.

What would settle it

Ask software providers to predict how the specific farmers they work with would rate each of the four formats, then survey those same farmers: the claimed mismatch is falsified if predicted and actual ratings agree on average across providers.

Watch

Extended reading notes

Core claim

The central claim is that software providers tend to assume farmers' explanation preferences without verifying them, and those assumptions are not necessarily accurate. Concretely, almost half of the 13 providers chose the rule-based format as most suitable for farmers, but the farmers' own most favored format was the time series. Providers also rated the time series as less comprehensible for farmers (median 3.0) than farmers rated it themselves (median 4.5), and they believed farmers would trust it (median 4.0) while farmers reported low trust (median 2.5). The paper reads these gaps as evidence that the two stakeholder groups have divergent perceptions, and concludes that better user-requirements analysis could improve software adaptation and user acceptance.

Load-bearing premise

The paper's conclusion rests on treating the 14 farmers from a separate earlier survey and the 13 software providers from this survey as comparable groups, even though they were recruited through different channels at different times.

Editorial extensions

If this is right

  • If providers' assumptions are systematically off, then gathering farmers' requirements before building explanation features could directly improve acceptance of dairy decision support systems.
  • Because farmers favored the time series while providers preferred the rule-based format, systems that offer a single explanation format may underserve their users; offering several formats would better match actual preferences.
  • Provider beliefs that farmers want concise, time-saving text may be wrong in the other direction: some farmers found textual explanations not detailed enough.
  • The rule-based format's ease of implementation for providers may not translate into trust or clarity for all farmers, since at least one farmer found it unclear.
  • A more thorough requirements analysis is a plausible low-cost remedy worth testing before attributing low adoption to farmer conservatism or technology quality.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A likely consequence beyond this study is that the same perception gap appears in other agricultural AI domains, wherever developers' familiarity with rules and code shapes what they expect users to understand.
  • The direction of the gap is testable at scale: if provider predictions about farmer ratings are compared with actual farmer ratings within the same farms, the discrepancy should shrink or vanish, which would isolate the sampling-comparability problem.
  • The farmers' low trust score for the time series, paired with their high comprehensibility score, raises a distinct question the paper leaves open: understanding a format may not automatically produce trust, so trust needs its own design work.
  • Persona-based user profiles, which the paper proposes for future work, could be validated by A/B testing explanation formats in real herd management systems, turning the reported tendency into a measurable adoption effect.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper reports a small comparative survey study in the German dairy sector. In a prior study, 14 dairy farmers rated four explanation formats (textual, rule-based, herd comparison, time series) for a hypothetical mastitis-detection decision support system in terms of comprehensibility and trust. The present study repeats a similar survey with 13 software providers, who were asked to predict how comprehensible and trustworthy farmers would find each format. The authors compare median ratings and free-text justifications across the two groups, concluding that software providers tend to make inaccurate assumptions about farmers' explanation preferences, particularly overestimating rule-based formats and underestimating farmers' valuation of time-series comprehensibility. The paper is explicitly framed as exploratory and non-representative due to small sample sizes, and it recommends future user-requirements analysis and persona-based design.

Significance. The manuscript addresses a genuinely understudied question: whether XAI explanation formats designed by software providers actually match the preferences of agricultural end users. The use of a concrete, realistic mastitis-detection scenario, the four diverse explanation formats, and the comparison of two stakeholder groups are strengths. The authors are transparent about the small, non-representative samples and present both quantitative medians and qualitative free-text evidence, giving the reader direct access to how the conclusions were reached. If the discrepancy finding were supported by appropriate uncertainty quantification, it would provide a useful empirical motivation for user-centered XAI design in agriculture. The paper does not introduce new theory or methods, but it contributes a small, falsifiable, and potentially informative case study to the HCI/XAI literature.

major comments (3)
  1. [Section 4.3] The central claim that software providers 'tend to make assumptions about farmers' preferences that are not necessarily accurate' (Abstract, Section 4.3, Section 5) is supported only by descriptive median differences between two small independent samples (N_providers = 13, N_farmers = 14). No inferential test, confidence interval, or effect size is reported for the format-by-metric comparisons, such as time-series comprehensibility (median 4.5 vs 3.0) or time-series trust (2.5 vs 4.0). With sample sizes this small, these differences could plausibly arise under the null hypothesis of no true difference. The paper's own caveats about non-representativeness do not address this: the leap from 'medians differ in these samples' to 'providers tend to be inaccurate' requires at least some demonstration that the observed differences are not sampling noise. Please add permutation tests or Mann-Whitney U tests with exact p-values, or bootstrap confidence intervals for median differences, for every format-by-metric comparison, or explicitly restrict the conclusion to 'in this sample' and remove the generalizing 'tend' language.
  2. [Sections 3.2 and 4.1] The comparability of the farmer and provider samples is not established, which is load-bearing for the discrepancy interpretation. The farmer data come from a separate earlier study (Girmay et al., 2024) and are summarized in Section 4.1 without details of recruitment, inclusion criteria, survey administration, or exact question wording. The provider survey (Section 3.2) was administered through different channels (associations and software companies) at a different time and asked participants to predict farmers' ratings rather than rate their own comprehension and trust. Consequently, the differences reported in Section 4.3 confound stakeholder group with timing, recruitment channel, and instrument wording. Please provide a side-by-side table of the two studies' procedures and question wordings, and explicitly discuss how these design differences could alternatively explain the observed median gaps.
  3. [Section 4.3] The comparison is described in terms of medians only, without reporting the distribution overlap or the number of observations underlying each median. The paper's Figure 3 does display quartiles and ranges, but the text often reads point estimates as categorical discrepancies (e.g., 'rated exceptionally highly by farmers, median 5.0' vs 'high by software providers, median 4.0'). Given the ordinal Likert scale and small N, it would be more informative to report the full response distributions (e.g., frequency tables) or at least the interquartile ranges alongside each median in the prose, so the reader can evaluate the strength of the claimed discrepancy.
minor comments (5)
  1. [Abstract and Section 3.2] The abstract says 'we repeat the survey with 13 software providers,' but the provider survey asks respondents to predict farmers' comprehensibility and trust, whereas the farmer survey asked farmers about their own comprehension and trust. The wording is therefore not identical. Please rephrase to 'we adapted the survey' or 'we administered a corresponding survey to software providers' to avoid implying identical items.
  2. [Figure 3] The caption lists the formats 'from left to right' but the figure itself has no labels identifying which boxplot corresponds to textual, rule-based, herd comparison, or time series. Please add direct labels to the boxplots or a legend with the format names.
  3. [Section 4.3] The sentence 'The textual format was rated as well comprehensive by both groups' contains an awkward construction; recommend 'as comprehensible' or 'as easy to understand.' Similar minor grammar issues appear elsewhere in the comparison paragraphs.
  4. [Section 3.2] The paper does not mention ethics approval, informed consent, or data protection procedures beyond stating that participation was anonymous. For survey research with human participants, a statement about consent and ethical compliance should be added.
  5. [References] The reference for Caldiera and Rombach appears as 'Victor R Basili1 Gianluigi Caldiera and H Dieter Rombach' with an apparent formatting artifact, and 'K ¨ohlet al. [2019]' has a spacing issue from the LaTeX source. Please check the reference list for formatting consistency.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the claim rests on a straightforward comparison of two independently collected survey datasets.

full rationale

The paper's central claim is that software providers' assumptions about farmers' preferences are not necessarily accurate. This claim is supported by comparing two independent survey measurements: farmer ratings of four explanation formats collected in a prior study [Girmay et al., 2024] and software providers' beliefs about how farmers would rate those same formats, collected in the present study. The provider ratings are not fitted to, derived from, or defined in terms of the farmer ratings; they are separate measurements that are then compared descriptively. The four explanation formats were designed once for the prior study and reused as stimuli, but reuse of a stimulus is not circularity. The self-citation to the authors' previous work is a normal reference to an external data source, and the comparison in Section 4.3 is not forced by any equation, definition, or fitted parameter. The statistical limitations of the study, such as small samples and the absence of inferential tests, are threats to the strength of the conclusion, not evidence that the conclusion is built into the inputs by construction. Consequently, no circular step can be identified under the specified criteria.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No parameters were fitted to data; the analysis is descriptive. The main assumptions are imported from the cited literature and the study design choices.

assumptions (4)
  • domain assumption Comprehensibility and trust are the two relevant dimensions of explainability for this comparison.
    Adopted from Atf and Lewis 2023 in Section 2; the entire survey is built around these two metrics.
  • domain assumption Five-point Likert self-reports validly capture comprehensibility and trust.
    Section 3.2 uses Likert ratings without validation against behavioral measures or objective comprehension tests.
  • domain assumption The four explanation formats adequately represent the explanation types relevant to a mastitis decision support system.
    Section 3.1 selects classes from Vilone and Longo 2021; no prior task analysis with farmers is used to justify the coverage.
  • domain assumption Software providers can meaningfully predict farmer preferences when asked to do so.
    RQ1 and RQ2 assume the providers' predictions are interpretable and meaningful, and the study then checks accuracy against farmer ratings.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Perspectives on Explanation Formats From Two Stakeholder Groups in Germany: Software Providers and Dairy Farmers." pith.science (2026). https://pith.science/paper/PUBBM6KT

@misc{pith2026250611665,
  author       = {Pith},
  title        = {Pith review of: Perspectives on Explanation Formats From Two Stakeholder Groups in Germany: Software Providers and Dairy Farmers},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PUBBM6KT}},
  note         = {Machine review of arXiv:2506.11665}
}
read the original abstract

This paper examines the views of software providers in the German dairy industry with regard to dairy farmers' needs for explanation of digital decision support systems. The study is based on mastitis detection in dairy cows using a hypothetical herd management system. We designed four exemplary explanation formats for mastitis assessments with different types of presentation (textual, rule-based, herd comparison, and time series). In our previous study, 14 dairy farmers in Germany had rated these formats in terms of comprehensibility and the trust they would have in a system providing each format. In this study, we repeat the survey with 13 software providers active in the German dairy industry. We ask them how well they think the formats would be received by farmers. We hypothesized that there may be discrepancies between the views of both groups that are worth investigating, partly to find reasons for the reluctance to adopt digital systems. A comparison of the feedback from both groups supports the hypothesis and calls for further investigation. The results show that software providers tend to make assumptions about farmers' preferences that are not necessarily accurate. Our study, although not representative due to the small sample size, highlights the potential benefits of a thorough user requirements analysis (farmers' needs) to improve software adaptation and user acceptance.

Figures

Figures reproduced from arXiv: 2506.11665 by the authors.

Figure 1
Figure 1. Research questions and related metrics comparison a particularly useful tool. Overall, however, in the farmer survey, no format stood out significantly from the others. All formats received positive and negative ratings, and for each format there were farmers who saw value in it. The conclusion was that a good decision support system should offer several explanation formats and allow farmers to switch between them d… view at source ↗
Figure 2
Figure 2. Four explanation formats (translated from the original German survey [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Quantitative analysis results for the textual format, rule [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

30 extracted references · 27 canonical work pages

  1. [1]

    Interpretable machine learn- ing in healthcare

    [Ahmadet al., 2018 ] Muhammad Aurangzeb Ahmad, Carly Eckert, and Ankur Teredesai. Interpretable machine learn- ing in healthcare. InProceedings of the 2018 ACM in- ternational conference on bioinformatics, computational biology, and health informatics, pages 559–560,

  2. [9]

    Towards a rigorous science of interpretable machine learning.arXiv preprint arXiv:1702.08608,

    [Doshi-Velez and Kim, 2017] Finale Doshi-Velez and Been Kim. Towards a rigorous science of interpretable machine learning.arXiv preprint arXiv:1702.08608,

  3. [11]

    Exploring explainability formats to aid decision-making in dairy farming systems

    [Girmayet al., 2024 ] Mengisti Berihu Girmay, Felix M¨ohrle, and Jens Henningsen. Exploring explainability formats to aid decision-making in dairy farming systems. In44. GIL-Jahrestagung, Biodiversit ¨at f ¨ordern durch digitale Landwirtschaft, pages 269–274. Gesellschaft f ¨ur Informatik eV ,

  4. [12]

    Trustworthy versus explainable ai in autonomous vessels

    [Glomsrudet al., 2019 ] Jon Arne Glomsrud, Andr ´e Ødeg˚ardstuen, Asun Lera St Clair, and Øyvind Smo- geli. Trustworthy versus explainable ai in autonomous vessels. InProceedings of the International Seminar on Safety and Security of Autonomous Vessels (ISSAV) and European STAMP Workshop and Conference (ESWC), volume 37,

  5. [14]

    The potential of explainable artificial intel- ligence in precision livestock farming

    [Hoxhallariet al., 2022 ] K Hoxhallari, W Purcell, and T Neubauer. The potential of explainable artificial intel- ligence in precision livestock farming

  6. [15]

    Explainability as a non-functional require- ment

    [K¨ohlet al., 2019 ] Maximilian A K ¨ohl, Kevin Baum, Markus Langer, Daniel Oster, Timo Speith, and Dimitri Bohlender. Explainability as a non-functional require- ment. In2019 IEEE 27th International Requirements Engineering Conference (RE), pages 363–368. IEEE,

  7. [16]

    Ex- plainable ai for safe and trustworthy autonomous driving: A systematic review.arXiv preprint arXiv:2402.10086,

    [Kuznietsovet al., 2024 ] Anton Kuznietsov, Balint Gyevnar, Cheng Wang, Steven Peters, and Stefano V Albrecht. Ex- plainable ai for safe and trustworthy autonomous driving: A systematic review.arXiv preprint arXiv:2402.10086,

  8. [17]

    The mythos of model interpretability: In machine learning, the concept of in- terpretability is both important and slippery.Queue, 16(3):31–57,

    [Lipton, 2018] Zachary C Lipton. The mythos of model interpretability: In machine learning, the concept of in- terpretability is both important and slippery.Queue, 16(3):31–57,

Show all 30 references
  1. [18]

    Explanation in artificial intelli- gence: Insights from the social sciences.Artificial intelli- gence, 267:1–38,

    [Miller, 2019] Tim Miller. Explanation in artificial intelli- gence: Insights from the social sciences.Artificial intelli- gence, 267:1–38,

  2. [19]

    [Mohr and K¨uhl, 2021] Svenja Mohr and Rainer K ¨uhl. Ac- ceptance of artificial intelligence in german agriculture: an application of the technology acceptance model and the theory of planned behavior.Precision Agriculture, 22(6):1816–1844,

  3. [20]

    Desiderata for explainable ai in statistical production systems of the european central bank

    [Navarroet al., 2021 ] Carlos Mougan Navarro, Georgios Kanellos, and Thomas Gottron. Desiderata for explainable ai in statistical production systems of the european central bank. InJoint European Conference on Machine Learning and Knowledge Discovery in Databases, pages 575–59...

  4. [21]

    Understanding the public attitudi- nal acceptance of digital farming technologies: a nation- wide survey in germany.Agriculture and Human Values, 38(1):107–128,

    [Pfeifferet al., 2021 ] Johanna Pfeiffer, Andreas Gabriel, and Markus Gandorfer. Understanding the public attitudi- nal acceptance of digital farming technologies: a nation- wide survey in germany.Agriculture and Human Values, 38(1):107–128,

  5. [22]

    Stakeholders in explainable ai.arXiv preprint arXiv:1810.00184,

    [Preeceet al., 2018 ] Alun Preece, Dan Harborne, Dave Braines, Richard Tomsett, and Supriyo Chakraborty. Stakeholders in explainable ai.arXiv preprint arXiv:1810.00184,

  6. [23]

    Dissemi- nation of precision farming in germany: acceptance, adop- tion, obstacles, knowledge transfer and training activities

    [Reichardtet al., 2009 ] Maike Reichardt, Carsten J ¨urgens, Ulrike Kl¨oble, Joachim H¨uter, and Klaus Moser. Dissemi- nation of precision farming in germany: acceptance, adop- tion, obstacles, knowledge transfer and training activities. Precision Agriculture, 10:525–545,

  7. [25]

    Schoonderwoerd, Wiard Jorritsma, Mark A

    [Schoonderwoerdet al., 2021 ] Tjeerd A.J. Schoonderwoerd, Wiard Jorritsma, Mark A. Neerincx, and Karel van den Bosch. Human-centered xai: Developing design pat- terns for explanations of clinical decision support sys- tems.International Journal of Human-Computer Studies, 154:102684,

  8. [26]

    The effects of explainability and causability on perception, trust, and acceptance: Implica- tions for explainable ai.International Journal of Human- Computer Studies, 146:102551,

    [Shin, 2021] Donghee Shin. The effects of explainability and causability on perception, trust, and acceptance: Implica- tions for explainable ai.International Journal of Human- Computer Studies, 146:102551,

  9. [27]

    Evalu- ating xai: A comparison of rule-based and example-based explanations.Artificial Intelligence, 291:103404,

    [van der Waaet al., 2021 ] Jasper van der Waa, Elisabeth Nieuwburg, Anita Cremers, and Mark Neerincx. Evalu- ating xai: A comparison of rule-based and example-based explanations.Artificial Intelligence, 291:103404,

  10. [28]

    How to choose an explainability method? towards a me- thodical implementation of xai in practice

    [Vermeireet al., 2021 ] Tom Vermeire, Thibault Laugel, Xavier Renard, David Martens, and Marcin Detyniecki. How to choose an explainability method? towards a me- thodical implementation of xai in practice. InJoint Eu- ropean Conference on Machine Learning and Knowledge Discove...

  11. [29]

    Classification of explainable artificial intelligence meth- ods through their output formats.Machine Learning and Knowledge Extraction, 3(3):615–661,

    [Vilone and Longo, 2021] Giulia Vilone and Luca Longo. Classification of explainable artificial intelligence meth- ods through their output formats.Machine Learning and Knowledge Extraction, 3(3):615–661,

  12. [30]

    Springer Science & Business Media,

    [Wohlinet al., 2012 ] Claes Wohlin, Per Runeson, Martin H¨ost, Magnus C Ohlsson, Bj ¨orn Regnell, and An- ders Wessl´en.Experimentation in software engineering. Springer Science & Business Media,

  13. [1994]

    Analyzing and assessing ex- plainable ai models for smart agriculture environments

    [Cartolanoet al., 2024 ] Andrea Cartolano, Alfredo Cuz- zocrea, and Giovanni Pilato. Analyzing and assessing ex- plainable ai models for smart agriculture environments. Multimedia Tools and Applications, pages 1–22,

  14. [2009]

    Explainable artificial intelli- gence and interpretable machine learning for agricultural data analysis.Artificial Intelligence in Agriculture, 6:257– 265,

    [Ryo, 2022] Masahiro Ryo. Explainable artificial intelli- gence and interpretable machine learning for agricultural data analysis.Artificial Intelligence in Agriculture, 6:257– 265,

  15. [2017]

    Adoption of digital technologies in agricul- ture—an inventory in a european small-scale farming re- gion.Precision Agriculture, 24(1):68–91,

    [Gabriel and Gandorfer, 2023] Andreas Gabriel and Markus Gandorfer. Adoption of digital technologies in agricul- ture—an inventory in a european small-scale farming re- gion.Precision Agriculture, 24(1):68–91,

  16. [2018]

    Hu- man centricity in the relationship between explainability and trust in ai.IEEE Technology and Society Magazine, 42(4):66–76,

    [Atf and Lewis, 2023] Zahra Atf and Peter R Lewis. Hu- man centricity in the relationship between explainability and trust in ai.IEEE Technology and Society Magazine, 42(4):66–76,

  17. [2019]

    Can requirements engineering support explainable artificial intelligence? towards a user-centric approach for explainability requirements

    7 [Habibaet al., 2022 ] Umm-E Habiba, Justus Bogner, and Stefan Wagner. Can requirements engineering support explainable artificial intelligence? towards a user-centric approach for explainability requirements. In2022 IEEE 30th International Requirements Engineering Conference...

  18. [2020]

    The goal question met- ric approach.Encyclopedia of software engineering, pages 528–532,

    [Caldiera and Rombach, 1994] Victor R Basili1 Gianluigi Caldiera and H Dieter Rombach. The goal question met- ric approach.Encyclopedia of software engineering, pages 528–532,

  19. [2021]

    Combin- ing machine learning and simulation modelling for better predictions of crop yield and farmer income

    [Bergeret al., 2020 ] Thomas Berger, A Bernardi, D Martini, A M¨unzberg, J Parussis, T Streck, and C Troost. Combin- ing machine learning and simulation modelling for better predictions of crop yield and farmer income. InProceed- ings 10th International Congress on Environment...

  20. [2022]

    [D¨orr and Nachtmann, 2022] J¨org D¨orr and Matthias Nacht- mann.Handbook Digital Farming

    Association for Computing Machinery. [D¨orr and Nachtmann, 2022] J¨org D¨orr and Matthias Nacht- mann.Handbook Digital Farming. Springer,

  21. [2023]

    Explainable ai (xai) models applied to planning in financial markets

    [Benhamouet al., 2021 ] Eric Benhamou, Jean-Jacques Ohana, David Saltiel, and Beatrice Guez. Explainable ai (xai) models applied to planning in financial markets

  22. [2024]

    how can we develop explain- able systems? insights from a literature review and an in- terview study

    [Chazetteet al., 2022 ] Larissa Chazette, Jil Kl ¨under, Merve Balci, and Kurt Schneider. how can we develop explain- able systems? insights from a literature review and an in- terview study. InProceedings of the International Confer- ence on Software and System Processes and ...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.