Pith. sign in

REVIEW 4 major objections 7 minor 5 references

Good intentions, unintended consequences: exploring forecasting harms

T0 review · 4 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper claims forecasting has its own four-type harm taxonomy, and that even accurate forecasts can harm.

desk verdict Useful empirical map of forecasting harms, but the 'emergent taxonomy' claim overstates what the deductive method can support. read the letter →

arxiv 2411.16531 v3 pith:A3UYJIHP submitted 2024-11-25 stat.OT

classification stat.OT
keywords forecastingharmtaxonomytimeseriesethicsresponsibleAIinterviewsmitigation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that time series forecasting creates harms distinct from the better-documented harms of machine learning classification, and it offers a taxonomy of four harm types: risk of injury, denial of consequential services, infringement on human rights, and erosion of social and democratic structures. The taxonomy is built from 21 interviews with forecasters and researchers in retail, healthcare, humanitarian response, government, and software, analyzed with both human-led coding and AI-assisted thematic analysis. The paper's central insight is that forecast harms arise not only from inaccurate predictions but also from accurate ones, and not only from malicious actors but from well-intentioned forecasting workflows. If the taxonomy holds, organizations can use it, together with the paper's ten mitigation strategies, to identify and reduce harm before a forecast is published. The paper also sets out a research agenda for responsible forecasting, including fairness metrics for numerical prediction, model cards, and stronger standards of care.

What carries the argument

The machinery is a two-part classification system. The first part is a harm typology with four top-level categories and eight subcategories, adapted from the Microsoft Azure technology-harm framework to forecasting contexts through interview coding. The second part is a forecast harm matrix that crosses forecast accuracy with intent, separating morally distinct harm scenarios. The typology identifies where harm can occur, while the matrix identifies who bears responsibility and what kind of ethical scrutiny applies. The supporting analytical device is a semi-structured interview protocol combined with human-led inductive coding and AI-assisted theme extraction, where the AI's proposed categories and quotes were manually verified against the transcripts.

What would settle it

A concrete test would be to collect a broad sample of documented forecasting incidents across many sectors and check whether all reported harms fall into the four categories; if, for instance, a clearly identified forecasting harm involving environmental damage or privacy cannot be assigned to any listed subcategory, or if a sector-specific harm such as election manipulation is fundamentally different in kind, the taxonomy fails to generalize.

Watch

Extended reading notes

Core claim

The central claim is that forecasting produces a distinct set of harms that existing machine learning harm taxonomies do not capture. Using the Microsoft Azure responsible-innovation taxonomy as a deductive starting point, the authors code 21 expert interviews and arrive at four harm types: risk of injury, denial of consequential services, infringement on human rights, and erosion of social and democratic structures. They further organize harm along two dimensions, forecast accuracy and intent, producing four scenarios: unintentional harm from inaccurate forecasts, unintentional harm from accurate forecasts, intentional harm from inaccurate forecasts, and intentional harm from accurate forecasts. A key point is that accurate forecasts can still harm, for example by changing behavior, causing panic, or enabling malicious actors, and that responsibility for harm extends beyond the forecaster to decision-makers and audiences.

Load-bearing premise

The taxonomy's generality rests on the assumption that forecasting workflows are similar enough across sectors for one set of harm categories to apply; if retail, healthcare, humanitarian, and election forecasting differ substantially in how they are built and used, some domain-specific harms will be missed.

Editorial extensions

If this is right

  • Organizations that publish forecasts can use the four-type taxonomy to audit their own forecasting workflow before release, rather than only after a failure.
  • The ten mitigation strategies, from forecasting model cards and uncertainty communication to access control and bias audits, give concrete steps that forecasters and decision-makers can adopt.
  • High-risk domains, especially healthcare, humanitarian response, and politics, warrant proportionally stronger scrutiny and safeguards under a risk-tier approach.
  • The accuracy-versus-intent matrix shows that improving forecast accuracy alone will not eliminate harm; communication, interpretation, and use matter just as much.
  • The proposed research agenda, including forecasting-specific fairness metrics, explainability, and contestability, could reshape how forecasters are trained and regulated.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension, not claimed by the paper, would be to apply the taxonomy to a large corpus of documented forecasting failures and check whether every reported harm maps to one of the four categories; a harm that fits none would show the taxonomy needs refinement.
  • The paper's harm matrix implies that forecasters could face moral or legal responsibility even for accurate forecasts when they foresee harmful misuse, a corollary the authors mention but do not develop in legal detail.
  • The taxonomy likely extends to adjacent predictive analytics tasks, since the underlying workflows are similar, even though the paper restricts its claims to time series forecasting.
  • The proposed forecasting model cards could be turned into a standardized reporting template and empirically evaluated in real deployments to see whether they change stakeholder understanding or behavior.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper reports a qualitative interview study of 21 forecasting practitioners and academics, combined with an LLM-assisted analysis, to identify harms specific to time series forecasting. It proposes a taxonomy of four main harm types (risk of injury, denial of consequential services, infringement on human rights, and erosion of social and democratic structures) with subcategories, a two-by-two matrix of harm by forecast accuracy and intent, ten mitigation strategies, and a research agenda. The paper claims that these four themes are 'emergent' from the interview data and that the taxonomy is forecasting-specific.

Significance. If the claimed forecasting-specific taxonomy were empirically established, it would fill a genuine gap: forecasting ethics is understudied compared with other ML harms, and the paper collects rich interview material spanning retail, public health, humanitarian, government, and academic settings. The authors are transparent about their hybrid inductive-deductive design and about the simplifying assumption of a stable forecasting workflow. The mitigation strategies in Section 5.3 and the research agenda in Section 6 are plausible and practically useful. However, the central claim of a novel, emergent, forecasting-specific taxonomy is not supported by the reported method: the four top-level categories and most subcategories coincide with the Microsoft Azure seed framework that was explicitly used to structure both the human and the LLM coding. As a result, the contribution is better characterized as an exploratory mapping of forecasting harms onto an existing generic taxonomy, not as evidence that forecasting gives rise to distinct harm categories.

major comments (4)
  1. [Sections 3.1, 3.3, Table 2, Figure 2, Section 7] The central claim that the taxonomy consists of 'four emergent themes' (Section 7) is not supported by the reported method. Section 3.1 states that the analysis 'began with a deductive analysis, focusing on the categories of harms defined in the previously-mentioned Microsoft Azure framework'; Section 3.3 says the human-led thematic analysis was 'based on the topics and harm categories outlined in the Microsoft Azure framework'; and Table 2 shows that the LLM was given the Azure table and asked to categorize responses into it, with the Azure table supplied 'to give you some ideas' for the new framework. Since Figure 2's four top-level types and most subcategories are identical to the Azure seed, the same output would be expected if the interviews merely supplied forecasting examples for a pre-existing generic framework. The manuscript should either present the contribution as a mapping and contextualization of the Azure taxonomy to forecasting, or provide explicit evidence of interview-driven categories that do not derive from the seed, such as categories that the human or LLM analysis identified as not fitting the Azure table.
  2. [Section 3.3] The qualitative evidence base is reported incompletely. Only the first author conducted the human-led coding in NVivo, with no inter-coder reliability, no saturation analysis, and no description of how disagreements were resolved. The full transcripts are only promised via a GitHub repository rather than included or linked, which makes it difficult to verify the coding. This underdetermines the claim of a 'systematic' identification of harms and leaves open the possibility that the coding merely reproduced the Azure categories. Please report coding procedures, provide a working link to the transcript repository, and give evidence of coding consistency or saturation, or explicitly state that this is an exploratory rather than confirmatory study.
  3. [Section 2.3] The simplifying assumption that 'the forecasting workflow is relatively stable across forecast types and applications' is load-bearing for the generalization of the taxonomy, but it is not tested or defended. Retail replenishment, public-health surveillance, humanitarian response, and election polling involve different stakeholders, feedback loops, publication norms, and decision horizons. If workflows differ substantially across these sectors, a single harm typology may miss domain-specific harms. The paper should either qualify the scope of the taxonomy to the domains represented in the interviews or provide evidence from the interviews that the experts converge on a common workflow description.
  4. [Section 2.2 vs. Section 4.4.2] There is an internal inconsistency between the claim in Section 2.2 that representational harms 'seem less relevant' to forecasting and the findings in Section 4.4.2, which identify 'stereotype reinforcement and loss of representation' as forecasting harms. The latter are representational harms in the sense of the cited literature (Barocas et al. 2017; Suresh & Guttag 2021). This tension is material because the paper argues that forecasting differs from other ML applications partly in this respect. Please reconcile the framing in Section 2.2 with the empirical findings, or explain why the interview-identified harms are not representational in the sense used there.
minor comments (7)
  1. [Section 3.1] The statement 'full interview transcripts will be made accessible via a GitHub repository' lacks a link or repository identifier; please provide one in the revised version.
  2. [Table 3] The participant table has a typo: 'Goverment' should be 'Government'.
  3. [Section 5.1] The city name 'L'Alquila' should be 'L'Aquila'.
  4. [Section 5.3] The sentence 'Model cards could a viable solution' should read 'Model cards could be a viable solution'.
  5. [Section 4.3.2] Placing 'Environmental impact' under 'Infringement on human rights' is non-obvious; please add a sentence explaining the rights-based framing or move the category to a more appropriate parent.
  6. [Figure 3] The two-by-two harm matrix is introduced without stating whether it is an empirically derived finding or an analytic synthesis; please state its status explicitly.
  7. [Section 6] Several research-agenda items begin with 'Many interviewers noted' or 'Many interviewers noted that managers...' but provide no counts or representative quotes; consider supporting these statements with evidence or softening the phrasing.

Circularity Check

1 steps flagged · score 6.0 of 10

The top-level taxonomy reproduces the Microsoft Azure seed framework's four harm categories, so the 'four emergent themes' are the input categories re-issued with forecasting examples.

  1. renaming known result [Section 3.2 (Table 2), Section 3.3, Section 4 (Figure 2); cf. Appendix C]
    "Before supplying Claude the transcripts, we provided it with the Microsoft Azure framework as an example of a harm typology. We then requested Claude to categorize the harms mentioned in the interviews into the Microsoft Azure framework categories, providing relevant quotes from the interviews. Finally, we requested Claude to create a new framework of forecasting harms, along with relevant quotes."

    The output taxonomy's four top-level types in Section 4 are exactly the four Microsoft Azure harm types used as the coding frame: risk of injury, denial of consequential services, infringement on human rights, and erosion of social and democratic structures. The LLM was first asked to categorize interview content into the Azure table, and the human-led thematic analysis was 'based on the topics and harm categories outlined in the Microsoft Azure framework.' Therefore the four 'emergent themes' are the seed categories reappearing by construction, not categories independently derived from the interviews.

full rationale

The paper is transparent about beginning deductively with the Microsoft Azure taxonomy, so this is not a case of hidden input. The circularity is structural: the same four harm categories were supplied to both analysis channels as the coding scheme, and they reappear as the paper's central 'four emergent themes.' This makes the central claim that the taxonomy is forecasting-specific and emergent from the interviews unsupported by the reported method. The interview quotes and adapted forecasting examples are genuine additions, so the finding is not wholly vacuous, but the top-level classification does reduce to the seed framework by construction. No load-bearing self-citation is present; the issue is seeding of the analytical frame, not author self-citation.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

No physical or mathematical entities are postulated; the taxonomy and mitigation list are qualitative constructs built from interviews and the Azure framework. The ledger records the domain assumptions the taxonomy depends on.

assumptions (7)
  • domain assumption Forecasting workflow is relatively stable across sectors and applications.
    Section 2.3 states this simplifying assumption; the taxonomy generalizes harms only if workflows are comparable.
  • domain assumption Harm can be treated as a cluster concept and defined as unjustly or wrongly setting back an interest.
    Section 2.1 adopts this definition; the taxonomy depends on this normative framing.
  • domain assumption Risk imposition and harm are interchangeable for the analysis.
    Section 2.4 states 'we treat them as interchangeable for now'; this collapses a contested distinction.
  • domain assumption Purposive sample of 21 experts provides sufficient breadth to identify common forecasting harms.
    Section 3.1 describes purposive sampling; no saturation analysis is reported.
  • domain assumption Interviewee self-reports accurately describe real forecasting harms.
    Section 3.1 and Appendix A rely on retrospective self-report without independent verification.
  • domain assumption Microsoft Azure harm taxonomy is an appropriate a priori frame for forecasting harms.
    Section 3.1 begins deductively from this framework, which shapes the top-level categories.
  • domain assumption LLM-assisted coding with Claude 2.1 produces reliable theme extraction after human validation.
    Section 3.2; no systematic validation beyond first author inspection is documented.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Good intentions, unintended consequences: exploring forecasting harms." pith.science (2026). https://pith.science/paper/A3UYJIHP

@misc{pith2026241116531,
  author       = {Pith},
  title        = {Pith review of: Good intentions, unintended consequences: exploring forecasting harms},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/A3UYJIHP}},
  note         = {Machine review of arXiv:2411.16531}
}
read the original abstract

Organizations worldwide that rely on data-driven approaches regularly employ forecasting methods to enhance their planning and decision-making processes. While extensive research has examined the harms associated with traditional machine learning applications, relatively little attention has been given to the ethical implications of time series forecasting. However, forecasting presents distinct ethical challenges due to its diverse organizational applications, varied objectives, and unique data processing, model development, and evaluation workflows. These distinctions complicate the direct application of existing machine learning harm taxonomies to common forecasting scenarios. To address this gap, we conduct multiple interviews with industry experts and academic researchers, systematically identifying and analyzing underexplored domains, use cases, and potential risks associated with forecasting. Our objective is to develop a novel taxonomy of forecasting-specific harms. Drawing inspiration from Microsoft Azure taxonomy for responsible innovation, we integrate a human-led inductive coding approach with AI-driven analysis to extract key categories of harm in forecasting. This taxonomy aims to support researchers and practitioners by fostering ethical reflection on their decision-making throughout the forecasting process. Additionally, we seek to establish a research agenda focused on identifying measures to mitigate potential harms in forecasting. By highlighting unique risks within forecasting, our work contributes to the broader discourse on machine learning ethics.

Figures

Figures reproduced from arXiv: 2411.16531 by the authors.

Figure 1
Figure 1. LLM use flowchart [PITH_FULL_IMAGE:figures/full_fig_p009_1.png] view at source ↗
Figure 2
Figure 2. Types of harm in forecasting based on our interviews. 4.1 Risk of injury This section explores how forecasting can potentially harm people, create hazardous environments, or contribute to significant emotional and psychological distress. 4.1.1 Physical injury Inadequate fail-safes Forecasts of natural disasters can have dire consequences. In the case of hurricane forecasts: “There have been instances in the last few… view at source ↗
Figure 3
Figure 3. Forecast harm matrix Unintentional harm from accurate forecasts Accurate forecasts, while informative, can unintentionally inflict harm due to their impact on human behavior and emotional well-being. For instance, a healthcare forecast predicting the progression of an inevitable illness or genetic condition may cause severe psychological distress, creating ethical dilemmas for individuals and their families and ofte… view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Mitigation strategies extracted from interview analysis 1. Tailor forecasts to the decision context Interviewees emphasized the importance of tailoring forecasts to the specific needs and contexts of decision-makers. Forecasts that are not aligned with decision-making …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

5 extracted references · 5 canonical work pages

  1. [1]

    & Van Den Hoven, J

    Aizenberg, E. & Van Den Hoven, J. (2020), ‘Designing for human rights in AI’, Big Data & Society 7(2), 2053951720949566. Arrieta, A. B., D´ıaz-Rodr´ıguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., Garc´ıa, S., Gil-L´opez, S., Molina, D., Benjamins, R. et al. (2020), ‘Explainable artificial intelligence (XAI): Concepts, taxonomies, opportuniti...

  2. [8]

    People tend to only look at the deterministic forecast and discard the uncertainty bounds that are around it which could be quite large

    Based on your experience, what are some common types of harm that can occur in the forecasting workflow prior to publishing, including stages like data collection, model selection, and accuracy evaluation? A Interview protocol 26 B Summary of participants Table 3: Summary of participants’ domain, expertise and interview length Code Perspective Role Sector...

  3. [29]

    (2009), Harming as causing harm, in M

    Harman, E. (2009), Harming as causing harm, in M. Roberts & D. Wasserman, eds, ‘Harming future persons: Ethics, genetics and the nonidentity problem’, Springer, pp. 137–154. Hayenhjelm, M. & Wolff, J. (2012), ‘The moral problem of risk impositions: A survey of the literature’,European Journal of Philosophy 20, E26–E51. Hewamalage, H., Bergmeir, C. & Banda...

  4. [963]

    & Davis, A

    Fischhoff, B. & Davis, A. L. (2014), ‘Communicating scientific uncertainty’, Proceedings of the National Academy of Sciences 111(supplement 4), 13664–13671. Floridi, L., Cowls, J., Beltrametti, M., Chatila, R., Chazerand, P., Dignum, V ., Luetge, C., Madelin, R., Pagallo, U., Rossi, F. et al. (2021), ‘An ethical framework for a good AI society: Opportunit...

  5. [1725]

    (2021), ‘Fighting artificial intelligence battles: Operational concepts for future AI-enabled wars’, Network 4(20), 1–100

    Layton, P. (2021), ‘Fighting artificial intelligence battles: Operational concepts for future AI-enabled wars’, Network 4(20), 1–100. Makridakis, S., Hyndman, R. J. & Petropoulos, F. (2020), ‘Forecasting in social settings: The state of the art’, International Journal of Forecasting36(1), 15–28. Microsoft (2023), ‘Types of harm’. Accessed on 2024-10-28. U...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.