{"id":"541b409e-aaf6-4ea5-a768-ef2a347e866b","arxiv_id":"2411.16531","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Forecasting harms can be grouped into four main types, and the paper adds ten mitigation strategies and a research agenda for responsible forecasting.","lead":"Researchers interviewed 21 forecasting experts and built a catalog of ways forecasts can hurt people, from panic and wasted resources to privacy invasion and election effects. The catalog gives forecasters and policymakers a shared vocabulary for spotting harm before it happens and for designing safer forecasting systems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The four top-level harm categories in Figure 2 are identical to the Microsoft Azure seed framework, so the paper's claim that these are 'emergent' forecasting-specific themes is not supported by the reported method.","rationale":"Good-faith reading: the paper is an exploratory, interview-based contribution to forecasting ethics. It has real strengths: 21 expert interviews across sectors, a transparent hybrid human/LLM procedure, concrete quotes, a useful harm-intent/accuracy matrix (Figure 3), and ten practical mitigation strategies. The central claim, however, is that the authors identified a novel taxonomy of forecasting-specific harms. That claim depends on the harms being grounded in the interview data rather than imported from the framework used to seed the analysis. The reported method makes that grounding impossible to verify: both the human coder and the LLM were given Azure's categories before coding, the LLM was explicitly asked to classify harms into those categories, and the final top-level taxonomy is identical to Azure's. Calling the result 'emergent' in Section 7 is therefore internally inconsistent with the method as described. The reader's weakest_assumption focused on cross-domain workflow stability and generalizability. That is a legitimate secondary concern, but it presupposes that the taxonomy is genuinely data-derived; the deductive-seeding issue attacks that precondition. If the taxonomy is a relabeling of Azure, then generalizability to forecasting domains is not the main problem—the lack of forecasting-specificity is. The reader's conditional verdict already flags the deductive role of Azure, so my concern is largely convergent rather than a new rejection. I do not think the paper should be rejected: the taxonomy may still be useful, and the promised transcripts plus independent coding would settle whether the categories are artifacts of the seed. Hence I recommend keeping the reader's CONDITIONAL verdict and adding this specific validation requirement. Agreement with the reader: partial—they identified the same deductive-role issue in the rationale, but their stated weakest assumption was the workflow-stability assumption, which I see as secondary to the emergence/specificity problem.","tokens_in":33711,"tokens_out":5762,"duration_ms":54368,"concrete_test":"Once the promised anonymized transcripts are released (GitHub repository cited in Section 3.1), run an independent, purely inductive thematic analysis: at least two coders who have never seen the Microsoft Azure type-of-harm table code all 21 transcripts for harms, without using any predefined categories, and then cluster the codes into top-level themes. Compare this emergent structure to Figure 2. If the independent analysis reproduces the same four top-level types (with high intercoder agreement), the taxonomy is robust to the seed concern. If it yields different top-level categories—for example, separating forecast-communication harms, self-fulfilling-prophesy effects, resource-misallocation harms, or domain-specific environmental harms—then Figure 2's structure was at least partly imposed by the Azure framework.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.1 states the analysis 'began with a deductive analysis, focusing on the categories of harms defined in the previously-mentioned Microsoft Azure framework.' Section 3.3 says the human-led thematic analysis was 'based on the topics and harm categories outlined in the Microsoft Azure framework,' and the LLM prompt (Table 2) explicitly gave Claude the Azure table, asked it to match responses to that table, and only then to create a new framework using the Azure table 'to give you some ideas.' Figure 2's top-level types—risk of injury, denial of consequential services, infringement on human rights, erosion of social and democratic structures—are exactly the four Azure types, and most subcategories (physical injury, opportunity loss, privacy loss, manipulation, stereotype reinforcement) map directly onto Azure subcategories. Given this near-total overlap with the seed, Section 7's description of 'four emergent themes' conflates categories supplied by the deductive starting point with categories that emerged from interview data. The central claim of a novel, forecasting-specific taxonomy therefore rests on a distinction the reported method cannot support: the same result would be expected if the interviews merely provided forecasting examples for a pre-existing generic harm framework. This is a methodological circularity risk, not a matter of authorial intent; the paper is transparent about drawing inspiration from Azure, but the 'emergent/inductive' language overstates what the design can establish. If the taxonomy is not demonstrably forecasting-specific, its main contribution reduces to an application of Azure's generic categories to forecasting workflows.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a qualitative interview study of 21 forecasting practitioners and academics, combined with an LLM-assisted analysis, to identify harms specific to time series forecasting. It proposes a taxonomy of four main harm types (risk of injury, denial of consequential services, infringement on human rights, and erosion of social and democratic structures) with subcategories, a two-by-two matrix of harm by forecast accuracy and intent, ten mitigation strategies, and a research agenda. The paper claims that these four themes are 'emergent' from the interview data and that the taxonomy is forecasting-specific.","tokens_in":33896,"tokens_out":4911,"duration_ms":46294,"significance":"If the claimed forecasting-specific taxonomy were empirically established, it would fill a genuine gap: forecasting ethics is understudied compared with other ML harms, and the paper collects rich interview material spanning retail, public health, humanitarian, government, and academic settings. The authors are transparent about their hybrid inductive-deductive design and about the simplifying assumption of a stable forecasting workflow. The mitigation strategies in Section 5.3 and the research agenda in Section 6 are plausible and practically useful. However, the central claim of a novel, emergent, forecasting-specific taxonomy is not supported by the reported method: the four top-level categories and most subcategories coincide with the Microsoft Azure seed framework that was explicitly used to structure both the human and the LLM coding. As a result, the contribution is better characterized as an exploratory mapping of forecasting harms onto an existing generic taxonomy, not as evidence that forecasting gives rise to distinct harm categories.","major_comments":[{"comment":"The central claim that the taxonomy consists of 'four emergent themes' (Section 7) is not supported by the reported method. Section 3.1 states that the analysis 'began with a deductive analysis, focusing on the categories of harms defined in the previously-mentioned Microsoft Azure framework'; Section 3.3 says the human-led thematic analysis was 'based on the topics and harm categories outlined in the Microsoft Azure framework'; and Table 2 shows that the LLM was given the Azure table and asked to categorize responses into it, with the Azure table supplied 'to give you some ideas' for the new framework. Since Figure 2's four top-level types and most subcategories are identical to the Azure seed, the same output would be expected if the interviews merely supplied forecasting examples for a pre-existing generic framework. The manuscript should either present the contribution as a mapping and contextualization of the Azure taxonomy to forecasting, or provide explicit evidence of interview-driven categories that do not derive from the seed, such as categories that the human or LLM analysis identified as not fitting the Azure table.","section":"Sections 3.1, 3.3, Table 2, Figure 2, Section 7"},{"comment":"The qualitative evidence base is reported incompletely. Only the first author conducted the human-led coding in NVivo, with no inter-coder reliability, no saturation analysis, and no description of how disagreements were resolved. The full transcripts are only promised via a GitHub repository rather than included or linked, which makes it difficult to verify the coding. This underdetermines the claim of a 'systematic' identification of harms and leaves open the possibility that the coding merely reproduced the Azure categories. Please report coding procedures, provide a working link to the transcript repository, and give evidence of coding consistency or saturation, or explicitly state that this is an exploratory rather than confirmatory study.","section":"Section 3.3"},{"comment":"The simplifying assumption that 'the forecasting workflow is relatively stable across forecast types and applications' is load-bearing for the generalization of the taxonomy, but it is not tested or defended. Retail replenishment, public-health surveillance, humanitarian response, and election polling involve different stakeholders, feedback loops, publication norms, and decision horizons. If workflows differ substantially across these sectors, a single harm typology may miss domain-specific harms. The paper should either qualify the scope of the taxonomy to the domains represented in the interviews or provide evidence from the interviews that the experts converge on a common workflow description.","section":"Section 2.3"},{"comment":"There is an internal inconsistency between the claim in Section 2.2 that representational harms 'seem less relevant' to forecasting and the findings in Section 4.4.2, which identify 'stereotype reinforcement and loss of representation' as forecasting harms. The latter are representational harms in the sense of the cited literature (Barocas et al. 2017; Suresh & Guttag 2021). This tension is material because the paper argues that forecasting differs from other ML applications partly in this respect. Please reconcile the framing in Section 2.2 with the empirical findings, or explain why the interview-identified harms are not representational in the sense used there.","section":"Section 2.2 vs. Section 4.4.2"}],"minor_comments":[{"comment":"The statement 'full interview transcripts will be made accessible via a GitHub repository' lacks a link or repository identifier; please provide one in the revised version.","section":"Section 3.1"},{"comment":"The participant table has a typo: 'Goverment' should be 'Government'.","section":"Table 3"},{"comment":"The city name 'L'Alquila' should be 'L'Aquila'.","section":"Section 5.1"},{"comment":"The sentence 'Model cards could a viable solution' should read 'Model cards could be a viable solution'.","section":"Section 5.3"},{"comment":"Placing 'Environmental impact' under 'Infringement on human rights' is non-obvious; please add a sentence explaining the rights-based framing or move the category to a more appropriate parent.","section":"Section 4.3.2"},{"comment":"The two-by-two harm matrix is introduced without stating whether it is an empirically derived finding or an analytic synthesis; please state its status explicitly.","section":"Figure 3"},{"comment":"Several research-agenda items begin with 'Many interviewers noted' or 'Many interviewers noted that managers...' but provide no counts or representative quotes; consider supporting these statements with evidence or softening the phrasing.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely suitable for a journal that welcomes qualitative work on the ethics and societal implications of statistical practice. The main issue is calibration of claims: the authors have done a transparent hybrid analysis, but the 'emergent' and 'forecasting-specific' claims outrun the evidence because the seed taxonomy is the output taxonomy. I would not reject, because the interview corpus and mitigation strategies are valuable and the circularity can be addressed by reframing the contribution as an exploratory mapping study supplemented by evidence of any genuinely new categories. The paper's own statements in Sections 2.3 and 3.1 acknowledge the deductive starting point, so the revision is achievable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth knowing: this is a competent qualitative study with real interview data, and it will be useful to forecasting practitioners and ethics researchers. The novelty, though, is smaller than the paper claims. The four top-level harm categories come straight from Microsoft Azure's framework; the interviews and the LLM were seeded with that framework, so calling the results 'four emergent themes' in Section 7 goes beyond what the method supports. That is the main weakness, and it is real.\n\nWhat the paper does well: 21 interviews across a genuinely wide range of sectors, with extensive quoting in the appendix. The forecasting-specific examples (inventory waste, family-planning supply disruptions, election polling bans, targeted aid misallocation) are the genuine contribution; they ground generic AI harm categories in concrete forecasting workflows. The harm matrix in Figure 3 (intent x accuracy) is a nice conceptual addition, and the ten mitigation strategies in 5.3 are practical and mostly actionable. The research agenda in Section 6, especially fair forecasting metrics and model cards, is sensible.\n\nSoft spots, in order of importance. First, the 'emergent' language is not defensible. Section 3.1 says the authors began deductively from Azure; Table 2 shows the LLM was given the Azure table, asked to match responses to it, and only then to build a new framework 'using the table... to give you some ideas'. The top-level types and most subcategories in Figure 2 are the Azure ones, so the design would produce this shape even if respondents merely supplied forecasting illustrations of a pre-existing generic taxonomy. That does not make the paper fraudulent—the authors are transparent about drawing inspiration from Azure—but it makes 'forecasting-specific taxonomy' an overclaim and the contribution more of a domain-mapped adaptation than a new framework.\n\nSecond, the method section reports no inter-coder reliability or saturation analysis, and the full transcripts are promised but not linked. For a qualitative paper claiming to consolidate themes, that is a reproducibility gap. Third, the simplifying assumption of a stable forecasting workflow across sectors is acknowledged but limits generalizability; retail replenishment and humanitarian migration forecasting may not share enough workflow structure for a single taxonomy to capture both well.\n\nBottom line: the paper is worth engaging, and the interview evidence is a useful resource. It does not, however, establish a de novo forecasting harm taxonomy, and the authors should revise the framing accordingly. I would send it to peer review, because the empirical material and the mitigation list have real value, but I would require the 'emergent' claims to be softened, the transcripts and codebook to be linked, and the deductive role of Azure to be front-and-center.","headline":"Useful empirical map of forecasting harms, but the 'emergent taxonomy' claim overstates what the deductive method can support.","tokens_in":34481,"tokens_out":2051,"would_cite":true,"duration_ms":22613,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims forecasting has its own four-type harm taxonomy, and that even accurate forecasts can harm.","keywords":["forecasting","harm","taxonomy","time series","ethics","responsible AI","interviews","mitigation"],"falsifier":"A concrete test would be to collect a broad sample of documented forecasting incidents across many sectors and check whether all reported harms fall into the four categories; if, for instance, a clearly identified forecasting harm involving environmental damage or privacy cannot be assigned to any listed subcategory, or if a sector-specific harm such as election manipulation is fundamentally different in kind, the taxonomy fails to generalize.","tokens_in":33472,"feed_emoji":"📉","tokens_out":5555,"duration_ms":56299,"temperature":0.7,"pith_summary":"This paper tries to establish that time series forecasting creates harms distinct from the better-documented harms of machine learning classification, and it offers a taxonomy of four harm types: risk of injury, denial of consequential services, infringement on human rights, and erosion of social and democratic structures. The taxonomy is built from 21 interviews with forecasters and researchers in retail, healthcare, humanitarian response, government, and software, analyzed with both human-led coding and AI-assisted thematic analysis. The paper's central insight is that forecast harms arise not only from inaccurate predictions but also from accurate ones, and not only from malicious actors but from well-intentioned forecasting workflows. If the taxonomy holds, organizations can use it, together with the paper's ten mitigation strategies, to identify and reduce harm before a forecast is published. The paper also sets out a research agenda for responsible forecasting, including fairness metrics for numerical prediction, model cards, and stronger standards of care.","feed_headline":"21 interviews map forecasting harms into four types","feed_subtitle":"A new taxonomy covers injury, service denial, rights harms, and democratic erosion, plus ten mitigations.","key_machinery":"The machinery is a two-part classification system. The first part is a harm typology with four top-level categories and eight subcategories, adapted from the Microsoft Azure technology-harm framework to forecasting contexts through interview coding. The second part is a forecast harm matrix that crosses forecast accuracy with intent, separating morally distinct harm scenarios. The typology identifies where harm can occur, while the matrix identifies who bears responsibility and what kind of ethical scrutiny applies. The supporting analytical device is a semi-structured interview protocol combined with human-led inductive coding and AI-assisted theme extraction, where the AI's proposed categories and quotes were manually verified against the transcripts.","core_discovery":"The central claim is that forecasting produces a distinct set of harms that existing machine learning harm taxonomies do not capture. Using the Microsoft Azure responsible-innovation taxonomy as a deductive starting point, the authors code 21 expert interviews and arrive at four harm types: risk of injury, denial of consequential services, infringement on human rights, and erosion of social and democratic structures. They further organize harm along two dimensions, forecast accuracy and intent, producing four scenarios: unintentional harm from inaccurate forecasts, unintentional harm from accurate forecasts, intentional harm from inaccurate forecasts, and intentional harm from accurate forecasts. A key point is that accurate forecasts can still harm, for example by changing behavior, causing panic, or enabling malicious actors, and that responsibility for harm extends beyond the forecaster to decision-makers and audiences.","pith_inferences":["A natural extension, not claimed by the paper, would be to apply the taxonomy to a large corpus of documented forecasting failures and check whether every reported harm maps to one of the four categories; a harm that fits none would show the taxonomy needs refinement.","The paper's harm matrix implies that forecasters could face moral or legal responsibility even for accurate forecasts when they foresee harmful misuse, a corollary the authors mention but do not develop in legal detail.","The taxonomy likely extends to adjacent predictive analytics tasks, since the underlying workflows are similar, even though the paper restricts its claims to time series forecasting.","The proposed forecasting model cards could be turned into a standardized reporting template and empirically evaluated in real deployments to see whether they change stakeholder understanding or behavior."],"forward_implications":["Organizations that publish forecasts can use the four-type taxonomy to audit their own forecasting workflow before release, rather than only after a failure.","The ten mitigation strategies, from forecasting model cards and uncertainty communication to access control and bias audits, give concrete steps that forecasters and decision-makers can adopt.","High-risk domains, especially healthcare, humanitarian response, and politics, warrant proportionally stronger scrutiny and safeguards under a risk-tier approach.","The accuracy-versus-intent matrix shows that improving forecast accuracy alone will not eliminate harm; communication, interpretation, and use matter just as much.","The proposed research agenda, including forecasting-specific fairness metrics, explainability, and contestability, could reshape how forecasters are trained and regulated."],"supporting_citations":[{"why":"Supplies the four harm categories that anchor the deductive coding and the final taxonomy.","marker":"Microsoft 2023"},{"why":"Provides the sociotechnical AI harm taxonomy that the paper argues is insufficient for forecasting.","marker":"Shelby et al. 2023"},{"why":"Supplies the philosophical definition of harm as the unjust defeating of an interest, which the paper adopts.","marker":"Feinberg 1984"},{"why":"Prior forecasting-specific ethics work in marine ecology that motivates the harm mitigation strategies.","marker":"Hobday et al. 2019"},{"why":"Frames harms across the machine learning lifecycle, informing the paper's discussion of representational and allocative harms in forecasting.","marker":"Suresh & Guttag 2021"},{"why":"Proposes model cards, which the paper adapts as a forecasting harm mitigation strategy.","marker":"Mitchell et al. 2019"}],"fun_headline_variants":["Forecasting harms: new taxonomy lists four distinct types","Accurate forecasts can harm too—study defines four types","Interview study maps forecasting-specific harms into four categories","Beyond ML: four harm types unique to time-series forecasting","21 interviews reveal four ways forecasts backfire"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The taxonomy's generality rests on the assumption that forecasting workflows are similar enough across sectors for one set of harm categories to apply; if retail, healthcare, humanitarian, and election forecasting differ substantially in how they are built and used, some domain-specific harms will be missed.","fun_headline_variants_meta":{"raw":{"variants":["Forecasting harms: new taxonomy lists four distinct types","Accurate forecasts can harm too—study defines four types","Interview study maps forecasting-specific harms into four categories","Beyond ML: four harm types unique to time-series forecasting","21 interviews reveal four ways forecasts backfire"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000623,"raw_usage":{"total_tokens":2865,"prompt_tokens":902,"completion_tokens":1963,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":1888}},"tokens_in":518,"tokens_out":1963,"duration_ms":14521,"temperature":1.0,"reasoning_tokens":1888,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:59:38.518259+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would be to collect a broad sample of documented forecasting incidents across many sectors and check whether all reported harms fall into the four categories; if, for instance, a clearly identified forecasting harm involving environmental damage or privacy cannot be assigned to any listed subcategory, or if a sector-specific harm such as election manipulation is fundamentally different in kind, the taxonomy fails to generalize.","supporting_citations":[],"review_version":1}