{"id":"4ef1671a-c486-49bd-bb53-a55ba2179b69","arxiv_id":"2501.04902","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In concurrent field trials, a regulator and an advocacy group verified the same AI tool's manure-spreading detections at similar rates, yet judged its value differently due to their organizational mandates.","lead":"A satellite AI model for detecting winter manure spreading was field-tested with a state regulator and an environmental advocacy group; both verified detections at similar rates, but they disagreed on the tool's usefulness because the regulator cared about clear legal violations while the group sought to document environmental risk. Most confirmed spreading (82%) fell outside existing regulations, illustrating how organizational context shapes AI deployment.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 82% 'regulatory cracks' figure and WDNR's 'few clear violations' assessment rest on unverifiable AFO attributions; if even a subset of those 26 events were actually CAFO applications, the organizational-context finding weakens.","rationale":"The reader identified the same soft spot I would: the AFO attribution is the load-bearing assumption. The paper's central claim is not just that organizations differ in goals, but that this difference—not model performance—explains their divergent assessments. That argument depends on WDNR seeing few clear violations; WDNR's stated reason for lukewarm enthusiasm is the low yield of actionable violations. If the 26 AFO attributions are wrong, the violation yield could be much higher (up to 37/64, 58%), which would undercut the empirical foundation for the Rashomon effect. The 82% figure is also a headline result. The paper is honest about the gap in §3.2, and Appendix A.3 offers plausible reasons to trust WDNR's diligence, but plausibility is not verification. A sensitivity analysis or documentation audit would settle it. I do not see a more serious internal inconsistency; the field study is well-conducted, the data and code are promised to be public, and the qualitative findings are reported with nuance. The conditional verdict is appropriate: the central organizational-context claim is likely right, but the precise magnitude of the regulatory gap and the strength of the WDNR-calculus argument should be presented with uncertainty until the AFO attribution is checked.","tokens_in":20715,"tokens_out":8223,"duration_ms":79946,"concrete_test":"Using the released WDNR follow-up logs in the Hugging Face dataset, audit the 26 AFO-attributed detections: require documented evidence per event (e.g., direct contact with the smaller farm, manure type inconsistent with the nearby CAFO, or field ownership excluding the permitted CAFO). Recompute Figure 4 treating any AFO attribution lacking documentation as an unclassified CAFO violation in a sensitivity analysis. If the non-compliance rate among confirmed events rises above ~50%, the 'few clear violations' premise for WDNR's assessment and the 82% figure are unsupported; if documentation is present for ≥80% of cases, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.2 reports that of 64 WDNR-confirmed spreading events, only 11 were clear CAFO violations; 27 were pre-February (18 substantiated by historical imagery) and the remaining ~26 were attributed to smaller AFOs, yielding the headline 82% figure. The paper states the AFO explanation 'cannot be evaluated from the available data as directly' (§3.2). This attribution is load-bearing for two reasons. First, the 82% regulatory-gap claim is a central contribution. Second, WDNR's lukewarm assessment—the empirical basis for the 'regulatory Rashomon effect'—is explicitly tied to seeing 'relatively few... clear regulatory violations.' If a substantial share of the 26 AFO attributions are actually CAFO applications (e.g., because specialists relied on local knowledge or farm self-reports rather than documented evidence), then the non-compliance rate would rise, the 82% figure would shrink, and WDNR's cost-benefit calculus might change. The paper's defense—interviews describing calls to smaller farms and manure-type observations—is anecdotal and not systematically recorded for each detection. Appendix A.3 argues against undercounting due to regulatory capture, but that argument does not resolve the missing verification of AFO status. This is not a disagreement with consensus; it is an internal evidentiary gap that the authors explicitly concede.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a concurrent field trial of a satellite-imagery computer vision model for detecting winter manure spreading, run with the Wisconsin Department of Natural Resources (WDNR) and the Environmental Law and Policy Center (ELPC) during February–March 2023. Both organizations verified detections at broadly similar rates and the model's confidence scores rank-ordered with ground truth, but the organizations diverged in their assessment of the tool: WDNR saw relatively few confirmed events as clear regulatory violations within its purview (11 of 64), while ELPC valued the documentation of environmental risk, including events that were compliant under current regulations. The paper introduces the label 'regulatory Rashomon effect' for this divergence and argues that the tool exposes gaps in existing law, reporting that 82% of confirmed WDNR detections fell 'between regulatory cracks' due to the February 1 temporal cutoff and the 1,000-animal-unit CAFO threshold.","tokens_in":21015,"tokens_out":5335,"duration_ms":53788,"significance":"If the headline findings hold, this is a valuable empirical contribution to the emerging literature on AI in environmental enforcement and organizational context in human-AI systems. The study is unusual in having two concurrent, independent field trials with the same model, with real verification effort (ELPC alone drove ~4,300 miles and spent 175 hours), and the authors have made the detection data and code public. The 'regulatory Rashomon effect' is a useful framing that goes beyond model-accuracy metrics. However, the central quantitative claim—the 82% 'regulatory cracks' figure, and WDNR's 'few clear violations' assessment—rests on attributions (pre-February timing and AFO vs. CAFO status) that the paper itself concedes cannot be fully verified. The qualitative organizational-context finding is plausible and interesting, but the paper's most cited number needs to be placed on firmer evidentiary ground.","major_comments":[{"comment":"The paper counts all 27 detections 'reported to have been spread prior to February 1' as compliant, but its own manual imagery review substantiated only 'at least 18 of the 27 cases' as applied by February 1. Supplementary Table S2 lists 5 as 'February Application' and 4 as 'Unsure'. Counting all 27 as compliant inflates the 82% figure. Please recompute the compliant share excluding, or explicitly justifying, the 5 February applications and 4 unsure cases; at a minimum the paper should acknowledge that up to 9 of the 27 may not be pre-February events.","section":"§3.2 and Supplementary Table S2"},{"comment":"The paper's headline 82% figure and WDNR's lukewarm assessment depend on the attribution of roughly 26 confirmed events to smaller, unregulated AFOs rather than CAFOs. The paper states this 'cannot be evaluated from the available data as directly' (§3.2). The interview-based evidence (calls to smaller farms, manure-type observations) is anecdotal and not recorded per detection. Because a substantial misattribution would change both the 82% figure and WDNR's cost-benefit calculus, the authors should either provide per-detection documentation of the AFO determination and its basis, or perform a sensitivity analysis showing how the headline percentages and the organizational-context interpretation change under alternative misattribution rates (e.g., 10%, 25%, 50%).","section":"§3.2, Figure 4, and Discussion"},{"comment":"The reported arithmetic does not appear to reconcile. If 64 events were confirmed, 11 were non-compliant, 27 were pre-February, and the remainder were AFOs, then the AFO share of post-February detections is roughly 26/(26+11) ≈ 70%, not the stated 62%. Conversely, 62% of the 37 post-February detections implies about 23 AFO events, leaving 30 pre-February events rather than 27. Please clarify the exact counts underlying Figure 4 and ensure the percentages in the text and abstract are consistent.","section":"§3.2 and Discussion"}],"minor_comments":[{"comment":"Typo: 'Enviromental Law and Policy Center' should be 'Environmental Law and Policy Center'.","section":"§2.2.2"},{"comment":"The term 'dumping' is used for what is often a legal land application of manure; consider using 'manure spreading' or 'land application' in neutral descriptions to avoid conflating detection with illegality.","section":"Throughout"},{"comment":"The label 'Field Visible (Among Visited)' is clear, but the paper might also report how many visited detections could not be seen from public roads and whether that non-visibility is correlated with model confidence; the current panel suggests it is not, which is useful.","section":"Figure 2B"},{"comment":"The comparison between the model's predictive lift and 'random selection' (77x, 219x) is striking but would benefit from a clearer definition of the random-selection baseline, since the detection area is 6 km boxes around known CAFO sites, not a random statewide sample.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"The 82% figure is the most likely result to be picked up by press and policy audiences, so the authors should be held to a high standard on its evidentiary basis. The paper already concedes the key limitation in §3.2, which is good, but the current presentation treats the AFO attribution as settled. A sensitivity analysis or per-detection documentation would substantially strengthen the paper. The arithmetic inconsistency in the 62% claim should also be fixed before the paper is suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: the genuinely new thing here is the concurrent field trial—the same satellite AI detection tool run through a regulator and an advocacy group at the same time, with both verifying detections at similar rates. That design earns the paper a serious referee. But the headline claim, that 82% of confirmed dumping fell between regulatory cracks, sits on an AFO-vs-CAFO attribution the authors concede they cannot verify directly. It needs a sensitivity analysis or a softer claim before I'd trust it as a headline.\n\nWhat's good: the 'regulatory Rashomon' framing is not just a label. The paper shows both organizations confirmed manure presence at comparable rates across confidence buckets, then diverged on usefulness: WDNR wanted clear violations and found the pilot underwhelming; ELPC wanted to document risk, compliant or not, and found it valuable. That contrast is supported by interview quotes, survey data, and the different filtering rules the two pilots used. The model also rank-orders ground truth, and the WDNR desk-review precision improvement is a nice, practical point about human-AI collaboration. The appendix is transparent about logistics and failure modes, and the data/code promise is appropriate.\n\nThe soft spot is real and the reader's concern is fair. Section 3.2 admits the AFO explanation 'cannot be evaluated from the available data as directly.' The evidence that the remaining ~26 events came from smaller farms is anecdotal: specialist calls, local knowledge, manure-type observations. It isn't recorded systematically per detection. The historical imagery check for the pre-February events is solid (18 of 27 substantiated), but that doesn't cover the AFO bucket. If a substantial share of those 26 were actually CAFO applications, the non-compliance rate would rise, the 82% figure would shrink, and WDNR's 'relatively few clear violations' assessment—the empirical basis for the Rashomon effect—would need to be reweighed. Appendix A.3 argues against undercounting via regulatory capture, but that is an argument, not data.\n\nIs this fatal? No. The organizational-context finding is more robust than the 82% number: the qualitative divergence in perceived usefulness stands even if some AFO attributions are wrong, because the two organizations' goals genuinely differed and the interviews support that directly. The weakness is an evidentiary gap, and the authors flag it themselves. That's a sign of honest work, not a reason to reject.\n\nWho should read it: people working on AI in public-sector and environmental enforcement, and anyone studying how organizational incentives shape tool adoption. I'd send it to peer review with a request to run sensitivity analyses on the AFO attribution (e.g., treat 25% or 50% as CAFOs and show the conclusions still hold) and to separate the robust 'context matters' finding from the shakier '82%' figure. I would not desk-reject it.","headline":"The dual-organization field trial is a genuine contribution, but the headline 82% 'regulatory cracks' figure rests on an unverifiable AFO attribution and needs a sensitivity analysis before the paper's claims are as strong as they look.","tokens_in":21507,"tokens_out":3697,"would_cite":true,"duration_ms":34588,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The same satellite-based AI tool was judged useful by an advocacy group and not by the regulator that verified the same detections.","keywords":["AI evaluation","organizational context","environmental enforcement","satellite remote sensing","CAFO regulation","regulatory thresholds","human-AI collaboration","field study"],"falsifier":"Count the animal units at every facility that WDNR attributed a compliant post-February application to (for example, by cross-referencing permit records or estimating barn capacity from satellite imagery); if a substantial share of those facilities exceed 1,000 animal units, the 82% regulatory-gap figure would collapse, even if the organizational-context finding still stands.","tokens_in":20570,"feed_emoji":"🛰️","tokens_out":5626,"duration_ms":54673,"temperature":0.7,"pith_summary":"The paper reports a field trial in Wisconsin where a satellite-imagery AI model that detects winter manure spreading was run concurrently by the state regulator (WDNR) and an environmental advocacy group (ELPC). Both organizations confirmed the model's detections at nearly identical rates, but they came to opposite conclusions about whether the tool was worth using: WDNR saw few clear violations and doubted the value, while ELPC valued the documentation of environmental risk even when the spreading was legal. The paper argues that organizational goals, not model accuracy, determine how AI tools are assessed in practice. It also finds that 82% of the spreading WDNR confirmed fell outside existing regulations, because of hard thresholds on facility size and application dates. The study matters because it shows AI can expose gaps in environmental law while raising the question of how to evaluate AI in institutional settings.","feed_headline":"Same AI tool, opposite verdicts from regulator and advocates","feed_subtitle":"Field study finds organizational goals, not model accuracy, decide whether AI is judged a success in environmental enforcement.","key_machinery":"The paper's conceptual engine is the 'regulatory Rashomon effect': concurrent field trials gave the same AI detections to a regulator and an advocacy group, so identical ground truth was filtered and judged by two different institutional mandates. Operationally, the tool is a fine-tuned YOLOv5 object detector that flags candidate land-application events in near-daily 3m/pixel Planet satellite imagery; detections were routed to WDNR only on permitted CAFO fields and after central-office desk review, while ELPC verifiers received the top-confidence detections within drivable range and field-checked them from public roads. The argument that regulation has arbitrary boundaries rests on two statutory thresholds—the February 1–March 31 winter ban and the 1,000-animal-unit CAFO definition—which are the specific legal machinery that lets most confirmed spreading escape enforcement.","core_discovery":"In concurrent February–March 2023 field trials, a YOLOv5-based computer vision model analyzing near-daily 3m/pixel Planet satellite imagery routed detections of land-applied manure to WDNR and to ELPC. Both organizations verified the presence of manure at similar rates (for example, about 35% for the highest-confidence detections), and on the few detections both followed up they agreed. Yet WDNR determined that only 11 of the 64 confirmed spreading events were clear violations of Wisconsin's winter-application rules; the rest were attributed to smaller AFOs below the 1,000-animal-unit CAFO threshold or to applications spread before February 1. ELPC, whose mission includes documenting environmental risk beyond current regulation, saw the tool as valuable for revealing the extent of winter spreading. The paper names this divergence the 'regulatory Rashomon effect': the same evidence, interpreted through different institutional mandates, yields opposite judgments of utility. It further reports that 82% of WDNR-confirmed dumping 'fell between these regulatory cracks,' and that a human pre-screening step substantially raised precision over the raw model.","pith_inferences":["If the organizational divergence generalizes, then AI procurement decisions in government should specify not just accuracy targets but the decision the tool is meant to inform; the same model could be a success for an advocacy group and a failure for a regulator by design, not by accident.","The 82% figure, if it survives better data on facility sizes, implies that the environmental harm from winter spreading is dominated by unpermitted smaller farms, a policy target that neither current regulation nor most remote-sensing research addresses.","A testable extension would be to run the same detection pipeline with other state environmental agencies or with a regulator whose statute includes risk-based, graduated thresholds; the Rashomon effect predicts their usefulness ratings would move in the direction of the advocacy group's.","The paper's own evidence that pre-February spreading can be confirmed from historical satellite imagery suggests an inexpensive upgrade: filter out detections whose imagery time series shows application before the ban, which would raise the share of detections that are violations and could change WDNR's cost-benefit calculus."],"forward_implications":["If AI tools are assessed through each organization's mandate, then accuracy benchmarks alone cannot predict whether a deployment will be judged a success; intended use must be part of the evaluation.","A human-in-the-loop pre-screening step, where domain experts review model outputs before field follow-up, raises precision above the raw model's, so deployed systems should budget for expert review.","Satellite-based detection can surface 450–1000% more violations than current complaint-driven processes, even when the absolute number of clear violations remains small.","Because most confirmed winter spreading is legally compliant under current thresholds, the same AI tool that helps enforcement also provides evidence that the law's bright lines (February 1; 1,000 animal units) may not match environmental risk.","For advocacy organizations, the tool can act as a force multiplier, turning confirmed detections into complaints or public-awareness campaigns that supplement limited regulator capacity."],"supporting_citations":[{"why":"Developed the land-application detection model and the training data that this field trial deploys; removing it removes the tool under study.","marker":"[14]"},{"why":"EWG's 2022 analysis of Wisconsin animal feeding operations supplies the statistic that only 3% of operations are permitted CAFOs, which supports attributing most spreading to smaller AFOs.","marker":"[1]"},{"why":"YOLOv5 is the pre-trained object-detection architecture that is fine-tuned for the manure-detection task, so it anchors the model's capabilities and limits.","marker":"[39]"},{"why":"Provides evidence of bunching just below the 1,000-animal-unit threshold in Iowa hog farms, which supports the paper's policy concern that hard size cutoffs create regulatory avoidance.","marker":"[73]"}],"fun_headline_variants":["AI tool: regulator says no, advocates say yes","Same AI, different verdicts: Wisconsin field trial","Why a regulator and advocates saw the same AI differently","Organizational goals trump model accuracy in AI deployment","Satellite AI: one model, two very different verdicts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that most confirmed winter spreading is legal depends on WDNR specialists' determinations that the post-February applications came from farms below the 1,000-animal-unit cutoff, which the paper could not verify directly from data.","fun_headline_variants_meta":{"raw":{"variants":["AI tool: regulator says no, advocates say yes","Same AI, different verdicts: Wisconsin field trial","Why a regulator and advocates saw the same AI differently","Organizational goals trump model accuracy in AI deployment","Satellite AI: one model, two very different verdicts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000203,"raw_usage":{"total_tokens":1419,"prompt_tokens":1009,"completion_tokens":410,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":625,"completion_tokens_details":{"reasoning_tokens":333}},"tokens_in":625,"tokens_out":410,"duration_ms":4175,"temperature":1.0,"reasoning_tokens":333,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:22:09.561430+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Count the animal units at every facility that WDNR attributed a compliant post-February application to (for example, by cross-referencing permit records or estimating barn capacity from satellite imagery); if a substantial share of those facilities exceed 1,000 animal units, the 82% regulatory-gap figure would collapse, even if the organizational-context finding still stands.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides evidence of bunching just below the 1,000-animal-unit threshold in Iowa hog farms, which supports the paper's policy concern that hard size cutoffs create regulatory avoidance."}],"review_version":1}