{"id":"533a3a65-8ff2-4a6d-9d67-b2f136a71aef","arxiv_id":"2606.12432","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Decommissioned AI systems produce residual socio-technical risks termed AI debris that continue to influence institutions, addressed via a proposed AI Debris Decommissioning Protocol.","lead":"This paper defines AI debris as the persistent socio-technical residue left by decommissioned AI systems and proposes a decommissioning protocol to manage it. A smart generalist might read it to see how AI oversight must extend past shutdown to cover lingering effects on workflows, accountability, and trust.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Causal attribution of listed socio-technical effects to prior AI systems (vs. pre-existing practices) is asserted but not demonstrated beyond a single vignette.","rationale":"The reader's weakest_assumption directly names the load-bearing premise. Because the manuscript (as described) advances a typology and protocol on the basis of conceptual argument plus one illustration, the absence of evidence isolating AI causation keeps the central claim provisional; this does not alter the UNVERDICTED status.","tokens_in":1714,"tokens_out":321,"duration_ms":17825,"concrete_test":"Select two matched organizations that each replaced a high-stakes screening process (one AI-based, one rule-based); apply the AIDP evidence requirements to both and measure the five debris domains at 6 and 18 months post-change while holding constant leadership turnover and data-system updates; if effect sizes and attribution patterns do not differ significantly, the AI-specific residual-risk claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that workflow dependency, data contamination, capability displacement, legitimacy erosion, and accountability breakdown are both widespread after AI withdrawal and specifically traceable to the decommissioned system rather than to baseline organizational routines, leadership changes, or non-AI policy shifts. The abstract and vignette (Amazon hiring tool) illustrate persistence via institutional memory and path dependency but supply no comparative cases, counterfactuals, or controls that would isolate the AI contribution. The proposed AIDP checklist operationalizes documentation but does not contain validation criteria or falsification tests for the causal premise itself.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that decommissioned AI systems generate persistent residual socio-technical risks termed 'AI debris,' encompassing workflow dependency, data contamination, capability displacement (deskilling), legitimacy erosion, and accountability breakdown. These effects continue to shape institutional behavior, accountability, and trust after model removal via mechanisms such as institutional memory, path dependency, blame avoidance, and data feedback loops. The manuscript develops a typology of debris domains, identifies persistence mechanisms, and proposes the AI Debris Decommissioning Protocol (AIDP) as an evaluator-ready stepwise checklist for auditable evidence on decision footprints, incident review, remediation, contestability, and post-withdrawal accountability. A vignette of Amazon's discontinued hiring tool is used to illustrate persistence of algorithmic categories and heuristics.","tokens_in":1822,"tokens_out":493,"duration_ms":24387,"significance":"If the causal mechanisms and prevalence of AI debris are substantiated, the framing could extend AI governance beyond development and deployment to the full lifecycle, offering regulators, auditors, and organizations a practical instrument (AIDP) to mitigate long-term institutional risks and avoid paper compliance in high-stakes domains.","major_comments":[{"comment":"The central claim that workflow dependency, data contamination, capability displacement, legitimacy erosion, and accountability breakdown are both widespread after AI withdrawal and specifically causally traceable to the decommissioned system (rather than pre-existing organizational practices, leadership changes, or non-AI policy shifts) rests on definitional assertion and a single vignette without comparative cases, counterfactuals, or controls. This is load-bearing for the residual-risk thesis. (Abstract; vignette description)","section":"Abstract and vignette"},{"comment":"The AIDP checklist operationalizes documentation requirements for freezing decision footprints, incident review, remediation, contestability, and accountability assignment but contains no validation criteria, falsification tests, or empirical grounding to confirm that observed effects are attributable to prior AI systems rather than baseline routines. This limits its utility as an evaluator-ready protocol. (AIDP proposal section)","section":"AIDP proposal"}],"minor_comments":[{"comment":"The boundaries of the 'AI debris' concept could be clarified to distinguish it more sharply from general organizational path dependency or change management effects, reducing potential overlap.","section":"Definition of AI debris"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments, which help clarify the scope and evidentiary basis of our conceptual contribution. We respond to each major point below and indicate planned revisions.","responses":[{"response":"The manuscript is a conceptual paper that introduces the 'AI debris' framing and typology rather than an empirical study claiming prevalence or strict causal attribution across cases. The Amazon vignette is presented explicitly as an illustration of persistence mechanisms (institutional memory, path dependency) drawn from publicly documented events, not as proof of generalizability. We agree the abstract and vignette section could more sharply delimit the contribution as definitional and illustrative. In revision we will (1) rephrase the abstract to state that the paper proposes a conceptual lens and protocol rather than demonstrating widespread effects, and (2) add explicit language in the vignette section noting the absence of comparative controls and the need for future empirical work to test attribution. This addresses the load-bearing concern without altering the core argument.","revision_made":"yes","referee_comment":"[Abstract and vignette] The central claim that workflow dependency, data contamination, capability displacement, legitimacy erosion, and accountability breakdown are both widespread after AI withdrawal and specifically causally traceable to the decommissioned system (rather than pre-existing organizational practices, leadership changes, or non-AI policy shifts) rests on definitional assertion and a single vignette without comparative cases, counterfactuals, or controls. This is load-bearing for the residual-risk thesis. (Abstract; vignette description)"},{"response":"We concur that the AIDP, as currently presented, is a proposed stepwise checklist without accompanying validation criteria or falsification tests. Its purpose at this stage is to translate the debris concept into auditable documentation requirements for practitioners and regulators. In the revised manuscript we will add a dedicated limitations subsection under the AIDP proposal that (a) acknowledges the absence of empirical validation, (b) outlines how future studies could develop falsification tests (e.g., pre/post withdrawal comparisons or matched organizational controls), and (c) positions the protocol as a starting point for evaluator use rather than a fully validated instrument. This maintains the practical intent while clarifying its provisional status.","revision_made":"partial","referee_comment":"[AIDP proposal] The AIDP checklist operationalizes documentation requirements for freezing decision footprints, incident review, remediation, contestability, and accountability assignment but contains no validation criteria, falsification tests, or empirical grounding to confirm that observed effects are attributable to prior AI systems rather than baseline routines. This limits its utility as an evaluator-ready protocol. (AIDP proposal section)"}],"tokens_in":1420,"tokens_out":550,"duration_ms":21436,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main move is to label the leftover socio-technical effects of retired AI systems as 'AI debris' and to propose the AIDP checklist as a practical response. It rightly points out that most governance work ends at deployment and treats withdrawal as a simple shutdown.\n\nThe framing is useful. The typology covers workflow dependency, data contamination, deskilling, legitimacy loss, and accountability gaps, and the persistence mechanisms (institutional memory, path dependency, blame avoidance) connect to existing organizational research. The Amazon hiring tool example shows how decision rules can outlast the system itself.\n\nThe soft spot is the evidence for the central claim. The paper states that these effects are widespread and directly traceable to the decommissioned AI, yet it supplies only the single vignette and no comparative cases, baselines, or controls that would separate AI-specific residue from ordinary organizational change. The AIDP checklist is presented as ready for auditors, but it includes no validation steps or criteria for checking whether the listed items actually reduce the risks.\n\nThis is for governance researchers, regulators, and auditors who want to extend lifecycle oversight. A reader already working on post-deployment issues would get a clear label and a starting protocol, though they would still need to add the empirical grounding.\n\nI would send it to peer review. The gap it identifies is real, and referees could push for stronger causal evidence or a testable version of the checklist.","headline":"The paper names 'AI debris' and offers a decommissioning checklist, but the causal claims rest on one vignette without controls or data.","tokens_in":2292,"tokens_out":356,"would_cite":false,"duration_ms":27193,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Decommissioned AI systems leave behind persistent socio-technical residue that continues to shape institutional behavior and accountability.","keywords":["AI governance","decommissioning","residual risk","socio-technical residue","accountability","institutional memory","AI lifecycle"],"falsifier":"An organization that decommissioned an AI hiring or decision tool and, after independent audit, showed zero measurable continuation of the original screening heuristics, data patterns, or accountability gaps traceable to that system.","tokens_in":2597,"feed_emoji":"","tokens_out":474,"duration_ms":23252,"temperature":0.7,"pith_summary":"The paper argues that withdrawing an AI system does not end its influence. Instead, effects such as workflow dependency, contaminated data, deskilling, eroded legitimacy, and broken accountability chains remain active through institutional memory and path dependency. These residues, labeled AI debris, are presented as a distinct category of post-deployment risk. The author supplies a typology of debris domains and a concrete AI Debris Decommissioning Protocol that lists auditable steps for freezing records, reviewing incidents, and assigning ongoing responsibility. A reader would care because current governance rules stop at removal and therefore miss risks that keep operating inside organizations.","feed_headline":"Decommissioned AI leaves lasting debris in organizations","feed_subtitle":"Workflow habits, contaminated data, and unclear responsibility persist after models are removed, requiring a new decommissioning protocol.","key_machinery":"AI debris, the post-withdrawal socio-technical residue of AI systems, which carries the argument by showing how listed effects continue after model removal.","core_discovery":"Decommissioned AI systems generate residual risk, termed AI debris, defined as the post-withdrawal socio-technical residue consisting of workflow dependency, data contamination, capability displacement, legitimacy erosion, and accountability breakdown. This residue persists via institutional memory, path dependency, blame avoidance, and data feedback loops. The paper supplies a typology of debris domains and an evaluator-ready AI Debris Decommissioning Protocol that specifies auditable evidence for freezing decision footprints, incident review, remediation, contestability, and post-withdrawal accountability assignment.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["AI Debris Outlives Decommissioned Systems","Residual AI Risks Persist After Withdrawal","Decommissioned AI Leaves Organizational Residue","Workflow Dependencies Survive AI Rollback","Protocol Targets Post-Removal AI Accountability Gaps"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The listed socio-technical effects are both widespread and directly caused by the prior AI system rather than by pre-existing organizational practices.","fun_headline_variants_meta":{"raw":{"variants":["AI Debris Outlives Decommissioned Systems","Residual AI Risks Persist After Withdrawal","Decommissioned AI Leaves Organizational Residue","Workflow Dependencies Survive AI Rollback","Protocol Targets Post-Removal AI Accountability Gaps"]},"model":"grok-4.3","cost_usd":0.003461,"raw_usage":{"total_tokens":1833,"prompt_tokens":682,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":34612000,"prompt_tokens_details":{"text_tokens":682,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1087,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":682,"tokens_out":64,"duration_ms":10754,"temperature":1.0,"reasoning_tokens":1087,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T19:18:10.352062+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An organization that decommissioned an AI hiring or decision tool and, after independent audit, showed zero measurable continuation of the original screening heuristics, data patterns, or accountability gaps traceable to that system.","supporting_citations":[],"review_version":1}