{"id":"8e5294ac-0a6f-4a01-ac0a-1de862c11d0f","arxiv_id":"2608.08601","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Workplace AI agent risks cluster at the human-agent boundary, and augmentation can erode worker skills even when automation shifts the visible failures to the organization.","lead":"The authors generated 8,356 AI-agent risk scenarios from 2,078 O*NET job tasks using a large language model, then built a 15-category workplace risk taxonomy and validated it with workers and a second LLM. The report matters because it gives organizations a job-level vocabulary for anticipating AI agent risks before they happen, including the quiet deskilling that can come from helpful augmentation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Corpus frequencies are likely artifacts of the generation prompt; Findings 1–3 describe the corpus, not workplace risk prevalence.","rationale":"The reader's conditional verdict rests on the same weakest assumption, and I agree with it. The paper is honest and methodical; it includes a limitations section that explicitly disclaims frequency interpretations. That candor is credit to the authors, but it also exposes the gap: the abstract and conclusion do not honor the disclaimer. The strongest claim in the abstract is that findings 'show' that risks depend on how people work with agents and how agents are deployed. For a taxonomy paper, the contribution is the classification and the validated plausibility of scenarios; those are real. The 45-worker study and 26-participant usability experiment provide independent support for plausibility and usability, and the 233 incident/failure cases ground the taxonomy in documented harms. None of this, however, makes the frequency distribution of an LLM-generated corpus an empirical measurement of workplace risk. The severity and deployment-mode labels are also outputs of the same generator, so the augmentation/automation asymmetry in Findings 1 and 3 and the severity concentration in Finding 2 are not independently verified. I see no internal inconsistency or fraud; the issue is an interpretive overreach that a prompt-ablation experiment would resolve. If the ablation reproduces the distribution, the findings gain real support; if not, the paper should be accepted conditionally as a taxonomy-plus-hypotheses contribution, with Findings 1–3 clearly labeled as corpus descriptions. This does not change the reader's verdict, so verdict_should_be is UNCHANGED, with agreement_with_reader = agree.","tokens_in":20628,"tokens_out":4686,"duration_ms":45583,"concrete_test":"Run a controlled regeneration study: for a random sample of 500 of the same 2,078 O*NET tasks, generate risk scenarios with gpt-4o-mini using a neutral prompt that omits the multi-layer framework, the layer/pathway instructions, and the severity/deployment labeling requirements, then map the outputs to the 15-category taxonomy with the same mapping procedure. Compare category shares and augmentation/automation splits with Figure 3. If Erroneous Agent Actions and Human Capability Erosion still dominate and the augmentation split is similar, the findings are robust; if the distribution shifts materially (e.g., a different category dominates), the reported frequencies are artifacts of the prompt design and Findings 1–3 should be reframed as hypotheses about plausible risks rather than empirical results.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's quantitative findings depend on treating the distribution of the 8,356 LLM-generated scenarios as informative about workplace AI agent risks. That dependence is load-bearing because Findings 1–3 are stated in terms of which risks are 'most common,' 'most severe,' and which deployment mode is riskier. Section 3.3 describes a prompt that 'explicitly instructed the model to consider risks at each framework layer.' The prompt also required every scenario to be labeled by component/interaction pathway, severity tier, and deployment mode. This does not measure risk distribution; it instructs the model to populate each cell of the authors' framework. The taxonomy categories were then extended to cover the same generated scenarios, so the reported shares (Erroneous Agent Actions 30.6%, Human Capability Erosion 21.3%, 97.0% of erosion under augmentation, etc.) partly reflect the instrument that created them. Section 5.3 concedes that frequencies 'describe only the composition of our risk corpus' and 'should not be interpreted as estimates of how often these risks occur in real workplaces.' Yet the Conclusion uses those same frequencies to assert that risks 'do not arise from agents alone.' The worker validation (450/8,356 scenarios) and LLM-judge ratings concern plausibility and task alignment, not distributional representativeness; severity and deployment labels were not human-validated at all. The taxonomy itself can stand as a useful classification, but the empirical framing of Findings 1–3 is not secured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper develops a multi-layer framework for workplace AI agent risks, uses it to generate 8,356 risk scenarios from 2,078 O*NET job tasks with an LLM, validates a sample with workers and an LLM judge, and builds a 15-category workplace AI agent risk taxonomy extended from an existing generative AI risk taxonomy. The paper reports four findings: augmentation carries deskilling and erosion risks; Erroneous Agent Actions is the most frequent and severe category; automation concentrates risk in organizations while augmentation concentrates it in workers; and the proposed taxonomy is more usable than two comparison taxonomies. The authors conclude that workplace AI agent risks depend not only on agents but also on human-agent interaction and deployment choices.","tokens_in":20801,"tokens_out":5840,"duration_ms":63929,"significance":"If the quantitative claims are properly scoped, this is a useful and timely contribution: it provides a structured, job-task-grounded way to anticipate workplace AI agent risks, fills a genuine gap in existing AI risk taxonomies for organizational and agent-execution risks, and ships a concrete artefact (the 15-category taxonomy) that practitioners can adopt for risk registers and impact assessments. The strengths include the grounding in O*NET task descriptions, the explicit two-part validation of scenario plausibility and task alignment, the coverage comparison against ten established frameworks, and the comparative usability study. The main weakness is that the headline frequency and severity findings rest on LLM-generated labels whose representativeness is acknowledged in Section 5.3 to be limited, yet those same frequencies are used in Findings 1-3 and the Conclusion as empirical evidence. The taxonomy and framework can stand as a useful contribution, but the quantitative findings need to be re-scoped or independently validated.","major_comments":[{"comment":"The generation prompt explicitly instructs the model to consider risks at each framework layer and to label every scenario by component or interaction pathway, and the taxonomy in Section 3.5 is then built from those labeled scenarios. As a result, the reported corpus shares (e.g., Erroneous Agent Actions at 30.6%, Human Capability Erosion at 21.3%, and the 97.0% augmentation concentration of the latter) are partly an artifact of the elicitation instrument rather than an estimate of workplace risk prevalence. Section 5.3 concedes that these frequencies 'describe only the composition of our risk corpus' and 'should not be interpreted as estimates of how often these risks occur in real workplaces,' but Findings 1-3 and the Conclusion in Section 6 use exactly these frequencies to support the paper's central claim. This is an internal inconsistency in a load-bearing part of the argument. Please either rephrase every quantitative finding as a statement about the generated corpus, or provide an independent validation of representativeness (for example, a generation-baseline comparison and a human-annotated random sample of severity and deployment labels).","section":"Section 3.3, Section 4, Section 5.3"},{"comment":"The severity and deployment-mode labels on which Findings 2 and 3 rely were not human-validated. The worker study covers 450 of 8,356 scenarios (about 5.4%) and evaluates plausibility and task alignment, not severity or deployment mode. The LLM-as-a-judge reliability check reports band agreement only for plausibility (98.9%) and connection-to-task (98.1%); no agreement is reported for severity or deployment mode. Claims such as '88.2% of Erroneous Agent Actions are rated high or critical,' 'Financial Losses are 90.9% from automation,' and 'Employment Displacement is 92.7% high or critical' therefore rest on unvalidated LLM labels. Please add a human annotation study for severity and deployment mode on a random sample, or explicitly re-label these percentages as LLM-generated hypotheses about the scenario corpus rather than validated measurements.","section":"Section 3.4, Findings 2-3"},{"comment":"The taxonomy extension is performed by using gpt-4o-mini to map generated scenarios to Li et al.'s generative AI risk taxonomy, with unmatched scenarios inspected manually, but the paper reports no inter-annotator agreement for this mapping or for the manual conversion of incidents and failure cases. Because the same model family that generated the scenarios also performed the mapping, the resulting 15-category structure and the statement that the taxonomy 'covers all our risk scenarios' could reflect the generator's systematic mapping preferences. Reporting a sample-based agreement study with a second annotator or a different model, together with the number and definitions of unmatched scenarios before extension, would make the taxonomy construction reproducible and less dependent on a single model.","section":"Section 3.5"}],"minor_comments":[{"comment":"In the typeset copy, the category labels in panel A are not visually aligned with their corresponding bars and percentages, making it difficult to read which frequency belongs to which category. Please re-render the figure with explicit row labels.","section":"Figure 3"},{"comment":"The structural integrity check uses a Jaccard overlap threshold of 0.30 and reports a mean semantic similarity of 0.24, but the threshold choice and the absence of a comparison baseline are not justified. Please report the full distribution of pairwise similarities and the rationale for the threshold in the supplement.","section":"Section 3.6"},{"comment":"The usability comparison is based on 26 participants and 130 classification tasks, with a pairwise p-value of .027 and a Friedman test p-value of .042; no multiple-comparison correction is applied. The wording 'outperforms' is stronger than this evidence supports; please temper the conclusion and report confidence intervals for the effect sizes.","section":"Section 4, Finding 4"},{"comment":"The phrase 'They can therefore overlooks socio-technical AI risks' contains a typographical error ('overlooks' should be 'overlook'), and the paragraph would benefit from a space or hyphen between 'overlook' and 'socio-technical'.","section":"Section 2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's primary contribution is the taxonomy and the job-task-grounded generation method, not the frequency estimates. The authors should be asked to reconcile the strong quantitative language in Findings 1-3 and the Conclusion with the explicitly acknowledged limitation in Section 5.3. I do not see this as a reject: the taxonomy fills a real gap, and the study design is transparent enough that the claims can be re-scoped without changing the core contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the 15-category taxonomy and the 8,356-scenario corpus are worth having, but the headline findings are framed as if they describe real workplace risk when they mostly describe what one LLM was prompted to write. The paper's own Section 5.3 admits the frequencies 'describe only the composition of our risk corpus,' yet the abstract and Conclusion use those same frequencies to claim that risks 'do not arise from agents alone.' That inference goes beyond the evidence.\n\nWhat's actually new and good: the integration of O*NET job tasks with LLM-based scenario generation, extended into a workplace-agent-specific taxonomy grounded in both generated scenarios and documented incidents. The coverage comparison against ten existing frameworks is useful and shows a real gap: organizational and agent-execution risks are poorly covered by prior taxonomies. The usability study with 26 workers is a genuine empirical contribution, even if the effect is modest (d=0.20). The validation effort is real: 45 workers rated 450 scenarios, an independent judge from a different model family rated the rest, and the paper reports cross-judge agreement. The limitations section is unusually candid.\n\nThe soft spots are real but addressable. The generation prompt explicitly instructed the model to consider each framework layer and interaction pathway, so the category shares and augmentation/automation splits are partly an artifact of the instrument. Severity and deployment labels came from the same generator model and were not human-validated; only 5.4% of scenarios received worker ratings. The paper sometimes phrases findings carefully as corpus-composition statements, but the Conclusion and abstract drop that qualifier. This is not a fatal flaw—it is a framing problem plus a missing validation step. Releasing the prompt and data, adding human validation for a sample of labels, and reframing Findings 1–3 as corpus composition would put it on solid ground.\n\nWho this is for: researchers and practitioners building AI risk taxonomies, running impact assessments, or doing red-teaming exercises for workplace agents. The taxonomy is a plausible scaffold and the scenario corpus could seed practical work. I'd send it to peer review; a serious referee should get past the framing issue and push for the fixes above. I'd cite it for the taxonomy and corpus, not for the distributional claims.","headline":"The taxonomy and scenario corpus are genuinely useful, but the distributional findings describe what the authors prompted an LLM to generate, not workplace risk prevalence; the paper's own limitations section says as much, yet the Conclusion overreaches.","tokens_in":21461,"tokens_out":2186,"would_cite":true,"duration_ms":26469,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that workplace AI agent risk is shaped by the human-agent boundary and deployment mode, not just by the agent itself.","keywords":["AI agents","workplace risk","risk taxonomy","augmentation vs automation","skill erosion","human oversight","job tasks","sociotechnical risk"],"falsifier":"Collect documented incidents from actual AI agent deployments in the same job roles and compare the category distribution and augmentation-versus-automation split against the paper's 8,356 scenarios; if real incidents concentrate in different categories, such as more technical model failures and fewer human-capability-erosion scenarios, the four findings and the taxonomy's claimed priorities would be called into question.","tokens_in":1803,"feed_emoji":"🤖","tokens_out":1788,"duration_ms":84508,"temperature":0.7,"pith_summary":"The paper maps the risks of AI agents at the level of individual job tasks. It embeds a three-layer framework of agents, goals, environments, and their interactions into a structured prompt, applies it to 2,078 computer-based job tasks, and generates 8,356 risk scenarios labelled by severity and by whether the agent automates or augments the task. The corpus yields a 15-category taxonomy and four findings: augmentation is not inherently safe because it can erode skills and oversight; erroneous agent actions form the largest and most severe category, mostly arising where humans interpret and act on agent output; automation shifts risks to organizations; and workers rated the taxonomy easier to use than two alternatives. The paper concludes that workplace AI agent risks depend as much on how people work with agents and how agents are deployed as on the agents themselves.","feed_headline":"Keeping a human in the loop does not erase AI-agent risk","feed_subtitle":"Largest and most severe risks: agents acting wrongly and workers' fading skills.","key_machinery":"The load-bearing mechanism is a multi-layer framework of an agentic AI system, defined as the agents, their goals, and their environment, together with the interaction pathways among them and with human workers, organized under three layers: technical capability, human interaction, and systemic impact. The framework is embedded in a structured prompt that labels each generated risk scenario with its component or interaction pathway and with a deployment mode, augmentation or automation. This same framework, extended through the generated scenarios plus documented incidents, produces the 15-category workplace AI agent risk taxonomy, and the taxonomy is then tested for structural distinctness, coverage against existing frameworks, and usability with workers.","core_discovery":"The central claim, on the paper's own terms, is that workplace AI agent risks do not arise from agents alone; they also depend on how people work with agents and how agents are deployed. The supporting evidence is a corpus of 8,356 risk scenarios grounded in 2,078 real job tasks: Erroneous Agent Actions is the largest and most severe class, at 30.6% of scenarios with 88.2% rated high or critical, and it is dominated by misinterpretation and incorrect recommendations rather than outright fabrication. The second-largest class, Human Capability Erosion at 21.3%, is overwhelmingly an augmentation phenomenon, with 97.0% of those scenarios arising when an agent assists rather than replaces a worker, and with 1,420 scenarios specifically describing erosion of workers' decision and oversight capability. By contrast, automation shifts the dominant categories to Operational Failures, Financial Losses, and Employment Displacement. The paper interprets these patterns as evidence that safer workplaces require not only safer agents but also carefully designed human-agent collaboration.","pith_inferences":["A direct testable extension would measure whether structured worker-contestability mechanisms, such as override channels and appeal routes, reduce the severity of Erroneous Agent Actions in field deployments, which would directly test the paper's human-agent-boundary thesis.","The paper's own caveat that its frequencies describe only the corpus implies a robustness prediction: repeating the scenario generation with different models and prompt designs should change the category shares substantially, so the taxonomy structure may persist even while the ranking of risk categories shifts.","The sharp augmentation-versus-automation split is likely a simplification, because real deployments often mix both modes; a finer-grained hybrid category could reveal additional risk patterns at the transition between human oversight and full automation.","If the paper is right about gradual skill erosion, long-horizon workplace studies are the natural next step: regular skill assessments of workers who use agents daily should show measurable decline in unaided task performance over months, something incident databases cannot capture."],"forward_implications":["Keeping a human in the loop does not automatically make a deployment safe: 21.3% of risk scenarios are human capability erosion and 97.0% of those occur under augmentation, including 1,420 scenarios of eroded decision and oversight capability.","Risk assessment should target the human-agent boundary: the largest category, Erroneous Agent Actions, is mostly misinterpretation and wrong recommendations rather than fabrication, and 69.5% of those risks arise under augmentation.","Automation and augmentation need different governance: automation concentrates risk in organizational categories such as operational failures, financial losses, and employment displacement, while augmentation concentrates risk in the worker, including skill erosion and psychological and social risks.","A workplace-specific taxonomy fills gaps left by broader AI risk classifications, since organizational risks such as operational and strategic management failures and agent-execution risks are absent or only partially covered in the ten frameworks compared.","Workers using the proposed taxonomy classified 97.7% of scenarios correctly and preferred it over a generative AI risk taxonomy in 64% of non-tied comparisons, supporting its practical usability."],"supporting_citations":[{"why":"Provides the sociotechnical three-layer evaluation framework that the paper extends to workplace agents.","marker":"(Weidinger et al. 2023)"},{"why":"Supplies the generative AI risk taxonomy whose categories are extended into the 15-category workplace taxonomy.","marker":"(Li et al. 2025b)"},{"why":"Filters the occupational database down to the 2,078 computer-based job tasks used as the grounding corpus.","marker":"(Shao et al. 2025b)"},{"why":"Frames labor-market exposure to AI and motivates the job-task-level perspective on automation and augmentation.","marker":"(Eloundou et al. 2024)"},{"why":"Supplies the ironies-of-automation mechanism underlying the human capability erosion finding.","marker":"(Bainbridge 1983)"},{"why":"Grounds the use of language-model-based generation and red-teaming to anticipate risks before documented incidents.","marker":"(Perez et al. 2022)"},{"why":"A multi-agent risk taxonomy compared against the proposed one and a source for interaction-pathway risk thinking.","marker":"(Hammond et al. 2025)"},{"why":"A comparison taxonomy whose single catch-all category motivates the usability evaluation and the preference result.","marker":"(Slattery et al. 2025)"}],"fun_headline_variants":["Augmentation can quietly erode worker skills and oversight","Delegation, not automation, is the risky part of AI agents","Erroneous AI actions and fading skills top workplace risk list","Overreliance on AI agents slowly erodes human oversight","Human-AI collaboration, not agents alone, creates workplace risks"],"cache_read_input_tokens":23424,"weakest_assumption_plain":"The entire risk map rests on whether the language model's imagined risk scenarios fairly represent real workplace risks, and the paper itself cautions that the category frequencies describe only the corpus, not how often risks occur in real workplaces.","fun_headline_variants_meta":{"raw":{"variants":["Augmentation can quietly erode worker skills and oversight","Delegation, not automation, is the risky part of AI agents","Erroneous AI actions and fading skills top workplace risk list","Overreliance on AI agents slowly erodes human oversight","Human-AI collaboration, not agents alone, creates workplace risks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001066,"raw_usage":{"total_tokens":4538,"prompt_tokens":1083,"completion_tokens":3455,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":699,"completion_tokens_details":{"reasoning_tokens":3370}},"tokens_in":699,"tokens_out":3455,"duration_ms":25003,"temperature":1.0,"reasoning_tokens":3370,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:29:30.804752+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect documented incidents from actual AI agent deployments in the same job roles and compare the category distribution and augmentation-versus-automation split against the paper's 8,356 scenarios; if real incidents concentrate in different categories, such as more technical model failures and fewer human-capability-erosion scenarios, the four findings and the taxonomy's claimed priorities would be called into question.","supporting_citations":[],"review_version":1}