Pith. sign in

REVIEW 2 major objections 2 minor 4 cited by

Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem

T0 review · 2 major / 2 minor · reviewed 2026-06-29 · grok-4.3

Pith's one-line read Analysis of 3,984 AI agent skills identifies 76 confirmed malicious payloads, with 13.4% of skills showing critical security issues.

desk verdict The report scans nearly 4000 agent skills and claims 76 malicious ones plus 13.4% critical issues, but the confirmation methods stay too thin to judge the numbers. read the letter →

arxiv 2605.28588 v1 pith:CQNL5SKN submitted 2026-05-27 cs.CR cs.AI

classification cs.CRcs.AI
keywords AIagentskillsmaliciouspayloadssecuritythreatscredentialtheftdataexfiltrationthreattaxonomyskillmarketplacesbackdoorinstallation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper analyzes thousands of AI agent skills available on marketplaces and identifies a substantial number containing malicious code. It documents 76 confirmed cases of payloads designed for credential theft, backdoor installation, and data exfiltration. The report also finds that 13.4% of skills have at least one critical security issue, with some malicious ones still publicly available. A reader should care because AI agents increasingly handle sensitive tasks and credentials, making these threats a direct risk to users and systems. The authors provide a taxonomy of observed attack patterns and call for automated security analysis as the ecosystem expands.

What carries the argument

Empirical scanning and confirmation of malicious payloads in AI agent skills from marketplaces using manual and automated methods.

What would settle it

An independent re-analysis of the same 3,984 skills or the marketplaces that finds substantially fewer than 76 malicious payloads or a much lower rate of critical issues would challenge the reported scale of the threat.

Watch

Extended reading notes

Core claim

Through examination of 3,984 AI agent skills from major marketplaces, the authors confirm 76 malicious payloads and determine that 13.4% of all skills contain at least one critical-level security issue. They present a threat taxonomy based on real-world samples and detail the attack patterns, noting that at least 8 malicious skills remain available on one platform.

Load-bearing premise

The manual and automated methods used to confirm malicious payloads are accurate with low false-positive rates, and the sampled marketplaces represent the broader agent skill ecosystem.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The manuscript reports an empirical study analyzing 3,984 AI agent skills collected from major marketplaces. It identifies 76 confirmed malicious payloads involving credential theft, backdoor installation, and data exfiltration, states that 13.4% of skills contain at least one critical-level security issue, and notes that at least 8 malicious skills remain publicly available on clawhub.ai. The report documents the methodology, presents a threat taxonomy derived from real-world samples, and describes observed attack patterns, concluding that automated security analysis is required as AI agents increasingly access sensitive credentials and systems.

Significance. If the classification pipeline is shown to be reliable, the work would provide concrete, large-scale evidence of security risks in the emerging AI agent skill ecosystem and supply a useful threat taxonomy grounded in observed samples. The scale of the dataset (3,984 skills) and the explicit listing of still-available malicious items are strengths that could inform standards and tooling for agent marketplaces.

major comments (2)
  1. [Abstract / Methodology] Abstract and Methodology section: The headline figures (76 confirmed malicious payloads; 13.4% critical issues) are presented without explicit criteria for what constitutes a 'confirmed' payload, without static/dynamic analysis rules, without inter-rater agreement metrics for manual review, and without reported false-positive rates for the automated components. These omissions directly undermine assessment of the central empirical claims.
  2. [Results / Threat Taxonomy] Results section (threat taxonomy and attack patterns): The classification of payloads into categories such as credential theft or backdoor installation rests on the same unvalidated pipeline; without a concrete validation protocol or error analysis, the taxonomy cannot be treated as reliably grounded in the sampled data.
minor comments (2)
  1. [Data Collection] The manuscript should include a table or appendix listing the exact marketplaces sampled and the collection dates to allow reproducibility of the 3,984-skill corpus.
  2. [Figures] Figure captions and axis labels in any security-issue distribution plots should be expanded for clarity; current presentation leaves some category definitions ambiguous.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for highlighting the need for greater methodological transparency. We agree that the current version omits key validation details and will revise the manuscript to address both major comments. The revisions will strengthen the credibility of the empirical claims without altering the core findings.

read point-by-point responses
  1. Referee: [Abstract / Methodology] Abstract and Methodology section: The headline figures (76 confirmed malicious payloads; 13.4% critical issues) are presented without explicit criteria for what constitutes a 'confirmed' payload, without static/dynamic analysis rules, without inter-rater agreement metrics for manual review, and without reported false-positive rates for the automated components. These omissions directly undermine assessment of the central empirical claims.

    Authors: We agree that these details are missing and that their absence weakens the central claims. In the revised manuscript we will add a new subsection in Methodology that (1) defines the exact criteria used to label a payload as 'confirmed' (automated detection thresholds plus mandatory dual-author manual review), (2) enumerates the static and dynamic analysis rules applied, (3) reports inter-rater agreement (Cohen's kappa) from the manual review, and (4) provides false-positive rate estimates obtained from a sampled validation set. These additions will directly support the headline numbers. revision: yes

  2. Referee: [Results / Threat Taxonomy] Results section (threat taxonomy and attack patterns): The classification of payloads into categories such as credential theft or backdoor installation rests on the same unvalidated pipeline; without a concrete validation protocol or error analysis, the taxonomy cannot be treated as reliably grounded in the sampled data.

    Authors: The taxonomy was derived exclusively from the 76 manually confirmed samples. We acknowledge the lack of an explicit validation protocol or error analysis. The revision will include a dedicated paragraph describing the classification protocol (mapping rules from observed behaviors to taxonomy categories) together with a brief error analysis based on the same validation set used for false-positive estimation. This will make the grounding of the taxonomy explicit. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: direct empirical report with no derivations

full rationale

The paper is a technical report presenting counts of malicious skills (76 confirmed payloads, 13.4% critical issues) from direct analysis of 3,984 marketplace items. No equations, fitted parameters, predictions, or derivation chains exist. Methodology for confirmation is described as manual/automated review but is not reduced to self-definition or prior self-citations. Claims rest on external data sampling rather than internal construction, making the work self-contained against the listed circularity patterns.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Empirical security analysis with no mathematical derivations, fitted parameters, or postulated entities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem." pith.science (2026). https://pith.science/paper/CQNL5SKN

@misc{pith2026260528588,
  author       = {Pith},
  title        = {Pith review of: Technical Report: Exploring the Emerging Threats of the Agent Skill Ecosystem},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CQNL5SKN}},
  note         = {Machine review of arXiv:2605.28588}
}
read the original abstract

We analyzed 3,984 AI agent skills from major marketplaces and found 76 confirmed malicious payloads, including credential theft, backdoor installation, and data exfiltration. 13.4% of all skills contain at least one critical-level security issue and at least 8 manually confirmed malicious skills remain publicly available on clawhub.ai as of the date of publication. This report documents our methodology, presents a threat taxonomy based on real-world samples, and details the attack patterns we observed. As skill marketplaces grow rapidly and AI agents gain access to sensitive credentials and systems, automated security analysis is no longer optional.

Figures

Figures reproduced from arXiv: 2605.28588 by the authors.

Figure 1
Figure 1. Number of agent skills published every day throughout 2026. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. moltbook.com’s heartbeat prompt (a skill that runs every few hours in an unsupervised [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Local LLM Agents as Vulnerable Runtimes:A Source-Code Audit of the Agent Runtime Layer

    cs.CR 2026-06 unverdicted novelty 7.0 of 10

    CLAWAUDIT applies a STRIDE-derived taxonomy and 47 Semgrep plus 30 CodeQL rules to local LLM agent code, lifting recall on held-out OpenClaw advisories from 21.7% and 13.8% baselines to 66.8% and 75.1%.

  2. SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents

    cs.CR 2026-05 unverdicted novelty 7.0 of 10

    SkCC compiles LLM skills via SkIR to achieve portability across agent frameworks, reduce adaptation effort from O(m×n) to O(m+n), and enforce security with reported gains in task success rates and token efficiency.

  3. SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents

    cs.CR 2026-05 unverdicted novelty 7.0 of 10

    SkCC introduces a typed intermediate representation and compiler pipeline to make LLM agent skills portable across frameworks and enforce security constraints before deployment.

  4. SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agents

    cs.CR 2026-05 unverdicted novelty 6.0 of 10

    SkCC compiles LLM agent skills through a strongly-typed IR and static security checks, cutting adaptation complexity from O(m×n) to O(m+n) and raising pass rates by 12-13 points on tested platforms.

Reference graph

Works this paper leans on

8 extracted references · 1 canonical work pages · cited by 2 Pith papers

  1. [1]

    toxic flows,

    Luca Beurer-Kellner, Marco Milanta, and Marc Fischer. Invariant labs exposes novel prompt injection attack vulnerabilities, “toxic flows,” in agentic systems & mcp servers. https:// invariantlabs.ai/blog/toxic-flow-analysis, July 2025. Accessed: 2026-02-05

  2. [2]

    Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection

    Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. InProceedings of the 16th ACM workshop on artificial intelligence and security, pages 79–90, 2023

  3. [3]

    Github mcp exploited: Accessing private repositories via mcp

    Marco Milanta and Luca Beurer-Kellner. Github mcp exploited: Accessing private repositories via mcp. https://invariantlabs.ai/blog/mcp-github-vulnerability, May 2025. Accessed: 2026-02-05. 9

  4. [4]

    Clawdbot skills ganked your crypto

    OpenSourceMalware. Clawdbot skills ganked your crypto. https://opensourcemalware.com/ blog/clawdbot-skills-ganked-your-crypto, 2026. Accessed 2026-02-05

  5. [5]

    Superhuman AI exfiltrates emails

    PromptArmor Threat Intelligence Team. Superhuman AI exfiltrates emails. https://www. promptarmor.com/resources/superhuman-ai-exfiltrates-emails , January 2026. Coordi- nated disclosure date: 2026-01-12. Accessed: 2026-02-05

  6. [6]

    Schmotz, S

    David Schmotz, Sahar Abdelnabi, and Maksym Andriushchenko. Agent skills enable a new class of realistic and trivially simple prompt injections.arXiv preprint arXiv:2510.26328, 2025

  7. [7]

    Inside the ‘clawdhub’ malicious campaign: AI agent skills drop reverse shells on OpenClaw marketplace

    Liran Tal. Inside the ‘clawdhub’ malicious campaign: AI agent skills drop reverse shells on OpenClaw marketplace. https://snyk.io/articles/ clawdhub-malicious-campaign-ai-agent-skills/ , 2026. Snyk Security Labs. Accessed 2026-02-05

  8. [8]

    The lethal trifecta

    Simon Willison. The lethal trifecta. https://simonwillison.net/2025/Jun/16/ the-lethal-trifecta/, June 2025. Accessed: 2026-02-05. 10

Pith tools

Reviewed June 29, 2026 · model on record in the stance chip above.