{"id":"cc902724-31c2-4803-be3f-11efbad44f77","arxiv_id":"2509.07131","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematization of knowledge that proposes a taxonomy and reference architecture for blockchain AI agents, and catalogs security and privacy threats.","lead":"This paper organizes and surveys AI agents that interact with blockchains, from chatbots that read data to autonomous trading agents, and catalogs their security and privacy risks. It is useful for researchers and builders who want a shared vocabulary and a map of the threat landscape for Web3 AI agents.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Corpus completeness is the load-bearing assumption: the 'first' and gap claims (T-2) rest on an undocumented literature search, and the paper's own Section VI.A.3 hints at deployment-support agents, so the negative claim needs verification.","rationale":"The paper is a competent, clearly written survey, and its taxonomy and reference architecture are plausible organizational contributions. It also provides useful pointers to existing systems and datasets. However, neither the 'first' claim nor the negative gap claim can be accepted without knowing how the corpus was assembled. SoKs in security routinely include a search methodology section; its absence is a standard correctness issue because the primary contribution is coverage and categorization, not a new algorithm. The internal tension in Section VI.A.3 makes the deployment gap especially fragile: a reader can reasonably conclude from that paragraph that there are agents assisting deployment, only to read in T-2 that none exist. The remedy is straightforward: document the search protocol, give inclusion/exclusion criteria, and either revise T-2 or remove the deployment-support wording. This reinforces, rather than changes, the reader's conditional verdict.","tokens_in":18366,"tokens_out":5880,"duration_ms":50516,"concrete_test":"Conduct a formal literature search to test both claims: use arXiv, Scopus, IEEE Xplore, and Google Scholar with a pre-registered query combining ('AI agent' OR 'LLM agent' OR 'autonomous agent') AND ('blockchain' OR 'smart contract') AND ('survey' OR 'systematization' OR 'SoK'), plus a second query for ('smart contract deployment') AND ('agent' OR 'LLM' OR 'autonomous') covering 2018–2025. If a prior domain-specific SoK exists, the 'first' claim fails; if a published deployment-focused AI agent exists, T-2 fails. Report the number of screened and included papers, and map each included system to the proposed taxonomy to check whether any system is unclassifiable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — 'the first comprehensive systematization of AI agents within blockchain environments' — and the accompanying gap analysis stand or fall on whether the surveyed corpus is complete and representative. The paper never documents how works were selected: Section III discusses only five related surveys, and Tables I and II list 'representative' systems with no inclusion/exclusion criteria, no search strings, no databases, no time window, and no screening protocol. Consequently, the universal negative in T-2 ('no AI agents have yet addressed' smart-contract deployment) is inferred from the absence of such systems from an unstated sample, not from an exhaustive review. This is not a stylistic quibble: the contribution is precisely the completeness and organization of the field, so a biased or incomplete corpus would invalidate both the 'first' claim and the taxonomy's coverage. There is also an internal tension: Section VI.A.3 states that AI agents 'can serve as valuable tools' for deployment by suggesting gas optimizations and recommending adjustments to ensure successful deployment, yet T-2 asserts no agents have addressed deployment. The authors should either identify concrete deployment-focused agents or clarify that the cited works are not agents; as written, the gap claim is under-supported. Because the authors provide no method for the reader to audit coverage, the 'comprehensive' label is currently unverifiable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a systematization of knowledge on AI-driven agents for blockchain, with a focus on security and privacy. It proposes a taxonomy of conversational, instruction-following, and goal-directed agents, introduces a four-layer reference architecture, surveys representative systems for blockchain interaction and smart contract development, discusses threats and risks, and identifies open challenges. The abstract and conclusion claim that this is the first comprehensive survey of AI agents within blockchain environments.","tokens_in":18777,"tokens_out":3090,"duration_ms":27684,"significance":"If the surveyed corpus is indeed complete and representative, the paper makes a useful contribution to a young and fragmented field: the taxonomy and four-layer reference architecture provide a shared vocabulary for system design, the application and threat surveys are clearly organized, and the dataset review in Section VI.C is practically informative. The paper also identifies concrete research gaps, such as the lack of common benchmarks and the absence of deployment-focused agents. The main risk is that the 'first' and 'comprehensive' claims depend on an unverified literature corpus, and the gap analysis includes a negative claim that is under-supported by the paper's own text.","major_comments":[{"comment":"The paper never documents a systematic literature search protocol: no search engines, databases, keywords, time window, inclusion/exclusion criteria, or screening process are provided. Because the paper's central contribution is the completeness of the corpus, asserted as 'first' in the abstract and 'comprehensive' in the conclusion, the absence of a reproducible selection method makes that claim unverifiable and weakens the negative gap claim in T-2.","section":"Section III, Tables I and II"},{"comment":"There is an internal tension: Section VI.A.3 states that AI agents 'can serve as valuable tools' for deployment by suggesting gas optimizations and platform-compatibility adjustments (citing [63] and [73]), yet T-2 asserts that 'no AI agents have yet addressed' deployment. The authors should either identify concrete deployment-focused agent systems or explicitly explain why the cited works are not agents under their definition; as written, the takeaway is contradicted by their own text.","section":"Section VI.A.3 and T-2"},{"comment":"The taxonomy is presented as exhaustive and the reference architecture as applicable to all surveyed systems, but no evidence is given that the three archetypes (conversational, instruction-following, goal-directed) are orthogonal or that each surveyed system maps cleanly onto one category. A mapping table, or at least a worked classification of every entry in Table I, would make this load-bearing contribution auditable.","section":"Section IV.A"},{"comment":"The 'first' claim is not substantiated beyond a brief review of five related surveys in Section III, and the abstract/conclusion assert the claim without qualification. The authors should either provide evidence from a broader related-work search or hedge the claim explicitly (e.g., 'to the best of our knowledge'), which is especially important given the undocumented corpus selection.","section":"Abstract and Section VIII"}],"minor_comments":[{"comment":"The text 'for many sections' should read 'for many sectors'.","section":"Section II.A"},{"comment":"The heading 'Non-Normative Example' is unusual for a SoK; consider renaming it to 'Worked Example' or 'Illustrative Example'.","section":"Section IV.B.2"},{"comment":"The description of Nguyen et al. [38] is repeated immediately after being introduced; consider consolidating the two passages.","section":"Section V.A.1"},{"comment":"The caption lists components such as 'Natural Language Processing' and 'Security & Validation' that are not all visible in the figure's layer diagram; align the caption with the figure content.","section":"Figure 2 caption"},{"comment":"Rust is mentioned as a low-resource smart contract language, but Rust was not discussed earlier in the paper's language coverage; add a brief explanation or keep the language list consistent.","section":"Section VII.A"},{"comment":"The abstract claims 'first Systematization of Knowledge' while the conclusion claims 'first comprehensive systematization'; the statements should be unified and hedged consistently with the corpus limitations.","section":"Abstract and Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The main reason for major revision is the load-bearing corpus-completeness issue: a SoK's contributions are the comprehensiveness of its coverage and the reliability of its gap analysis, and the manuscript provides no reproducibility basis for either. If the authors can document their search process and reconcile the T-2 tension with Section VI.A.3, the paper could be acceptable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent and genuinely useful SoK, the first I know of that focuses specifically on AI agents for blockchain with a security and privacy lens. The taxonomy (conversational, instruction-following, goal-directed) and the four-layer reference architecture are sensible organizing devices, and the survey does a real service by pulling representative systems, LLMs, datasets, and threat categories into two readable tables. It deserves a serious referee, but not without revision.\n\nWhat is new: nobody has put this intersection together in a SoK before. The authors take existing agent frameworks (Cheng et al., Eliza, Nguyen et al.) and map them onto a blockchain-specific structure, then attach security and privacy concerns to that structure. The dataset discussion in Section VI.C is more concrete than most surveys, with real caveats about Solidity bias, noise in crawled datasets, and reproducibility. The descriptive claims I checked are consistent with the cited works.\n\nThe main soft spot is methodology. There is no search protocol, no inclusion/exclusion criteria, no databases or time window. Tables I and II say \"representative\" but don't say how representativeness was determined. That matters because a SoK's value is its coverage, and because the paper makes negative claims. T-2 states that \"no AI agents have yet addressed\" smart contract deployment. Yet Section VI.A.3 says AI agents \"can serve as valuable tools\" for deployment by suggesting gas optimizations and adjustments, citing [63] and [73]. That is a tension the authors need to resolve: either identify concrete deployment-focused agents or explain why those cited works are not agents. As written, the gap claim is under-supported.\n\nI'd also soften \"first\" and \"comprehensive\" until the search is documented. It may well be the first; but those are empirical claims about the literature, and the reader cannot audit them from the manuscript. The threat catalog is competent but mostly synthesizes other people's findings, which is fine for a SoK; just don't expect new attack analysis. The citation pattern looks fine; the self-citation to [2] is not a problem.\n\nBottom line: worth engaging with. It fills a real gap, and the taxonomy and architecture will help researchers and developers entering this area. A serious referee should ask for a methodology section, a reconciliation of the deployment contradiction, and toned-down novelty claims. I would accept after revision.","headline":"Useful, timely SoK with a sensible taxonomy, but the unverified 'comprehensive' claim and an internal contradiction in the deployment gap need fixing before it can be trusted as authoritative.","tokens_in":19112,"tokens_out":1768,"would_cite":true,"duration_ms":17436,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents the first systematization of AI agents for blockchain, organizing the field by agent autonomy and a four-layer architecture with security and privacy as the central lens.","keywords":["AI agents","blockchain","large language models","smart contracts","security","privacy","taxonomy","DeFi"],"falsifier":"A literature search that finds a peer-reviewed AI-agent system for smart-contract deployment published before this paper would contradict takeaway T-2; more broadly, applying explicit inclusion and exclusion criteria to the same search and showing that the taxonomy's cells leave a substantial fraction of systems unclassifiable would test the completeness claim.","tokens_in":18170,"feed_emoji":"🤖","tokens_out":5628,"duration_ms":49177,"temperature":0.7,"pith_summary":"AI agents that converse with users and move money or deploy contracts on blockchains are proliferating, but the paper argues no survey has treated them as their own domain with security and privacy as the lens. It claims to deliver the first Systematization of Knowledge for AI-driven blockchain systems, organizing the field with a three-level taxonomy of agents (conversational, instruction-following, goal-directed) and a four-layer reference architecture. The survey then maps representative systems and known threats onto that structure and identifies concrete gaps, including the absence of AI agents that handle smart-contract deployment and the lack of common benchmark datasets. A sympathetic reader would take the paper's contribution to be an organizing frame: once agents are sorted by autonomy and by layer, the field's open security problems become visible and comparable.","feed_headline":"Survey maps AI agents for blockchain into three types","feed_subtitle":"A new taxonomy and reference architecture sort the field and expose open security gaps.","key_machinery":"The organizing device is the autonomy-based taxonomy overlaid on a modular reference architecture. The taxonomy splits AI4B systems into conversational, instruction-following, and goal-directed agents; the architecture supplies four layers (Application, AI Agent, Blockchain Interaction, Blockchain) and names the internal modules — planner, memory/context, validator, tool controller, evaluator/observer, wallet integration, human-in-the-loop — that security mechanisms attach to. This combined classification does the work of the paper: it positions every surveyed system in a shared structure, makes threat patterns legible by component, and turns 'there is no agent for X' into a checkable statement about a cell in the taxonomy.","core_discovery":"The paper's central claim is that the intersection 'AI agents for blockchain' is a distinct research area that existing surveys miss, and that it can be systematically organized. Its proposed taxonomy classifies agents by how they treat user input — as a question (conversational agents, read-only), as an instruction (instruction-following agents that build and submit transactions), or as a goal (goal-directed agents that plan and execute multi-step strategies autonomously). The accompanying reference architecture places these agents in four layers — Application, AI Agent, Blockchain Interaction, and Blockchain — with internal components such as planner, validator, memory, tool controller, and human-in-the-loop mechanisms. The paper uses this frame to review applications from blockchain data analysis and supply-chain traceability to DeFi portfolio management, DAO governance, and smart-contract development and auditing, and to enumerate threats such as prompt injection, fake-memory/context manipulation, privacy leakage, and autonomy-induced market risk. It concludes with gaps: no common benchmark for auditing agents, low-resource languages underserved, and no AI agent yet addressing smart-contract deployment.","pith_inferences":["The autonomy axis implies a testable risk gradient: moving from conversational to instruction-following to goal-directed shifts the dominant failure mode from wrong information to unauthorized or irreversible transactions, a hypothesis one could quantify by measuring failure costs across the surveyed systems.","The paper's negative claim that no AI agent handles smart-contract deployment is timestamped; as tool-use and wallet-integration standards mature, that cell is a likely place for rapid filling, so the gap should be rechecked periodically rather than treated as permanent.","Privacy leakage through logs and over-access to wallet data may turn out to be the binding constraint for regulated use; the architecture's explicit placement of privacy modules suggests a design rule that agents should request the minimum chain data needed for the current task.","The taxonomy could be extended into a maturity scale for blockchain agents, with each autonomy level associated with required safety controls; the paper does not propose such a scale, but its categories make it straightforward to construct."],"forward_implications":["Future work on blockchain agents can position new systems in the proposed taxonomy and architecture instead of re-describing them from scratch.","The identified gaps become a research agenda: smart-contract deployment support, low-resource languages like Vyper and Move, and community benchmark datasets are the paper's named open problems.","Security defenses can be mapped to architecture components — validator and trust-score modules, human-in-the-loop signing, memory integrity — giving designers a checklist rather than an ad hoc list of attacks.","If the taxonomy is adopted, comparing auditing or trading agents becomes a matter of comparing systems inside the same cell, which would sharpen evaluation."],"supporting_citations":[{"why":"General AI-agent protocol survey used to show prior overviews are not blockchain-specific.","marker":"[11]"},{"why":"Security/privacy-focused agent survey that mentions blockchain only as an audit tool, establishing the gap this paper fills.","marker":"[14]"},{"why":"Surveys blockchain for LLM security, showing the nearest related work centers on the model rather than on agents.","marker":"[35]"},{"why":"Provides the Planning, Memory, Rethinking, Environments, Action framework the paper aligns with its AI Agent Layer.","marker":"[36]"},{"why":"ElizaOS supplies the reference architecture's example of validator/trust-score components and blockchain API integration.","marker":"[37]"},{"why":"Multi-agent chatbot for blockchain APIs; used as the non-normative instruction-following example with human-in-the-loop wallet approval.","marker":"[38]"},{"why":"Goal-directed portfolio-management agent with dual validation via confidence scoring and agent voting; grounds the top taxonomy class.","marker":"[41]"},{"why":"LLM-SmartAudit multi-agent auditing system; representative of the smart-contract auditing application and its dataset issues.","marker":"[6]"}],"fun_headline_variants":["AI agents for blockchain: taxonomy of three, with security gaps","Three types of AI agents for blockchain, from queries to goals","First SoK exposes prompt injection and other AI-blockchain risks","AI blockchain agents: a new reference architecture and threat map"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"That the surveyed set of representative systems is complete enough to support 'first' and 'comprehensive,' and to support negative gap claims such as no AI agent addressing smart-contract deployment.","fun_headline_variants_meta":{"raw":{"variants":["AI agents for blockchain: taxonomy of three, with security gaps","Three types of AI agents for blockchain, from queries to goals","First SoK exposes prompt injection and other AI-blockchain risks","AI blockchain agents: a new reference architecture and threat map"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000209,"raw_usage":{"total_tokens":1390,"prompt_tokens":912,"completion_tokens":478,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":408}},"tokens_in":528,"tokens_out":478,"duration_ms":4558,"temperature":1.0,"reasoning_tokens":408,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:12:26.743999+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A literature search that finds a peer-reviewed AI-agent system for smart-contract deployment published before this paper would contradict takeaway T-2; more broadly, applying explicit inclusion and exclusion criteria to the same search and showing that the taxonomy's cells leave a substantial fraction of systems unclassifiable would test the completeness claim.","supporting_citations":[{"cited_title":"Multi-agent Chatbot for Efficient Interaction with Blockchain APIs,","cited_arxiv_id":null,"evidence_quote":"Multi-agent chatbot for blockchain APIs; used as the non-normative instruction-following example with human-in-the-loop wallet approval."}],"review_version":2}