{"id":"24abdab4-cd43-42b1-b547-cbed38f017f5","arxiv_id":"2508.01186","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review that classifies 24 agent workflow systems along functional and architectural axes and argues for standardization, optimization, and security work.","lead":"This preprint surveys LLM-based agent workflow systems, sorting 24 frameworks by functional capabilities and architectural features. It is a useful map of an emerging field, but it contributes a taxonomy and observations rather than new measurements or results.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing comparison tables lack a defined sample and contain concrete annotation errors, so the standardization claim rests on unverified data.","rationale":"The paper's unique contribution is its structured comparison; without the tables, the survey is mostly background and opinion. The standardization claim is supported by observing that many independently built systems do not share a common workflow representation, but that observation only carries weight if the 24 rows are a legitimate sample and the annotations are accurate. The manuscript gives no inclusion criteria, and the rows are heterogeneous in kind: prompting strategies, a no-code engine, a commercial product, compilers, and SDKs are all lumped together. Applying a uniform feature checklist across these artifact types conflates 'the system supports X' with 'the paper demonstrates X' and makes counts, percentages, and generalizations such as 'most systems adopt self-defined DSLs' unreliable. Concrete annotation errors (ReWoo 2024 vs. the cited 2023 paper; duplicated n8n reference; the uncertain Agno/Phidata distinction) show that the cells were not checked against primary sources. The Section VIII.A.1 'current system / APO' passage is an additional signal that portions of the paper were carried over from another project; it does not directly refute the survey's thesis, but it weakens confidence in the paper's self-assessment. These issues are serious but fixable, so the reader's CONDITIONAL verdict remains appropriate: the survey's high-level message is plausible, but its empirical backbone needs a reproducibility audit before the comparison can be treated as reliable.","tokens_in":15415,"tokens_out":8007,"duration_ms":102339,"concrete_test":"Run a reproducibility audit: reconstruct the 24-row table from primary sources (official repositories and cited papers), state explicit inclusion criteria for 'agent workflow system', and have two independent annotators re-annotate every cell with a citation. Specifically verify ReWoo's year, whether Phidata and Agno are distinct systems, and whether ReAct/ReWoo meet the inclusion criterion. If any of these checks confirms an error, or the annotators disagree on more than 5% of cells, the tables need revision before the standardization conclusion can be relied on.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that agent workflows lack a unified specification is supported empirically by the claim that 24 representative systems have been compared in Tables 1 and 2. For that comparison to carry weight, the rows must be a well-defined sample and the √/◑/× annotations must be reproducible. Neither condition is met. The paper never defines 'agent workflow system' or gives inclusion/exclusion criteria, and the table mixes prompting strategies (ReAct is a prompting framework, ReWoo a reasoning-without-observation strategy), a non-LLM automation shell (n8n), a closed commercial product (DeepResearch), a compiler (DSPy), and SDKs (AutoGen, CrewAI). A uniform checklist of GUI, API, Open Source, Year, and deployment is then applied across those artifact types, so aggregate statements such as 'most systems adopt self-defined DSLs or configuration formats' are not supported by a defined corpus. Concrete annotation problems confirm that the cells were not checked against primary sources: ReWoo is listed as 2024 although the cited paper is arXiv:2305.18323 from 2023; reference [68] for n8n is duplicated; Phidata [71] and Agno [56] are both credited to 'Agno Contributors' with sibling repository URLs and need to be checked for double counting. Section VIII.A.1 also contains a system-specific sentence about 'the current system' and 'the Automated Prompt Optimization (APO) process' that has no counterpart in this survey, making the self-reported limitations unreliable. These issues do not disprove the high-level standardization message, but they mean the comparison tables cannot yet be treated as verified evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews the emerging area of agent workflow systems, arguing that workflows are becoming the backbone of the AI ecosystem while lacking a unified specification. It proposes a taxonomy based on functional capabilities and architectural features, compares 24 systems in two large tables, and discusses workflow optimization, application domains, security, limitations, and future directions such as standardization and multimodal integration.","tokens_in":15690,"tokens_out":3628,"duration_ms":47148,"significance":"If the comparative analysis were backed by a reproducible annotation protocol, the paper would provide a useful structured overview of a fast-moving field. The high-level taxonomy is plausible, the coverage of optimization and security issues is reasonable, and the survey draws on a broad set of academic and industrial systems. The main value lies in the two comparison tables, but their reliability is currently the weakest point: the sample is not defined and several concrete annotation errors are visible. The central claim about the lack of a unified workflow specification is plausible, but as presented it rests on an unauditable comparison.","major_comments":[{"comment":"The selection of the 24 systems is not defined. The paper never states inclusion or exclusion criteria for what counts as an \"agent workflow system,\" and the tables mix heterogeneous artifact types: ReAct and ReWoo are prompting/reasoning strategies, n8n is a non-LLM workflow automation shell, DeepResearch is a closed commercial product, DSPy is a compiler, and AutoGen/CrewAI are SDKs. A uniform checklist of GUI, API, Open Source, Year, and deployment is then applied across these different kinds of artifacts. Consequently, aggregate statements such as \"most systems adopt self-defined DSLs or configuration formats\" (Section VIII.B) are not supported by a well-defined corpus. Please add a methodology subsection describing the corpus construction, artifact-type classification, annotation rules, and per-cell evidence, or provide an appendix with justifications for each row.","section":"Section IV, Tables 1 and 2"},{"comment":"The table annotations contain concrete errors that undermine reproducibility. ReWoo is listed as 2024 in Table 1, but the cited reference [23] is arXiv:2305.18323 from 2023. Reference [68] for n8n is duplicated in the bibliography. Phidata [71] and Agno [56] are both credited to \"Agno Contributors\" and point to sibling github.com/agno-agi repositories, so the two rows may double-count a single project or need explicit clarification. These issues indicate that at least some cells were not checked against primary sources, and they weaken confidence in the comparison as a whole.","section":"Table 1 and reference list"},{"comment":"The first listed limitation, \"Lack of Environmental Feedback,\" states that \"the current system does not consider incorporating environmental feedback into the Automated Prompt Optimization (APO) process.\" No such system and no APO process are described anywhere in this survey, so this sentence appears to be carried over from a different paper. This makes the self-reported limitations section unreliable as a description of the survey's own limitations. Please either rewrite the item so it applies to the surveyed field as a whole, or remove it.","section":"Section VIII.A.1"}],"minor_comments":[{"comment":"There are several typographical and grammatical issues, including \"Muti-agent path finding\" (should be \"Multi-agent path finding\"), \"as it is a source project\" (should be \"as it is an open-source project\"), \"comprehend comparison\" (should be \"comprehensive comparison\"), \"To dealing with\" (should be \"To deal with\"), and \"are not inline with expectations\" (should be \"are not in line with expectations\").","section":"Throughout"},{"comment":"The section numbering is inconsistent: the text refers to \"Section 2,\" \"Section 3,\" etc., while the actual headings use Roman numerals (II, III, IV). Please unify the cross-referencing style.","section":"Section II vs Section III"},{"comment":"The same system is referred to as \"DeepResearch\" in Table 1 and \"Deep Research\" in Table 2; please standardize the name. Similarly, \"AgentUniverse\" appears with different capitalizations in the text and tables.","section":"Table 1 and Table 2"},{"comment":"The sentence \"Swift [74], VDL [75] use a functional-flavored scripting language to concisely describe large scale scientific workflows\" lacks a verb tense agreement and reads as an incomplete example; please revise for clarity.","section":"Section III.C.1"},{"comment":"The discussion of MAPF and CBS, while interesting, is not clearly tied to agent workflow systems as defined in the rest of the paper; a connecting sentence explaining why this path-finding strategy is a workflow-level concern would improve readability.","section":"Section VIII.A.5"}],"recommendation":"major_revision","confidential_remarks":"The paper has a plausible taxonomy and broad coverage, but its main empirical contribution—the two comparison tables—needs an auditable protocol and correction of the identified annotation errors. I would be willing to reconsider after a revision that provides selection criteria, per-cell evidence, and a cleanup of the reference list and the misplaced APO sentence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick take: this is a genuinely useful orientation piece for someone new to agent workflows, and the two-axis comparison (functional vs architectural) is a modest but real framing contribution. The coverage is broad—24 systems, with extra attention to optimization and security that many agent surveys skimp on. The standardization argument is plausible and consistent with what industry discussions (A2A, MCP) are already doing.\n\nThe soft spots are mostly in the comparison tables. There is no stated inclusion/exclusion criterion for the 24 systems, and the rows mix very different artifact types: prompting strategies (ReAct), automation shells (n8n), SDKs (AutoGen, CrewAI), a compiler (DSPy), and closed products (DeepResearch). That makes aggregate claims like 'most systems adopt self-defined DSLs' hard to interpret. More concretely, the tables have verification errors: ReWoo is listed as 2024 even though the cited paper is from 2023, reference [68] appears twice, and Phidata and Agno are both attributed to the same contributors with sibling URLs, looking like a possible double count. These are fixable, but they undermine confidence in the specific cells. The citation pattern is otherwise reasonable; the Swift/VDL self-citations are on-topic and not a concern.\n\nThere is also one leftover sentence in Section VIII.A.1 about 'the current system' and an 'Automated Prompt Optimization (APO) process' that has no place in a survey. That, plus the absence of any comparison with prior agent or workflow surveys, makes the paper feel less carefully positioned than it should be.\n\nThe high-level message—workflows are becoming the backbone of AI ecosystems and standardization is missing—stands on its own and doesn't depend on the flawed tables. So I wouldn't reject the paper on that basis. But the tables should be either properly defined and checked, or downgraded to 'illustrative examples' with a clear caveat.\n\nWho is this for? A newcomer wanting a structured map of the area, or an instructor building a course module. Not for someone who needs a reliable feature-by-feature comparison.\n\nI'd engage with it in peer review: the topic is timely, and with a methodology section and error fixes it could be a solid survey. I wouldn't cite the tables in their current form, though. Recommendation: send it to review with a request for a major revision that tightens the sample definition, verifies the table entries, and cleans up the stray text.","headline":"A useful broad map of agent workflows, but the comparison tables aren't yet auditable; fixable with methodology and corrections.","tokens_in":16229,"tokens_out":4636,"would_cite":false,"duration_ms":48909,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Agent workflows have become the AI ecosystem's backbone, and they lack a shared standard, a 24-system survey finds.","keywords":["agent workflow","multi-agent systems","LLM agents","orchestration","standardization","workflow optimization","workflow security","survey"],"falsifier":"Take two mature systems from the paper's tables that use different workflow representations and try to execute the same workflow—say a planner–executor loop with memory and a critic—in both, with the same external tools and no custom adapter; if one system can import and run the other's workflow definition directly, the claimed absence of a unified workflow framework would be seriously weakened.","tokens_in":15241,"feed_emoji":"🤖","tokens_out":9799,"duration_ms":110534,"temperature":0.7,"pith_summary":"This paper argues that agent workflows—structured orchestration layers in which LLM-driven agents plan, call tools, share memory, and cooperate—have become the backbone of the emerging AI ecosystem, yet the field has no unified specification. Its contribution is a systematic comparison of 24 agent workflow systems, scored on functional capabilities (planning, tool use, multi-agent collaboration, memory, GUI interaction, API calls, self-reflection, custom tools, cross-platform support, open-source status) and on architectural and mechanistic features (agent roles, flow type, representation, language, protocol, deployment). The comparison is meant to reveal common patterns and to pinpoint the bottleneck: most systems use self-defined DSLs, prompt fragments, or incompatible execution interfaces, making workflows hard to verify, debug, reuse, or orchestrate across frameworks. On that basis the paper derives a research agenda around standardization, workflow-level optimization, security, and open problems such as multimodal integration and autonomous pervasive agents.","feed_headline":"Agent workflows lack a shared standard, 24-system survey shows","feed_subtitle":"Comparing capabilities and architectures exposes why cross-platform agents still cannot easily interoperate.","key_machinery":"The central machinery is the two comparison tables that score 24 agent workflow systems with a four-value mark system (√ supported, × not supported, ◑ partially supported, ⚪ unspecified, with √* meaning only one specific API is supported). The capability table covers planning, tool use, multi-agent, memory, GUI, API, self-reflection, custom tools, cross-platform, and open-source status; the architecture table covers agent roles, flow type, representation, language, protocol, and deployment. These tables let the survey move from individual examples to aggregate patterns, such as near-universal support for planning and tool use alongside fragmented protocols and workflow representations. The supporting conceptual grid is Section III's taxonomy of workflow modes (chain, parallel, routing, orchestrator–workers, evaluator–optimizer), agent roles (planner, executor, critic, memory manager, communicator), and the three-layer architecture of UI/UX, workflow management, and agent collaboration.","core_discovery":"The claim this paper is trying to establish is that agent workflow systems are moving from scattered practices toward the structural core of AI applications, but are blocked by a missing shared framework for describing, executing, and verifying workflows. The paper supports this by classifying and comparing 24 systems along two dimensions: what the workflow can do (planning, tool use, multi-agent collaboration, memory, GUI, API, self-reflection, custom tools, cross-platform, open source) and how the workflow is built (agent roles, orchestration flow, representation formalism, language, protocol, deployment). Its reading of the comparison is that capability support is converging—most systems plan, use tools, and support multiple agents—while representation and protocol choices remain fragmented, which is exactly the gap the paper names the 'absence of a unified workflow framework.' From that diagnosis it moves to optimization strategies (manual reconstruction, heuristics, Bayesian optimization, LLM-based generative optimization), security concerns at the tool, protocol, MCP-server, LLM, memory, and multi-agent levels, and future directions centered on standardization and interoperability.","pith_inferences":["An implication the authors leave implicit is that systems offering formal, machine-readable workflow representations should compose more easily than prompt-fragment-based ones; a direct cross-framework porting experiment could test that ordering.","A natural extension would be to turn each mark in the comparison tables into a reproducible probe, so the taxonomy becomes a benchmark the community can re-run and extend rather than a static expert judgment.","If the standardization diagnosis is correct, public open-source activity should show convergence toward a small set of workflow description formats and communication protocols over the next few years; measuring that trend would independently validate the paper's central concern."],"forward_implications":["A unified workflow framework would let a workflow defined once run on different agent platforms, turning today's isolated agent systems into a composable ecosystem.","Standardized protocols for tool access and agent-to-agent communication would allow agents from different vendors to share tools and context, which the paper identifies as a precondition for scaling beyond single-vendor deployments.","Workflow-level optimization—token budgets, multi-task scheduling, resource-aware allocation, adaptive schedulers—would become a recognized research area in its own right.","Security would have to be enforced inside the orchestration layer, defending against tool poisoning, MCP-server impersonation, and memory-poisoning attacks rather than only at the model level.","Progress toward autonomous pervasive agents depends on workflows that adapt dynamically at execution time instead of following fixed predefined chains."],"supporting_citations":[{"why":"Supplies the paper's working definitions of agent and workflow, which anchor the entire taxonomy.","marker":"[4]"},{"why":"Provides the AutoGen multi-agent conversation case used in the comparison tables and the OptiGuide example.","marker":"[3]"},{"why":"Defines the ReAct interleaved reasoning-and-acting loop that the survey treats as a baseline pattern across many systems.","marker":"[14]"},{"why":"Gives the ReWoo planner–executor reasoning workflow, used in the performance snapshots and architecture comparison.","marker":"[23]"},{"why":"Prior survey of LLM-based multi-agent systems that frames the workflow, infrastructure, and challenges discussion.","marker":"[6]"},{"why":"Source for the workflow management system definition and formal workflow language background in Section III.","marker":"[10]"},{"why":"Google's A2A protocol cited as industry evidence of movement toward agent interoperability and standardization.","marker":"[76]"},{"why":"The Model Context Protocol specification used as a key protocol example and as the target of the security analysis.","marker":"[40]"},{"why":"Survey of AI agent security that underpins the tool-poisoning and cross-server attack claims.","marker":"[41]"}],"fun_headline_variants":["24 agent workflows, one missing standard","Survey finds no unified framework for agent workflows","Agent workflow chaos: 24 systems, zero shared standard","Why agent workflows can't interoperate: missing standard","Agent workflows: fragmented standards, converging capabilities"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 24 systems chosen for the comparison fairly represent the agent workflow landscape and that each feature mark in the tables reflects a consistent, correct reading of very different projects.","fun_headline_variants_meta":{"raw":{"variants":["24 agent workflows, one missing standard","Survey finds no unified framework for agent workflows","Agent workflow chaos: 24 systems, zero shared standard","Why agent workflows can't interoperate: missing standard","Agent workflows: fragmented standards, converging capabilities"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000672,"raw_usage":{"total_tokens":3045,"prompt_tokens":913,"completion_tokens":2132,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":2062}},"tokens_in":529,"tokens_out":2132,"duration_ms":18029,"temperature":1.0,"reasoning_tokens":2062,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T05:46:43.489035+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take two mature systems from the paper's tables that use different workflow representations and try to execute the same workflow—say a planner–executor loop with memory and a critic—in both, with the same external tools and no custom adapter; if one system can import and run the other's workflow definition directly, the claimed absence of a unified workflow framework would be seriously weakened.","supporting_citations":[{"cited_title":"Creating large language model appli-cations utilizing langchain: A primer on developing llm apps fast,","cited_arxiv_id":null,"evidence_quote":"Survey of AI agent security that underpins the tool-poisoning and cross-server attack claims."}],"review_version":1}