{"id":"98bce461-13bb-4b77-b169-790b6cc7729f","arxiv_id":"2604.19667","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A real-world industrial benchmark and agentic baseline show current LLMs often fail to produce correct, stable, deployable visual workflows from natural language, with only modest resolve-rate gains.","lead":"Chat2Workflow is a benchmark that tests whether language models can turn natural-language requests into executable visual workflows for industrial platforms like Dify and Coze. It matters because manual workflow engineering is costly, and a reliable automation path would change how business systems are built.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Resolve rate's claim to measure true industrial deployability hinges on an uninspectable transform-and-deploy pipeline whose fidelity to Dify/Coze is asserted but not evidenced in the abstract.","rationale":"The reader's weakest_assumption correctly isolates the single load-bearing premise: that the real-world instances plus transform-and-deploy pipeline validly measure industrial executability rather than benchmark-specific success. From the abstract alone no internal contradiction or circularity is visible; the risk is purely empirical validity of the metric. Because full methods, data, and harness remain unavailable, the UNVERDICTED/LOW-confidence status is appropriate and no adjustment is warranted. The proposed concrete test directly probes that premise once artifacts are obtained.","tokens_in":2045,"tokens_out":427,"duration_ms":13368,"concrete_test":"Clone the linked GitHub repo, execute the evaluation harness on a random sample of 20 'resolved' instances, then manually import the resulting workflows into Dify and Coze; record the fraction that run to completion without any manual edit. If that fraction falls below ~80%, the resolve-rate metric overstates industrial readiness and the headline gap claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that SOTA models capture high-level intent yet fail to produce correct, stable, executable workflows, and that an agentic baseline's 6.05% resolve-rate gain still leaves a large real-world gap—rests on the premise that each curated instance is constructed so a generated workflow can be transformed and directly deployed. The abstract supplies no specification of the transformation rules, platform schema constraints, scoring of partial failures (syntax vs. control flow vs. runtime), or how multi-round requirement evolution is simulated. If 'resolve' largely rewards benchmark-specific formatting rather than end-to-end executability under realistic platform APIs, both the reported gap and the modest gain cease to support the industrial-automation conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The manuscript introduces Chat2Workflow, a benchmark for multi-round natural-language generation of executable visual workflows, constructed from real-world business workflows and designed so that outputs can be transformed and deployed to industrial platforms such as Dify and Coze. It reports that state-of-the-art language models often capture high-level intent but struggle to produce correct, stable, and executable workflows under complex and evolving requirements, and proposes an agentic baseline that improves resolve rate by up to 6.05%, while arguing that a substantial real-world gap remains. Code is released at a public repository.","tokens_in":2216,"tokens_out":846,"duration_ms":15317,"significance":"If the benchmark construction, deployability pipeline, and evaluation protocol hold under full scrutiny, Chat2Workflow would be a practically relevant contribution for industrial automation of visual workflow authoring—an area where manual multi-round engineering is costly and error-prone. Explicit framing around evolving requirements, platform deployability (Dify/Coze), and a public code release are strengths. The reported modest agentic gain alongside a large residual gap would usefully position the resource as a foundation for subsequent work rather than a solved task.","major_comments":[{"comment":"The abstract’s central industrial claim—that generated workflows can be transformed and directly deployed to platforms such as Dify and Coze—depends on an uninspectable transform-and-deploy pipeline. Without a full-paper specification of transformation rules, platform schema constraints, and how syntax vs. control-flow vs. runtime failures are scored, it is not possible to verify that “resolve rate” measures true deployable correctness rather than benchmark-specific formatting success. This premise is load-bearing for both the reported SOTA gap and the 6.05% gain.","section":"Abstract"},{"comment":"Resolve rate is presented as the primary success measure for multi-round NL-to-workflow generation under evolving requirements, but the abstract does not define how multi-round requirement evolution is simulated, how partial credit is assigned, or what constitutes a resolved instance. Without these definitions (and associated dataset size, splits, and error bars), the 6.05% gain and the “large real-world gap” claim cannot be assessed for statistical or practical significance.","section":"Abstract"},{"comment":"The abstract asserts construction from a “large collection of real-world business workflows” and SOTA struggle under “complex and evolving requirements,” yet supplies no instance counts, complexity stratification, baseline inventory, or ablation of the agentic method. These elements are load-bearing for the claim that the residual gap is industrial rather than artifactual; they must appear with reproducible detail in the full manuscript.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract is generally clear, but “resolve rate” should be briefly glossed on first use (e.g., end-to-end deployable success under the stated transform pipeline) so readers can interpret the 6.05% figure without the full paper.","section":"Abstract"},{"comment":"Naming the concrete SOTA models and the agentic baseline architecture in the abstract would help readers gauge the strength of the comparison before consulting the full text.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"Only the abstract was available for this review; the full manuscript could not be inspected. The recommendation is therefore uncertain rather than a content verdict. The stress-test concern about resolve-rate fidelity to real Dify/Coze deployability is material and should be checked first when the full paper is obtained. If the full paper supplies a transparent transform pipeline, metric definitions, dataset statistics, and reproducible evaluation, the contribution may be suitable for major or minor revision rather than rejection; if those elements remain underspecified, the industrial-automation claim would not be supportable."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing to know: this is a new evaluation artifact for turning natural language into executable visual workflows that claim direct deployability to Dify/Coze, plus a small agentic baseline lift (up to 6.05% resolve rate) and an honest admission that a large gap remains. That is the contribution.\n\nWhat is actually new is the task framing and the industrial grounding. Manual workflow engineering is expensive and real; a benchmark built from real business workflows, with multi-round evolving requirements and a transform-and-deploy path to actual platforms, is a sensible measurement instrument for LLM agents and low-code automation. The abstract does not overclaim a paradigm shift. It reports that SOTA models often get high-level intent but fail on correctness, stability, and executability, and that their agentic baseline helps only modestly. That posture is useful. Shipping code is a plus.\n\nThe soft spots are exactly what you would expect from abstract-only. We cannot see dataset size, splits, metric definitions, how multi-round evolution is simulated, what “resolve” scores (syntax vs control flow vs runtime), or the fidelity of the transform pipeline to Dify/Coze schemas. The stress-test concern is fair as a risk: if resolve mostly rewards benchmark-specific formatting rather than end-to-end platform executability, both the gap and the 6.05% gain lose industrial force. That is not proven failure; it is missing evidence. Circularity looks ordinary for a benchmark paper, not load-bearing self-definition. Novelty and significance sit in the mid range—solid applied evaluation, not first-principles work.\n\nWho it is for: people building LLM agents for enterprise automation, low-code platforms, and anyone who needs a harder, more deployable workflow generation test than pure text-to-code. A serious referee should see the full paper, data, and harness. I would not desk-reject on the abstract; I would send it out and demand the metric and pipeline details. Bring it to reading group only if someone is actively working on agentic workflow tools; otherwise wait for the full artifact. I would not cite it yet from the abstract alone, but I would track the repo.","headline":"Useful industrial NL-to-workflow benchmark framing with a modest agentic gain, but abstract-only so the deployability claim is still uncheckable.","tokens_in":2846,"tokens_out":547,"would_cite":false,"duration_ms":4730,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Large language models still cannot reliably turn natural language into correct, stable, and deployable visual workflows, even with an agentic baseline.","keywords":["executable visual workflows","natural language to workflow","Chat2Workflow","agentic baseline","workflow automation","large language models","multi-round interaction","industrial deployment"],"falsifier":"Take a held-out set of live production workflows from Dify or Coze, generate candidates with the same models and agent, deploy them without manual repair, and check whether end-to-end execution success still matches the reported resolve rates.","tokens_in":2918,"feed_emoji":"⚙️","tokens_out":817,"duration_ms":18906,"temperature":0.7,"pith_summary":"The paper introduces Chat2Workflow, a benchmark built from real-world business workflows so that any generated workflow can be transformed and deployed directly onto industrial platforms such as Dify and Coze. It shows that state-of-the-art language models often capture high-level intent yet fail to produce correct, stable, and executable visual workflows when requirements are complex or evolve across multi-round interaction. An agentic baseline proposed by the authors raises resolve rate by as much as 6.05 percent, but a substantial practical gap remains. The work therefore positions the benchmark as a foundation for measuring and advancing industrial-grade automation of workflow construction that today is still done almost entirely by hand.","feed_headline":"LLMs still fail to emit deployable visual workflows","feed_subtitle":"A real-world benchmark shows only a 6% agent gain and a large gap for platforms like Dify and Coze.","key_machinery":"Chat2Workflow—the benchmark of real-world business workflow instances, each constructed so a generated workflow can be transformed and directly deployed to platforms such as Dify and Coze—together with the agentic baseline that iteratively refines the workflow under multi-round dialogue.","core_discovery":"State-of-the-art language models can often capture high-level intent from natural language descriptions of business processes, yet they struggle to generate correct, stable, and executable visual workflows—especially under complex and evolving requirements. The authors formalize this gap with Chat2Workflow, a benchmark whose instances are designed for direct deployment on platforms such as Dify and Coze, and show that an agentic baseline improves resolve rate by up to 6.05 percent while leaving a large real-world gap.","pith_inferences":["Closing the remaining gap may require tighter coupling of formal workflow verification with language-model generation rather than pure agentic prompting.","Success on this benchmark would cut the manual engineering cost of building production visual workflows.","Similar deployability-first benchmarks could be built for other visual or low-code programming domains.","The modest 6.05 percent gain suggests iterative agent scaffolding alone is insufficient and that new inductive biases may be needed."],"forward_implications":["Resolve-rate numbers on Chat2Workflow become a concrete yardstick for industrial workflow automation progress.","Future models or agents can be scored on whether they close the remaining gap before any auto-generated workflow is deployed.","The transform-and-deploy pipeline lets generated outputs be tested on live platforms rather than only in simulation.","Multi-round requirement evolution is treated as a first-class evaluation axis rather than a side concern."],"fun_headline_variants":["LLMs grasp intent but fail to emit deployable visual workflows","Chat2Workflow: SOTA models struggle with executable workflows","Only 6% agent gain leaves large gap for Dify-style workflows","High-level intent captured, yet workflows remain non-executable","Complex requirements expose LLMs' limit on stable visual workflows"],"cache_read_input_tokens":0,"weakest_assumption_plain":"That the curated real-world business instances and the transform-and-deploy pipeline to platforms such as Dify and Coze truly measure industrial executability, so resolve rate reflects deployable correctness rather than benchmark-specific formatting success.","fun_headline_variants_meta":{"raw":{"variants":["LLMs grasp intent but fail to emit deployable visual workflows","Chat2Workflow: SOTA models struggle with executable workflows","Only 6% agent gain leaves large gap for Dify-style workflows","High-level intent captured, yet workflows remain non-executable","Complex requirements expose LLMs' limit on stable visual workflows"]},"model":"grok-4.5","effort":"low","cost_usd":0.0043,"raw_usage":{"total_tokens":1322,"prompt_tokens":812,"num_sources_used":0,"completion_tokens":89,"cost_in_usd_ticks":43000000,"prompt_tokens_details":{"text_tokens":812,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":421,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":812,"tokens_out":89,"duration_ms":4899,"temperature":1.0,"reasoning_tokens":421,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T18:49:17.482433+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Take a held-out set of live production workflows from Dify or Coze, generate candidates with the same models and agent, deploy them without manual repair, and check whether end-to-end execution success still matches the reported resolve rates.","supporting_citations":[],"review_version":3}