{"id":"5ae7f30a-5a42-4250-9d4b-c5f3bd66bc5f","arxiv_id":"2412.03446","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Text2Workflow is a multi-prompt LLM system with human feedback that generates JSON workflows from natural language, scoring 71.3% average semantic accuracy on the authors' 60-request Process2JSON dataset, versus 64.2% for a single-prompt gpt-4o-mini baseline.","lead":"This paper presents Text2Workflow, a prompt-based pipeline that converts natural language requests into JSON workflow steps using large language models. The authors evaluate it on a self-built 60-example dataset and report higher accuracy on complex requests than a single-prompt baseline, at the cost of much higher token use and runtime.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported hard-level advantage is confounded by human-in-the-loop: without user feedback, Text2Workflow-NUA (26.3%) and GC (27.5%) underperform baseline-gpt-4o-mini (30%); only full system with author-provided feedback reaches 57.5%.","rationale":"The strongest positive evidence in the paper is the hard-level accuracy gap (57.5% vs 30%). The reader focused on gold-standard provenance and single-evaluator rubric. I agree those are serious, but the more direct threat to the central claim is the asymmetry in human involvement between conditions. The paper's own ablation (Table 4) shows the automated pipeline without user aid (NUA) scores 26.3% at hard level, below the gpt-4o-mini baseline (30%), and with only general clarification (GC) scores 27.5%, still below baseline. The human feedback loop (HFL) alone reaches 32.5%, and the full system (GC+HFL) reaches 57.5%. This is a 25-point jump that the paper does not explain mechanistically. Because the same single author both provides feedback in the loop and scores the outputs (Section 6), the headline comparison conflates model capability with human assistance and evaluator leniency. The paper's central claim as stated — that the multi-layer prompt pipeline outperforms single-prompt gpt-4o-mini — is therefore not established for the automated components; it is established only for a human-in-the-loop system evaluated by its own author. A fair test would give the baseline the same opportunity for user feedback and use blinded multi-evaluator scoring. This does not negate the value of the dataset or the prompt-engineering effort, but it changes what the 57.5% number means. The verdict remains CONDITIONAL: the claim could survive if the controlled experiment supports it, but as presented the evidence is insufficient.","tokens_in":21434,"tokens_out":6067,"duration_ms":57161,"concrete_test":"Re-run the 20 hard-level Process2JSON requests in three conditions: (1) full Text2Workflow, (2) baseline-gpt-4o-mini augmented with the same High-level Summary and Workflow Modification loop (up to two loops), and (3) Text2Workflow-NUA. Use multiple independent users for feedback and multiple blinded evaluators scoring with Table 2. If condition (2) closes the 27.5-point gap, the claimed advantage is due to human feedback rather than the prompt architecture. Additionally, check whether the 32.5-to-57.5 interaction between GC and HFL replicates across runs and users; if it does not, the headline result is unstable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that Text2Workflow 'achieves the highest performance across experiments' (Section 5, Figure 13), driven by a 57.5% hard-level accuracy versus 30% for baseline-gpt-4o-mini (Table 4). The most load-bearing weakness is the conflation of the automated pipeline with human-in-the-loop assistance in that headline comparison. The ablation in Table 4 shows that without user aid, Text2Workflow-NUA scores 26.3% at hard level, below the 30% of baseline-gpt-4o-mini; with only general clarification (GC) it is 27.5%, still below baseline. Only when the human feedback loop (HFL) is added does the method reach 32.5%, and the full system (GC+HFL) jumps to 57.5%. This 25-point jump from HFL to full system is not explained by any mechanism or analysis in the paper, and no significance testing or multiple-user variance is reported. Since the 'human' in the loop is a single author who both suggests edits and later evaluates the output (Section 6 concedes single-evaluator bias), the reported advantage may reflect human intelligence injected through the feedback loop plus evaluator leniency, rather than the multi-layer prompt architecture. Thus the paper's strongest quantitative evidence for the central claim is confounded: the baseline receives zero human feedback, while the proposed method receives up to two rounds of user-driven correction plus a user clarification step. Without controlling for this asymmetry, the claimed superiority of Text2Workflow over gpt-4o-mini is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Text2Workflow, a multi-stage LLM prompting pipeline that converts natural-language business requests into structured JSON workflows. The pipeline comprises seven prompt layers, including a user request clarification step, a master/expert decomposition into steps, a human feedback loop over the workflow skeleton, and a missing-parameter verification stage. The authors also introduce Process2JSON, a 60-request dataset labeled easy/medium/hard, and compare Text2Workflow against two single-prompt baselines (gpt-3.5-turbo-0125 and gpt-4o-mini). They report that Text2Workflow achieves 71.3% overall JSON accuracy versus 64.2% for the gpt-4o-mini baseline, with the largest gap on hard-level requests (57.5% versus 30%). An ablation separates no user aid (NUA), general clarification (GC), human feedback loop (HFL), and the full system.","tokens_in":21714,"tokens_out":4378,"duration_ms":42023,"significance":"If the reported results are sound, the paper makes a useful practical contribution: a modular, prompt-based workflow-generation pipeline with a reusable schema and a new benchmark dataset, plus a transparency note about the single-evaluator limitation. The strengths are the explicit architectural description, the ablation design that isolates the user-input mechanisms, and the candid acknowledgment in Section 6 that the evaluation is biased. However, the quantitative superiority claim is not yet established. The headline hard-level advantage is confounded with human-in-the-loop assistance, the evaluator is a single non-blinded author, no statistical significance or variance is reported, and the gold-standard dataset construction is undocumented. These issues are load-bearing because they directly support the paper's central claim that Text2Workflow outperforms a strong single-prompt baseline on complex requests.","major_comments":[{"comment":"The central claim that Text2Workflow outperforms baseline-gpt-4o-mini on hard-level requests (57.5% vs. 30%) is confounded by the human feedback loop. The ablation shows that without user aid (NUA) the hard-level accuracy is 26.3%, and with only general clarification (GC) it is 27.5%, both below the 30% baseline. Only with the human feedback loop alone (HFL, 32.5%) and especially the full system (57.5%) does the method surpass the baseline. Since the baseline receives no human feedback, the reported advantage cannot be attributed to the multi-layer prompt architecture alone; it may reflect the human corrections injected through the loop. The paper should add a control that gives the baseline the same human feedback (e.g., baseline-gpt-4o-mini with HFL), or otherwise explain the 25-point jump from HFL (32.5%) to the full system (57.5%) in terms of the specific interaction between GC and HFL, rather than attributing it to the pipeline generally.","section":"Section 5, Table 4"},{"comment":"The accuracy evaluation rests on the semantic rubric in Table 2 applied by a single author-evaluator, with no inter-annotator agreement, no confidence intervals, and no significance tests. The paper acknowledges this in Section 6, but the acknowledgment does not reduce the load-bearing role of the 57.5% vs. 30% comparison: the same author who supplies the human feedback in the full-pipeline runs also assigns the scores, so the evaluation is not blinded. The authors should either provide multiple independent evaluators with agreement statistics, or use an objective metric (e.g., schema validity plus executable-function correctness), and report variance or significance for the difficulty-stratified results.","section":"Section 5 and Section 6"},{"comment":"The construction of the Process2JSON gold-standard workflows is not documented. The paper does not state how the expected JSON outputs were produced—whether manually, by LLM, or by a hybrid process—nor whether the gold standard was validated for executability. If the expected JSONs were authored with the same schema and the same style conventions that the prompts enforce, the reported accuracies may measure consistency with the authors' format rather than executable correctness. The paper should describe the annotation process, the schema provenance, the qualifications of the annotator(s), and any checks performed on the gold-standard workflows.","section":"Section 4.1"},{"comment":"The claim that Text2Workflow 'achieves the highest performance across experiments' is based on an overall average that is dominated by the hard-level gap. At the easy level, baseline-gpt-4o-mini (92.5%) outperforms Text2Workflow (87.5%), and at the medium level they are close (70% versus 68.8%). The paper should present the difficulty-stratified accuracies with per-cell sample sizes (20 per difficulty level) and should temper the overall-performance claim accordingly, since the practical value of the system depends on where the differences are statistically reliable.","section":"Section 5, Figure 13"}],"minor_comments":[{"comment":"The text lists 'seven distinct layers of prompts' but then summarizes the pipeline into 'five main mechanisms'; the mapping between the seven layers and the five mechanisms is not made explicit and should be clarified.","section":"Section 3.2"},{"comment":"The abbreviation 'HLF' appears once ('the three ablation models (NUA, GC, HLF)') while the rest of the paper uses 'HFL' for the human feedback loop; the abbreviation should be standardized.","section":"Section 5"},{"comment":"The sentence 'Text2Workflow surpasses baseline-gpt-3.5-0125 by 48.7%, and equally outperforming baseline-gpt-4o-mini (by 27.5%)' is grammatically awkward; also, the percentage differences should be labeled as percentage points to avoid confusion.","section":"Section 5"},{"comment":"Section 6 states that Text2Workflow 'sometimes generates incorrect keys in the JSON structure,' while Section 5 states that 'there are no structural errors in the JSON format'; these claims should be harmonized or qualified.","section":"Section 6 vs. Section 5"},{"comment":"The reference 'OpenIA, 2024' is a typo for OpenAI, and the citation to the structured-outputs guide should be formatted consistently with the other OpenAI citations.","section":"References"},{"comment":"The text refers to 'figs. D.57 and D.58' for the erroneous output, but the figures are captioned as (1/3), (2/3), and (3/3); the figure numbering appears inconsistent and should be corrected.","section":"Appendix D"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an application-oriented prompt-engineering paper. The main uncertainty is whether the evaluation supports the central claim; the human-in-the-loop confound and single-evaluator scoring are the key issues. I see no citation-pattern concerns or scope mismatch; the paper fits an applied AI journal. A major revision that adds a controlled baseline with human feedback, multiple evaluators, and documentation of the dataset construction would substantively strengthen the contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result does not withstand scrutiny. The 57.5% vs 30% hard-level advantage comes almost entirely from the human feedback loop: with no user aid, Text2Workflow scores 26.3%, and with general clarification alone it scores 27.5%, both below baseline-gpt-4o-mini's 30%. Only when the human feedback loop is added does the full system reach 57.5%, a 25-point jump that the paper does not explain. And the human in that loop is also the single evaluator who scores the outputs. You do not need to be a statistician to see the confound.\n\nThat said, the paper has real substance. Process2JSON, a dataset of 60 categorized user requests with expected JSON workflows, is a useful new artifact. The Master/Expert prompt decomposition with a JSON schema is a reasonable engineering contribution, and the ablation of user input mechanisms is the right empirical question to ask. The authors also deserve credit for candidly acknowledging the single-evaluator limitation, the prompt-maintenance burden, and the security concerns of relying on a cloud GPT service. They include failure examples in Appendix D, which is more than many papers do.\n\nThe soft spots are exactly what the reader flagged. The gold-standard dataset provenance is unreported, so the expected JSONs may have been written with the same schema the prompts enforce, which would make the accuracy scores measure format consistency rather than executable correctness. There is no released code or data, no comparison to FlowMind or ProcessGPT, and no significance testing or variance reporting. There is also a small internal inconsistency: the text says baseline-gpt-4o-mini never beats Text2Workflow by more than 3% on easy-to-medium requests, but at easy level the gap is 5 percentage points (92.5 vs 87.5).\n\nThe central claim—that this pipeline outperforms a single-prompt baseline on complex requests—is not established by the evidence as presented. But the paper is honest, the method is described in enough detail to replicate, and the dataset alone justifies referee time. I would send it to peer review with a strong request for an independent, multi-evaluator, blinded evaluation and a clear separation of automated and human-assisted performance. The paper is not ready to be accepted as is, but it is a legitimate piece of engineering work worth engaging with.","headline":"A useful dataset and plausible pipeline, but the headline accuracy advantage is a human-in-the-loop artifact: without feedback the method loses to the baseline.","tokens_in":22312,"tokens_out":1738,"would_cite":false,"duration_ms":18477,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Text2Workflow generates business workflows from natural language, hitting 71.3% accuracy and beating single-prompt GPT-4o mini by 27.5% on complex requests.","keywords":["Natural Language Processing","Generative Artificial Intelligence","Large Language Model","Intelligent Automation","Workflow Automation","Text2Workflow","Process2JSON","JSON workflow generation"],"falsifier":"Score the same Text2Workflow outputs with multiple independent evaluators and with gold-standard workflows produced by a separate, non-LLM process, then compare inter-rater agreement and absolute scores on the hard subset; if agreement is low or the scores drop materially, the claimed 27.5-point advantage over the gpt-4o-mini baseline would not survive an objective re-measurement.","tokens_in":21212,"feed_emoji":"⚙️","tokens_out":5637,"duration_ms":44324,"temperature":0.7,"pith_summary":"This paper argues that a multi-stage prompt pipeline with user checkpoints can convert natural-language business requests into structured JSON workflows far more reliably than a single large prompt, especially when requests are complex. To test this, the authors built Text2Workflow, a seven-layer prompt system that first screens the request, builds a workflow skeleton, lets a human approve or edit it, then fills in step-level parameters. They also introduce Process2JSON, a 60-request dataset split into easy, medium, and hard levels with expected JSON outputs. On hard requests, Text2Workflow scores 57.5% accuracy against 30% for a single-prompt gpt-4o-mini baseline and 8.8% for gpt-3.5, while overall it reaches 71.3% accuracy. If these numbers hold, the main value is that non-experts could generate executable process blueprints by describing what they want, without writing code.","feed_headline":"Multi-step LLM prompts turn requests into workflows at 71.3% accuracy","feed_subtitle":"On complex requests it beats a single-prompt GPT-4o mini by 27.5 points, easing automation for non-coders.","key_machinery":"The machine that carries the argument is a predefined JSON workflow schema combined with a seven-layer prompt pipeline. The schema separates a general process header (id, name, description, parameters, steps, defaultStartStepId, context) from step-specific structures, with step types such as Decision, Loop, Calculation, DataExtraction, API, Exception, and Unknown. The pipeline uses a Master Prompt to build the workflow skeleton, Expert Prompts to fill parameters step by step, a Parameter Expert for API-like steps, a Questions Prompt to surface missing fields, and a Workflow Modification Prompt to apply user edits. The human feedback loop, inspired by FlowMind, summarizes the skeleton in plain language for user approval or revision before details are filled in. That loop is the main mechanism by which the method improves on hard requests.","core_discovery":"The central claim is that decomposing the workflow-generation task into a sequence of specialized prompts, with a human feedback loop at the skeleton stage, yields substantially more accurate JSON workflows than asking one LLM to do the whole job in a single long prompt. The paper reports that Text2Workflow achieves the highest overall accuracy (71.3%) across all experiments, and that on the hardest requests it surpasses baseline-gpt-3.5-0125 by 48.7 percentage points and baseline-gpt-4o-mini by 27.5 points. The ablation study attributes part of the gain to the human feedback loop, which alone improves accuracy by over 10%, while the logic-screening step adds only a modest 2% improvement.","pith_inferences":["The paper's reported accuracy is a measure of semantic similarity to gold-standard JSONs, not of whether a generated workflow runs correctly; extending the pipeline to actually execute the generated workflows would be the natural next test, and the authors' own appendix shows a non-executable Loop example.","The single-evaluator scoring rubric likely inflates agreement; an independent multi-evaluator study on the same dataset could shift the reported 71.3% figure, especially at the hard level where scoring is most subjective.","The token and time costs (roughly twice the input tokens, up to 163 seconds per run) imply that the technique is practical for offline or semi-automated process design but may need optimization before it fits interactive real-time use.","The approach's reliance on OpenAI's cloud API brings security and cost constraints; adapting the same prompt decomposition to an open-weight model would show whether the gain comes from the pipeline or from the particular LLM."],"forward_implications":["If the results replicate, decomposing LLM tasks into expert sub-prompts becomes a viable alternative to fine-tuning for structured-output generation tasks.","The public Process2JSON dataset gives other researchers a benchmark to compare natural-language-to-JSON workflow generation on the same easy, medium, and hard split.","A business user could describe a process in a few sentences and receive a JSON blueprint ready for visualization and, ultimately, execution by an automation engine with minimal manual intervention.","Because the gains concentrate on hard requests, organizations with convoluted, multi-branch processes stand to benefit most, while simple requests are better served by a single prompt for cost and speed."],"supporting_citations":[{"why":"Introduces FlowMind and the human-in-the-loop prompt recipe that the paper's feedback loop is inspired by; also cited as the basis for the observed accuracy gain.","marker":"Zeng et al., 2024"},{"why":"Supplies gpt-4o-mini, the model underlying Text2Workflow and the stronger single-prompt baseline.","marker":"OpenAI, 2024"},{"why":"Supplies gpt-3.5-0125, the weaker baseline and earlier model used to assess the method.","marker":"OpenAI, 2023"},{"why":"Introduces ProcessGPT, the prior LLM workflow-generation approach that Text2Workflow positions itself against.","marker":"Beheshti et al., 2023"},{"why":"Lists the four LLM limitations that frame the paper's evaluation design and the conceded weaknesses of the approach.","marker":"Ruan et al., 2023"},{"why":"Provides evidence that long-context windows do not imply accurate comprehension, motivating the modular prompt decomposition.","marker":"Hosseini et al., 2024"}],"fun_headline_variants":["Text2Workflow: LLM prompts hit 71.3% accuracy","Decomposed LLM prompts beat single-shot for workflows","Human-feedback LLM loop lifts workflow accuracy to 71.3%","LLM auto-workflows: 71.3% accuracy, beats GPT-4o mini by 27.5"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported accuracy numbers assume the expected JSON workflows in Process2JSON are correct and that the semantic scoring rubric, applied by one author-evaluator, is a trustworthy measure of workflow quality; if those gold standards were produced by the same kind of model or are biased toward the authors' own schema, the accuracy comparisons would measure consistency with that format rather than true executability.","fun_headline_variants_meta":{"raw":{"variants":["Text2Workflow: LLM prompts hit 71.3% accuracy","Decomposed LLM prompts beat single-shot for workflows","Human-feedback LLM loop lifts workflow accuracy to 71.3%","LLM auto-workflows: 71.3% accuracy, beats GPT-4o mini by 27.5"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1322,"prompt_tokens":869,"completion_tokens":453,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":365}},"tokens_in":485,"tokens_out":453,"duration_ms":4821,"temperature":1.0,"reasoning_tokens":365,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:23:35.816331+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Score the same Text2Workflow outputs with multiple independent evaluators and with gold-standard workflows produced by a separate, non-LLM process, then compare inter-rater agreement and absolute scores on the hard subset; if agreement is low or the scores drop materially, the claimed 27.5-point advantage over the gpt-4o-mini baseline would not survive an objective re-measurement.","supporting_citations":[{"cited_title":"title Models - openai api","cited_arxiv_id":null,"evidence_quote":"Supplies gpt-3.5-0125, the weaker baseline and earlier model used to assess the method."},{"cited_title":", author Chen, Y","cited_arxiv_id":null,"evidence_quote":"Lists the four LLM limitations that frame the paper's evaluation design and the conceded weaknesses of the approach."}],"review_version":1}