A real-world industrial benchmark and agentic baseline show current LLMs often fail to produce correct, stable, deployable visual workflows from natural language, with only modest resolve-rate gains.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language
A real-world industrial benchmark and agentic baseline show current LLMs often fail to produce correct, stable, deployable visual workflows from natural language, with only modest resolve-rate gains.