A new execution-based benchmark of 126 Ansible tasks shows open-source LLMs reach at most 12% pass@10, with most failures coming from state-tracking and module-knowledge errors.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Large Language Models for IT Automation Tasks: Are We There Yet?
A new execution-based benchmark of 126 Ansible tasks shows open-source LLMs reach at most 12% pass@10, with most failures coming from state-tracking and module-knowledge errors.