COINS, a test-case-based Rocq evaluation framework, shows frontier LLMs generate few candidate formal specifications on HumanEval, with rates between 1.22% and 28.05%.
International Conference on Machine Learning , pages=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
How Powerful are LLMs in Generating Formal Program Specifications?
COINS, a test-case-based Rocq evaluation framework, shows frontier LLMs generate few candidate formal specifications on HumanEval, with rates between 1.22% and 28.05%.