LLMs achieve roughly 0.80 TypeSim similarity to human type annotations but far more mypy consistency errors than a coherent system should have, on a new 50-repo benchmark.
J., Bird, C., Barr, E
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
TypyBench: Evaluating LLM Type Inference for Untyped Python Repositories
LLMs achieve roughly 0.80 TypeSim similarity to human type annotations but far more mypy consistency errors than a coherent system should have, on a new 50-repo benchmark.