ML-Dev-Bench introduces 30 ML workflow tasks and finds agent success drops sharply as tasks become more open-ended, with OpenHands-Sonnet best at 50%.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows
ML-Dev-Bench introduces 30 ML workflow tasks and finds agent success drops sharply as tasks become more open-ended, with OpenHands-Sonnet best at 50%.