ML-Dev-Bench introduces 30 ML workflow tasks and finds agent success drops sharply as tasks become more open-ended, with OpenHands-Sonnet best at 50%.
Mlagentbench: Evaluating language agents on machine learning experimentation, 2024
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.SE 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows
ML-Dev-Bench introduces 30 ML workflow tasks and finds agent success drops sharply as tasks become more open-ended, with OpenHands-Sonnet best at 50%.