An experience report arguing that reliable AI evaluation needs structured volunteer cohorts, statistical error bars, and shared infrastructure rather than just software engineering.
Calculating optimal resampling for model evaluation
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Developing and Maintaining an Open-Source Repository of AI Evaluations: Challenges and Insights
An experience report arguing that reliable AI evaluation needs structured volunteer cohorts, statistical error bars, and shared infrastructure rather than just software engineering.