NLP papers commonly report annotator recruitment, expertise, and volume but frequently omit training, compensation, socio-demographics, adjudication, and agreement metrics, with reporting improving over time yet remaining uneven across tasks and venues.
arXiv preprint arXiv:2503.07575 , year =
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3verdicts
UNVERDICTED 3representative citing papers
SDGBiasBench reveals intrinsic SDG biases in VLMs driven by priors rather than evidence, and CADE mitigates them with up to 25% accuracy gains and 12-point MAE reductions.
The authors create a synthetic video auditing framework that detects statistically significant skin color biases in popular human action recognition models even when actions are identical.
citing papers explorer
-
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
NLP papers commonly report annotator recruitment, expertise, and volume but frequently omit training, compensation, socio-demographics, adjudication, and agreement metrics, with reporting improving over time yet remaining uneven across tasks and venues.
-
SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals
SDGBiasBench reveals intrinsic SDG biases in VLMs driven by priors rather than evidence, and CADE mitigates them with up to 25% accuracy gains and 12-point MAE reductions.
-
Identifying Ethical Biases in Action Recognition Models
The authors create a synthetic video auditing framework that detects statistically significant skin color biases in popular human action recognition models even when actions are identical.