Token-level clinical text errors below 10% barely change classifier AUC in this study, but error rates at or above 10% cause visible drops, so data cleaning is mainly justified for heavily noisy datasets.
The Secondary Use of Electronic Health Records for Data Mining: Data Characteristics and Challenges
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Assessing the Impact of the Quality of Textual Data on Feature Representation and Machine Learning Models
Token-level clinical text errors below 10% barely change classifier AUC in this study, but error rates at or above 10% cause visible drops, so data cleaning is mainly justified for heavily noisy datasets.