REVIEW 3 major objections 4 minor 38 references
Towards Automatic Detection of Misinformation in Online Medical Videos
T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A multimodal SVM can identify misinformative prostate cancer YouTube videos with 74% accuracy.
desk verdict A genuinely new dataset with a headline accuracy that is likely inflated by the query terms used to collect misinformative videos; deserves peer review with major revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is an early-fusion linear SVM over three feature families: six YouTube engagement statistics (views per day, comments, thumbs up and down, duration, and category), 6,988 linguistic features computed from automatically transcribed and punctuation-restored transcripts, and 384 Emo_IS09 acoustic features. Labels come from expert urologist ratings on a five-point misinformation scale collapsed into a binary trustworthy/misinformative split. The ablation design—each feature family alone, then combined—is what lets the paper attribute the final accuracy to complementarity across modalities.
What would settle it
Collect a fresh, unfiltered sample of prostate cancer videos from YouTube, including animations, voice-overs, non-English content, and videos longer than 30 minutes, annotate it with the same expert protocol, and run the published feature pipeline with the same SVM settings. If accuracy on that sample drops toward the 52.8% majority baseline, the original 74% result is an artifact of the curated search and filtering procedure.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that misinformation in medical videos is detectable from surface signals: a classifier that fuses all extracted linguistic features, six YouTube engagement attributes, and the Emo_IS09 acoustic feature set reaches 74.41% accuracy, with 76.51% precision and 73.15% recall on the misinformative class. The integrated model outperforms every single-modality model, and ngrams are the strongest individual linguistic feature, followed by syntax features and LIWC categories. The paper also establishes through ablation that the acoustic and engagement channels add complementary signal on top of language rather than simply duplicating it.
Load-bearing premise
The dataset was built by searching for obviously misinformative phrases like 'miracle cure' and by excluding animated videos, artificial voices, non-English speech, more than two speakers, and videos over 30 minutes, so the learned boundary may not hold for the full population of prostate cancer videos on YouTube.
Editorial extensions
If this is right
- A first-pass automated screener could rank YouTube prostate cancer videos by likelihood of misinformation and send only flagged or uncertain videos to medical experts, lowering the cost of manual review.
- Viewer engagement alone is not enough: it reaches 96% precision but only 21% recall on misinformative videos, so a deployed detector needs the linguistic and acoustic channels to avoid missing most bad content.
- The combined model's improvement over language-only features shows that tone of voice and audience reaction carry independent misinformation cues that transcripts alone miss.
- The annotated 250-video dataset, with fine-grained scores from 1 to 5, can support future work on graded misinformation severity instead of only binary flags.
Reading between the lines
- A natural next test is cross-topic transfer: because many ngram features are prostate-cancer-specific, the same pipeline may need topic-adapted features to detect misinformation about other diseases.
- The time-sensitive engagement features mean a deployed model trained on April 2019 statistics would age; a practical detector would need periodic retraining or should drop engagement features.
- If acoustic cues truly help, a targeted experiment could measure which acoustic dimensions—pitch range, loudness, speaking rate—separate anecdotal 'miracle cure' testimonials from clinical explanations.
- The declared filtering criteria imply the strong accuracy may be partly a curated-sample effect; the strongest testable extension is running the same pipeline on unfiltered YouTube search results.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses automatic detection of misinformation in YouTube videos about prostate cancer. The authors introduce a new dataset of 250 videos manually annotated on a 5-point misinformation scale and binarized into trustworthy and misinformative classes. They extract viewer engagement, linguistic (ngrams, LIWC, syntax, readability, lexical richness), and acoustic (openEAR) feature sets, and train linear SVM classifiers with five-fold cross-validation. The best combined model is reported to reach 74.41% accuracy, with 76.51% precision and 73.15% recall for the misinformative class, against a 52.8% majority-class baseline. The paper also reports ablation results showing that individual feature sets vary widely and that the combined model outperforms single-modality models.
Significance. If the reported accuracy transfers beyond the specific retrieval protocol, this is a useful early step toward automated screening of health misinformation in online video, a largely underexplored problem. The main strengths are the new expert-annotated dataset, the explicit majority-class baseline, and the systematic ablation across multimodal feature groups. The central risk is external validity: because misinformative videos were retrieved using trigger phrases such as 'miracle cure' and 'natural remedies' while trustworthy videos were retrieved using neutral medical queries, the high ngram-based accuracy may reflect query-specific vocabulary rather than generalizable misinformation cues. The paper's contribution is therefore conditional on additional robustness evidence showing that the model generalizes beyond its search-term vocabulary.
major comments (3)
- [Section 3 (Data Collection) and Table 3]
- [Table 3 and Section 4 (Classification Results)]
- [Section 3 (Data Annotation)]
minor comments (4)
- [Abstract and Section 1]
- [Table 3]
- [Section 4 (Viewer Engagement Features)]
- [Section 3 (Privacy Considerations)]
Circularity Check
No circularity: the reported accuracy is measured on held-out cross-validation folds, and the query-based dataset bias is an external-validity concern, not an in-loop derivation.
full rationale
The paper's central claim is that an SVM trained on engagement, linguistic, and acoustic features detects misinformative prostate cancer videos at 74.41% accuracy. This result is obtained through five-fold cross-validation (Section 4, 'Classification Results'), so the reported numbers are out-of-fold evaluations rather than fitted values renamed as predictions. The features are extracted from transcripts, audio, and YouTube metadata; none of them is defined in terms of the misinformation label, and the paper does not fit a parameter to the test folds and then call it a prediction. The dataset construction used trigger queries such as 'prostate cancer miracle cure' to oversample misinformative content, which may inflate separability or let ngram features memorize retrieval vocabulary, but that is a sampling and external-validity limitation, not a circularity step under the specified taxonomy. Section 5 explicitly acknowledges related limitations, including the exclusion of animated or voice-over videos and the overrepresentation of laypeople, and the paper does not claim the model transfers beyond the collected distribution. The only author-overlapping citation, Loeb et al. [23], is an externally published observational study used to motivate engagement features and the existence of misinformation; it is not the source of the classifier's measured performance and does not make the derivation circular. No uniqueness theorem, self-defined quantity, or ansatz is imported from prior work to force the reported outcome. Therefore no circular step is present.
Assumptions & free parameters
free parameters (3)
- N-gram and syntax frequency threshold =
10
- SVM regularization parameter C =
not reported
- Acoustic feature set selection =
Emo_IS09
assumptions (5)
- domain assumption The 5-point misinformation scale, anchored to DISCERN and expert urologist judgment, is a valid measure of misinformation.
- domain assumption Videos excluded by the collection criteria (animations, artificial voices, more than two speakers, non-English, longer than 30 minutes) are not needed for the central claim.
- domain assumption Automatic speech recognition and punctuation restoration produce transcripts of sufficient quality for linguistic feature extraction.
- ad hoc to paper The feature sets chosen (ngrams, LIWC, syntax, readability, lexical richness, openEAR acoustics, viewer engagement) capture content and style cues that vary with misinformation.
- standard math Cross-validation with a majority class baseline is an appropriate evaluation protocol for the claim of automatic detection.
Cite this review
Pith. "Pith review of Towards Automatic Detection of Misinformation in Online Medical Videos." pith.science (2026). https://pith.science/paper/NFITRGJD
@misc{pith2026190901543,
author = {Pith},
title = {Pith review of: Towards Automatic Detection of Misinformation in Online Medical Videos},
year = {2026},
howpublished = {\url{https://pith.science/paper/NFITRGJD}},
note = {Machine review of arXiv:1909.01543}
}
read the original abstract
Recent years have witnessed a significant increase in the online sharing of medical information, with videos representing a large fraction of such online sources. Previous studies have however shown that more than half of the health-related videos on platforms such as YouTube contain misleading information and biases. Hence, it is crucial to build computational tools that can help evaluate the quality of these videos so that users can obtain accurate information to help inform their decisions. In this study, we focus on the automatic detection of misinformation in YouTube videos. We select prostate cancer videos as our entry point to tackle this problem. The contribution of this paper is twofold. First, we introduce a new dataset consisting of 250 videos related to prostate cancer manually annotated for misinformation. Second, we explore the use of linguistic, acoustic, and user engagement features for the development of classification models to identify misinformation. Using a series of ablation experiments, we show that we can build automatic models with accuracies of up to 74%, corresponding to a 76.5% precision and 73.2% recall for misinformative instances.
Figures
Reference graph
Works this paper leans on
-
[1]
Corey H Basch, Anthony Menafro, Jennifer Mongiovi, Grace Clarke Hillyer, and Charles E Basch. 2017. A content analysis of YouTubeâĎć videos related to prostate cancer. American journal of men’s health 11, 1 (2017), 154–157
work page 2017
-
[2]
Peter C Black and David F Penson. 2006. Prostate cancer on the Internet-information or misinformation? The Journal of urology 175, 5 (2006), 1836–1842
work page 2006
-
[3]
Christina Boididou, Symeon Papadopoulos, Lazaros Apostolidis, and Yiannis Kompatsiaris. 2017. Learning to detect misleading content on twitter. In Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval. ACM, 278–286
work page 2017
-
[4]
JT Cassidy, E Fitzgerald, ES Cassidy, M Cleary, DP Byrne, BM Devitt, and JF Baker. 2018. YouTube provides poor information regarding anterior cruciate ligament injury and reconstruction. Knee Surgery, Sports Traumatology, Arthroscopy 26, 3 (2018), 840–845
work page 2018
-
[5]
J. T. Cassidy, E. Fitzgerald, E. S. Cassidy, M. Cleary, D. P. Byrne, B. M. Devitt, and J. F. Baker. 2018. YouTube provides poor information regarding anterior cruciate ligament injury and reconstruction. Knee Surgery, Sports Traumatology, Arthroscopy 26, 3 (01 Mar 2018), 840–845. https://doi.org/10.1007/s00167-017-4514-x
-
[6]
Deborah Charnock, Sasha Shepperd, Gill Needham, and Robert Gann
-
[7]
Danqi Chen and Christopher Manning. 2014. A fast and accurate de- pendency parser using neural networks. In Proceedings of the 2014 con- ference on empirical methods in natural language processing (EMNLP) . 740–750
work page 2014
- [8]
Show all 38 references
-
[9]
Brandy Drozd, Emily Couvillon, and Andrea Suarez. 2018. Medical YouTube videos and methods of evaluation: literature review. JMIR medical education 4, 1 (2018), e3
2018
-
[10]
Florian Eyben, Martin Wöllmer, and Björn Schuller. 2009. OpenEAR- introducing the Munich open-source emotion and affect recognition toolkit. In 2009 3rd international conference on affective computing and intelligent interaction and workshops . IEEE, 1–6
2009
-
[11]
Rong-En Fan, Kai-Wei Chang, Cho-Jui Hsieh, Xiang-Rui Wang, and Chih-Jen Lin. 2008. LIBLINEAR: A Library for Large Linear Clas- sification. J. Mach. Learn. Res. 9 (June 2008), 1871–1874. http: //dl.acm.org/citation.cfm?id=1390681.1442794
2008
-
[12]
Song Feng, Ritwik Banerjee, and Yejin Choi. 2012. Syntactic stylometry for deception detection. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Short Papers-Volume 2 . Association for Computational Linguistics, 171–175
2012
-
[13]
Susannah Fox. 2011. The social life of health information 2011 . Pew Internet & American Life Project Washington, DC
2011
-
[14]
Amira Ghenai and Yelena Mejova. 2018. Fake cures: user-centric modeling of health misinformation in social media. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 58
2018
-
[15]
Shan Jiang and Christo Wilson. 2018. Linguistic Signals Under Mis- information and Fact-Checking: Evidence from User Comments on Social Media. Proc. ACM Hum.-Comput. Interact. 2, CSCW, Article 82 (Nov. 2018), 23 pages. https://doi.org/10.1145/3274351 Towards Automatic Detectio...
2018 doi
-
[16]
Philipp Koehn. 2005. Europarl: A parallel corpus for statistical machine translation. In MT summit, Vol. 5. 79–86
2005
-
[17]
Gangeshwar Krishnamurthy, Navonil Majumder, Soujanya Poria, and Erik Cambria. 2018. A deep learning approach for multimodal decep- tion detection. arXiv preprint arXiv:1803.00344 (2018)
2018 arXiv
-
[18]
Lau, Elia Gabarron, Luis Fernandez-Luque, and Manuel Ar- mayones
Annie Y.S. Lau, Elia Gabarron, Luis Fernandez-Luque, and Manuel Ar- mayones. 2012. Social Media in Health âĂŤ What are the Safety Con- cerns for Health Consumers? Health Information Management Jour- nal 41, 2 (2012), 30–35. https://doi.org/10.1177/183335831204100204 arXiv:http...
2012 doi
-
[19]
Amanda Y Leong, Ravina Sanghera, Jaspreet Jhajj, Nandini Desai, Bikramjit Singh Jammu, and Mark J Makowsky. 2018. Is YouTube useful as a source of health information for adults with type 2 diabetes? A South Asian perspective. Canadian journal of diabetes 42, 4 (2018), 395–403
2018
-
[20]
Xiao Liu, Bin Zhang, Anjana Susarla, and Rema Padman. 2018. YouTube for Patient Education: A Deep Learning Approach for Understand- ing Medical Knowledge from User-Generated Videos. arXiv preprint arXiv:1807.03179 (2018)
2018 arXiv
-
[21]
Alto S Lo, Michael J Esser, and Kevin E Gordon. 2010. YouTube: a gauge of public perception and awareness surrounding epilepsy. Epilepsy & Behavior 17, 4 (2010), 541–545
2010
-
[22]
Stacy Loeb, Matthew S Katz, Aisha Langford, Nataliya Byrne, and Shannon Ciprut. 2018. Prostate cancer and social media. Nature Reviews Urology (2018), 1
2018
-
[23]
Macaluso, Stefan W
Stacy Loeb, Shomik Sengupta, Mohit Butaney, Joseph N. Macaluso, Stefan W. Czarniecki, Rebecca Robbins, R. Scott Braithwaite, Lingshan Gao, Nataliya Byrne, Dawn Walter, and Aisha Langford. 2019. Dis- semination of Misinformative and Biased Information about Prostate Cancer on Y...
2019 doi
-
[24]
Xiaofei Lu. 2012. The relationship of lexical richness to the quality of ESL learnersâĂŹ oral narratives. The Modern Language Journal 96, 2 (2012), 190–208
2012
-
[25]
Kapil Chalil Madathil, A Joy Rivera-Rodriguez, Joel S Greenstein, and Anand K Gramopadhye. 2015. Healthcare information on YouTube: a systematic review. Health informatics journal 21, 3 (2015), 173–194
2015
-
[26]
Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J
Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky. 2014. The Stanford CoreNLP Natural Language Processing Toolkit. InAssociation for Computational Linguistics (ACL) System Demonstrations. 55–60. http://www.aclweb. org/antho...
2014
-
[27]
Pedregosa, G
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay
-
[28]
Verónica Pérez-Rosas, Bennett Kleinberg, Alexandra Lefevre, and Rada Mihalcea. 2017. Automatic detection of fake news. arXiv preprint arXiv:1708.07104 (2017)
2017 arXiv
-
[29]
Jiangwei Qi, T Trang, J Doong, S Kang, and Anna L Chien. 2016. Mis- information is prevalent in psoriasis-related YouTube videos. Derma- tology online journal 22, 11 (2016)
2016
-
[30]
Kai Shu, Amy Sliva, Suhang Wang, Jiliang Tang, and Huan Liu. 2017. Fake news detection on social media: A data mining perspective. ACM SIGKDD Explorations Newsletter 19, 1 (2017), 22–36
2017
-
[31]
RP Smith, P Devine, H Jones, A DeNittis, R Whittington, and JM Metz. 2003. Internet use by patients with prostate cancer undergoing radiotherapy. Urology 62, 2 (2003), 273–277
2003
-
[32]
American Cancer Society. 2019. Cancer Facts and Figures 2019. (2019)
2019
-
[33]
Peter L Steinberg, Shaun Wason, Joshua M Stern, Levi Deters, Brian Kowal, and John Seigne. 2010. YouTube as source of prostate cancer information. Urology 75, 3 (2010), 619–622
2010
-
[34]
Shabbir Syed-Abdul, Luis Fernandez-Luque, Wen-Shan Jian, Yu-Chuan Li, Steven Crain, Min-Huei Hsu, Yao-Chin Wang, Dorjsuren Khan- dregzen, Enkhzaya Chuluunbaatar, Phung Anh Nguyen, et al. 2013. Misleading health-related information promoted through video-based social media: ano...
2013
-
[35]
Ottokar Tilk and Tanel Alumäe. 2016. Bidirectional Recurrent Neural Network with Attention Mechanism for Punctuation Restoration. In Interspeech 2016
2016
-
[36]
Kristina Toutanova, Dan Klein, Christopher D Manning, and Yoram Singer. 2003. Feature-rich part-of-speech tagging with a cyclic de- pendency network. In Proceedings of the 2003 conference of the North American chapter of the association for computational linguistics on human l...
2003
-
[1999]
Journal of Epi- demiology & Community Health 53, 2 (1999), 105–111
DISCERN: an instrument for judging the quality of written consumer health information on treatment choices. Journal of Epi- demiology & Community Health 53, 2 (1999), 105–111
1999
-
[2011]
Journal of Machine Learning Research 12 (2011), 2825–2830
Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research 12 (2011), 2825–2830
2011
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.