Pith. sign in

REVIEW 3 major objections 4 minor 38 references

Towards Automatic Detection of Misinformation in Online Medical Videos

T0 review · 3 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A multimodal SVM can identify misinformative prostate cancer YouTube videos with 74% accuracy.

desk verdict A genuinely new dataset with a headline accuracy that is likely inflated by the query terms used to collect misinformative videos; deserves peer review with major revision. read the letter →

arxiv 1909.01543 v1 pith:NFITRGJD submitted 2019-09-04 cs.LG stat.ML

classification cs.LGstat.ML
keywords misinformationdetectionprostatecancerYouTubemultimodalclassificationsupportvectormachinehealthacousticfeatureslinguistic
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a computer can tell when a YouTube video about prostate cancer is spreading misinformation, without a doctor watching it. It introduces a new annotated dataset of 250 prostate cancer videos—132 trustworthy and 118 misinformative—and shows that a linear SVM combining viewer engagement statistics, transcript language features, and acoustic-prosodic features identifies misinformative videos with 74.4% accuracy, 76.5% precision, and 73.2% recall. That is a clear improvement over the 52.8% majority-class baseline, cutting the error rate roughly in half. If the result holds, it means cheap automated screening could flag suspicious health videos for expert review at scale.

What carries the argument

The load-bearing mechanism is an early-fusion linear SVM over three feature families: six YouTube engagement statistics (views per day, comments, thumbs up and down, duration, and category), 6,988 linguistic features computed from automatically transcribed and punctuation-restored transcripts, and 384 Emo_IS09 acoustic features. Labels come from expert urologist ratings on a five-point misinformation scale collapsed into a binary trustworthy/misinformative split. The ablation design—each feature family alone, then combined—is what lets the paper attribute the final accuracy to complementarity across modalities.

What would settle it

Collect a fresh, unfiltered sample of prostate cancer videos from YouTube, including animations, voice-overs, non-English content, and videos longer than 30 minutes, annotate it with the same expert protocol, and run the published feature pipeline with the same SVM settings. If accuracy on that sample drops toward the 52.8% majority baseline, the original 74% result is an artifact of the curated search and filtering procedure.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that misinformation in medical videos is detectable from surface signals: a classifier that fuses all extracted linguistic features, six YouTube engagement attributes, and the Emo_IS09 acoustic feature set reaches 74.41% accuracy, with 76.51% precision and 73.15% recall on the misinformative class. The integrated model outperforms every single-modality model, and ngrams are the strongest individual linguistic feature, followed by syntax features and LIWC categories. The paper also establishes through ablation that the acoustic and engagement channels add complementary signal on top of language rather than simply duplicating it.

Load-bearing premise

The dataset was built by searching for obviously misinformative phrases like 'miracle cure' and by excluding animated videos, artificial voices, non-English speech, more than two speakers, and videos over 30 minutes, so the learned boundary may not hold for the full population of prostate cancer videos on YouTube.

Editorial extensions

If this is right

  • A first-pass automated screener could rank YouTube prostate cancer videos by likelihood of misinformation and send only flagged or uncertain videos to medical experts, lowering the cost of manual review.
  • Viewer engagement alone is not enough: it reaches 96% precision but only 21% recall on misinformative videos, so a deployed detector needs the linguistic and acoustic channels to avoid missing most bad content.
  • The combined model's improvement over language-only features shows that tone of voice and audience reaction carry independent misinformation cues that transcripts alone miss.
  • The annotated 250-video dataset, with fine-grained scores from 1 to 5, can support future work on graded misinformation severity instead of only binary flags.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next test is cross-topic transfer: because many ngram features are prostate-cancer-specific, the same pipeline may need topic-adapted features to detect misinformation about other diseases.
  • The time-sensitive engagement features mean a deployed model trained on April 2019 statistics would age; a practical detector would need periodic retraining or should drop engagement features.
  • If acoustic cues truly help, a targeted experiment could measure which acoustic dimensions—pitch range, loudness, speaking rate—separate anecdotal 'miracle cure' testimonials from clinical explanations.
  • The declared filtering criteria imply the strong accuracy may be partly a curated-sample effect; the strongest testable extension is running the same pipeline on unfiltered YouTube search results.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper addresses automatic detection of misinformation in YouTube videos about prostate cancer. The authors introduce a new dataset of 250 videos manually annotated on a 5-point misinformation scale and binarized into trustworthy and misinformative classes. They extract viewer engagement, linguistic (ngrams, LIWC, syntax, readability, lexical richness), and acoustic (openEAR) feature sets, and train linear SVM classifiers with five-fold cross-validation. The best combined model is reported to reach 74.41% accuracy, with 76.51% precision and 73.15% recall for the misinformative class, against a 52.8% majority-class baseline. The paper also reports ablation results showing that individual feature sets vary widely and that the combined model outperforms single-modality models.

Significance. If the reported accuracy transfers beyond the specific retrieval protocol, this is a useful early step toward automated screening of health misinformation in online video, a largely underexplored problem. The main strengths are the new expert-annotated dataset, the explicit majority-class baseline, and the systematic ablation across multimodal feature groups. The central risk is external validity: because misinformative videos were retrieved using trigger phrases such as 'miracle cure' and 'natural remedies' while trustworthy videos were retrieved using neutral medical queries, the high ngram-based accuracy may reflect query-specific vocabulary rather than generalizable misinformation cues. The paper's contribution is therefore conditional on additional robustness evidence showing that the model generalizes beyond its search-term vocabulary.

major comments (3)
  1. [Section 3 (Data Collection) and Table 3]
  2. [Table 3 and Section 4 (Classification Results)]
  3. [Section 3 (Data Annotation)]
minor comments (4)
  1. [Abstract and Section 1]
  2. [Table 3]
  3. [Section 4 (Viewer Engagement Features)]
  4. [Section 3 (Privacy Considerations)]

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the reported accuracy is measured on held-out cross-validation folds, and the query-based dataset bias is an external-validity concern, not an in-loop derivation.

full rationale

The paper's central claim is that an SVM trained on engagement, linguistic, and acoustic features detects misinformative prostate cancer videos at 74.41% accuracy. This result is obtained through five-fold cross-validation (Section 4, 'Classification Results'), so the reported numbers are out-of-fold evaluations rather than fitted values renamed as predictions. The features are extracted from transcripts, audio, and YouTube metadata; none of them is defined in terms of the misinformation label, and the paper does not fit a parameter to the test folds and then call it a prediction. The dataset construction used trigger queries such as 'prostate cancer miracle cure' to oversample misinformative content, which may inflate separability or let ngram features memorize retrieval vocabulary, but that is a sampling and external-validity limitation, not a circularity step under the specified taxonomy. Section 5 explicitly acknowledges related limitations, including the exclusion of animated or voice-over videos and the overrepresentation of laypeople, and the paper does not claim the model transfers beyond the collected distribution. The only author-overlapping citation, Loeb et al. [23], is an externally published observational study used to motivate engagement features and the existence of misinformation; it is not the source of the classifier's measured performance and does not make the derivation circular. No uniqueness theorem, self-defined quantity, or ansatz is imported from prior work to force the reported outcome. Therefore no circular step is present.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central claim rests on a small, deliberately sampled expert-annotated dataset. The main free parameters are hyperparameters of the classifier and feature extraction (ngram threshold 10, acoustic feature set selected by performance). The axioms are domain assumptions about the validity of expert annotation, the sufficiency of ASR transcript quality, and the representativeness of the excluded-video subset.

free parameters (3)
  • N-gram and syntax frequency threshold = 10
    Frequency threshold for unigrams/bigrams and CFG rules was set to 10, 'obtained experimentally on a development set' (Section 4, Linguistic Features). This threshold is tuned on data and may lead to overfitting if the development set is small.
  • SVM regularization parameter C = not reported
    The paper uses LinearSVC from scikit-learn but does not report the C value or any hyperparameter search; if left at the library default, this is a hand-selected hyperparameter.
  • Acoustic feature set selection = Emo_IS09
    The paper selects Emo_IS09 over Emobase and Emo_large because it performed best in the initial experiments; this is a model selection choice based on the evaluation set.
assumptions (5)
  • domain assumption The 5-point misinformation scale, anchored to DISCERN and expert urologist judgment, is a valid measure of misinformation.
    The dataset labels all model training and evaluation. The paper relies on expert annotation as ground truth, without a formal validation of the rating rubric beyond within-one agreement.
  • domain assumption Videos excluded by the collection criteria (animations, artificial voices, more than two speakers, non-English, longer than 30 minutes) are not needed for the central claim.
    The paper deliberately restricts the dataset, which limits the scope of the claim to a subset of YouTube videos.
  • domain assumption Automatic speech recognition and punctuation restoration produce transcripts of sufficient quality for linguistic feature extraction.
    The pipeline uses YouTube ASR, Google Cloud Speech-to-Text, and a punctuation restoration model; the paper acknowledges 'some degree of error' but asserts it is small enough not to affect downstream analysis.
  • ad hoc to paper The feature sets chosen (ngrams, LIWC, syntax, readability, lexical richness, openEAR acoustics, viewer engagement) capture content and style cues that vary with misinformation.
    This is a design choice grounded in prior deception detection work, but the paper does not prove that these features are uniquely informative; it relies on empirical improvement over baseline.
  • standard math Cross-validation with a majority class baseline is an appropriate evaluation protocol for the claim of automatic detection.
    Standard practice, though the paper does not report variance across folds.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards Automatic Detection of Misinformation in Online Medical Videos." pith.science (2026). https://pith.science/paper/NFITRGJD

@misc{pith2026190901543,
  author       = {Pith},
  title        = {Pith review of: Towards Automatic Detection of Misinformation in Online Medical Videos},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NFITRGJD}},
  note         = {Machine review of arXiv:1909.01543}
}
read the original abstract

Recent years have witnessed a significant increase in the online sharing of medical information, with videos representing a large fraction of such online sources. Previous studies have however shown that more than half of the health-related videos on platforms such as YouTube contain misleading information and biases. Hence, it is crucial to build computational tools that can help evaluate the quality of these videos so that users can obtain accurate information to help inform their decisions. In this study, we focus on the automatic detection of misinformation in YouTube videos. We select prostate cancer videos as our entry point to tackle this problem. The contribution of this paper is twofold. First, we introduce a new dataset consisting of 250 videos related to prostate cancer manually annotated for misinformation. Second, we explore the use of linguistic, acoustic, and user engagement features for the development of classification models to identify misinformation. Using a series of ablation experiments, we show that we can build automatic models with accuracies of up to 74%, corresponding to a 76.5% precision and 73.2% recall for misinformative instances.

Figures

Figures reproduced from arXiv: 1909.01543 by the authors.

Figure 1
Figure 1. YouTube viewer engagement features for misinformative and trustworthy videos [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

38 extracted references · 34 canonical work pages

  1. [1]

    Corey H Basch, Anthony Menafro, Jennifer Mongiovi, Grace Clarke Hillyer, and Charles E Basch. 2017. A content analysis of YouTubeâĎć videos related to prostate cancer. American journal of men’s health 11, 1 (2017), 154–157

  2. [2]

    Peter C Black and David F Penson. 2006. Prostate cancer on the Internet-information or misinformation? The Journal of urology 175, 5 (2006), 1836–1842

  3. [3]

    Christina Boididou, Symeon Papadopoulos, Lazaros Apostolidis, and Yiannis Kompatsiaris. 2017. Learning to detect misleading content on twitter. In Proceedings of the 2017 ACM on International Conference on Multimedia Retrieval. ACM, 278–286

  4. [4]

    JT Cassidy, E Fitzgerald, ES Cassidy, M Cleary, DP Byrne, BM Devitt, and JF Baker. 2018. YouTube provides poor information regarding anterior cruciate ligament injury and reconstruction. Knee Surgery, Sports Traumatology, Arthroscopy 26, 3 (2018), 840–845

  5. [5]

    J. T. Cassidy, E. Fitzgerald, E. S. Cassidy, M. Cleary, D. P. Byrne, B. M. Devitt, and J. F. Baker. 2018. YouTube provides poor information regarding anterior cruciate ligament injury and reconstruction. Knee Surgery, Sports Traumatology, Arthroscopy 26, 3 (01 Mar 2018), 840–845. https://doi.org/10.1007/s00167-017-4514-x

  6. [6]

    Deborah Charnock, Sasha Shepperd, Gill Needham, and Robert Gann

  7. [7]

    Danqi Chen and Christopher Manning. 2014. A fast and accurate de- pendency parser using neural networks. In Proceedings of the 2014 con- ference on empirical methods in natural language processing (EMNLP) . 740–750

  8. [8]

    Wen-Ying Sylvia Chou, April Oh, and William M. P. Klein. 2018. Ad- dressing Health-Related Misinformation on Social Media. JAMA 320, 23 (12 2018), 2417–2418. https://doi.org/10.1001/jama.2018.16865

Show all 38 references
  1. [9]

    Brandy Drozd, Emily Couvillon, and Andrea Suarez. 2018. Medical YouTube videos and methods of evaluation: literature review. JMIR medical education 4, 1 (2018), e3

  2. [10]

    Florian Eyben, Martin Wöllmer, and Björn Schuller. 2009. OpenEAR- introducing the Munich open-source emotion and affect recognition toolkit. In 2009 3rd international conference on affective computing and intelligent interaction and workshops . IEEE, 1–6

  3. [11]

    Rong-En Fan, Kai-Wei Chang, Cho-Jui Hsieh, Xiang-Rui Wang, and Chih-Jen Lin. 2008. LIBLINEAR: A Library for Large Linear Clas- sification. J. Mach. Learn. Res. 9 (June 2008), 1871–1874. http: //dl.acm.org/citation.cfm?id=1390681.1442794

  4. [12]

    Song Feng, Ritwik Banerjee, and Yejin Choi. 2012. Syntactic stylometry for deception detection. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Short Papers-Volume 2 . Association for Computational Linguistics, 171–175

  5. [13]

    Susannah Fox. 2011. The social life of health information 2011 . Pew Internet & American Life Project Washington, DC

  6. [14]

    Amira Ghenai and Yelena Mejova. 2018. Fake cures: user-centric modeling of health misinformation in social media. Proceedings of the ACM on Human-Computer Interaction 2, CSCW (2018), 58

  7. [15]

    Shan Jiang and Christo Wilson. 2018. Linguistic Signals Under Mis- information and Fact-Checking: Evidence from User Comments on Social Media. Proc. ACM Hum.-Comput. Interact. 2, CSCW, Article 82 (Nov. 2018), 23 pages. https://doi.org/10.1145/3274351 Towards Automatic Detectio...

  8. [16]

    Philipp Koehn. 2005. Europarl: A parallel corpus for statistical machine translation. In MT summit, Vol. 5. 79–86

  9. [17]

    Gangeshwar Krishnamurthy, Navonil Majumder, Soujanya Poria, and Erik Cambria. 2018. A deep learning approach for multimodal decep- tion detection. arXiv preprint arXiv:1803.00344 (2018)

  10. [18]

    Lau, Elia Gabarron, Luis Fernandez-Luque, and Manuel Ar- mayones

    Annie Y.S. Lau, Elia Gabarron, Luis Fernandez-Luque, and Manuel Ar- mayones. 2012. Social Media in Health âĂŤ What are the Safety Con- cerns for Health Consumers? Health Information Management Jour- nal 41, 2 (2012), 30–35. https://doi.org/10.1177/183335831204100204 arXiv:http...

  11. [19]

    Amanda Y Leong, Ravina Sanghera, Jaspreet Jhajj, Nandini Desai, Bikramjit Singh Jammu, and Mark J Makowsky. 2018. Is YouTube useful as a source of health information for adults with type 2 diabetes? A South Asian perspective. Canadian journal of diabetes 42, 4 (2018), 395–403

  12. [20]

    Xiao Liu, Bin Zhang, Anjana Susarla, and Rema Padman. 2018. YouTube for Patient Education: A Deep Learning Approach for Understand- ing Medical Knowledge from User-Generated Videos. arXiv preprint arXiv:1807.03179 (2018)

  13. [21]

    Alto S Lo, Michael J Esser, and Kevin E Gordon. 2010. YouTube: a gauge of public perception and awareness surrounding epilepsy. Epilepsy & Behavior 17, 4 (2010), 541–545

  14. [22]

    Stacy Loeb, Matthew S Katz, Aisha Langford, Nataliya Byrne, and Shannon Ciprut. 2018. Prostate cancer and social media. Nature Reviews Urology (2018), 1

  15. [23]

    Macaluso, Stefan W

    Stacy Loeb, Shomik Sengupta, Mohit Butaney, Joseph N. Macaluso, Stefan W. Czarniecki, Rebecca Robbins, R. Scott Braithwaite, Lingshan Gao, Nataliya Byrne, Dawn Walter, and Aisha Langford. 2019. Dis- semination of Misinformative and Biased Information about Prostate Cancer on Y...

  16. [24]

    Xiaofei Lu. 2012. The relationship of lexical richness to the quality of ESL learnersâĂŹ oral narratives. The Modern Language Journal 96, 2 (2012), 190–208

  17. [25]

    Kapil Chalil Madathil, A Joy Rivera-Rodriguez, Joel S Greenstein, and Anand K Gramopadhye. 2015. Healthcare information on YouTube: a systematic review. Health informatics journal 21, 3 (2015), 173–194

  18. [26]

    Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J

    Christopher D. Manning, Mihai Surdeanu, John Bauer, Jenny Finkel, Steven J. Bethard, and David McClosky. 2014. The Stanford CoreNLP Natural Language Processing Toolkit. InAssociation for Computational Linguistics (ACL) System Demonstrations. 55–60. http://www.aclweb. org/antho...

  19. [27]

    Pedregosa, G

    F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay

  20. [28]

    Verónica Pérez-Rosas, Bennett Kleinberg, Alexandra Lefevre, and Rada Mihalcea. 2017. Automatic detection of fake news. arXiv preprint arXiv:1708.07104 (2017)

  21. [29]

    Jiangwei Qi, T Trang, J Doong, S Kang, and Anna L Chien. 2016. Mis- information is prevalent in psoriasis-related YouTube videos. Derma- tology online journal 22, 11 (2016)

  22. [30]

    Kai Shu, Amy Sliva, Suhang Wang, Jiliang Tang, and Huan Liu. 2017. Fake news detection on social media: A data mining perspective. ACM SIGKDD Explorations Newsletter 19, 1 (2017), 22–36

  23. [31]

    RP Smith, P Devine, H Jones, A DeNittis, R Whittington, and JM Metz. 2003. Internet use by patients with prostate cancer undergoing radiotherapy. Urology 62, 2 (2003), 273–277

  24. [32]

    American Cancer Society. 2019. Cancer Facts and Figures 2019. (2019)

  25. [33]

    Peter L Steinberg, Shaun Wason, Joshua M Stern, Levi Deters, Brian Kowal, and John Seigne. 2010. YouTube as source of prostate cancer information. Urology 75, 3 (2010), 619–622

  26. [34]

    Shabbir Syed-Abdul, Luis Fernandez-Luque, Wen-Shan Jian, Yu-Chuan Li, Steven Crain, Min-Huei Hsu, Yao-Chin Wang, Dorjsuren Khan- dregzen, Enkhzaya Chuluunbaatar, Phung Anh Nguyen, et al. 2013. Misleading health-related information promoted through video-based social media: ano...

  27. [35]

    Ottokar Tilk and Tanel Alumäe. 2016. Bidirectional Recurrent Neural Network with Attention Mechanism for Punctuation Restoration. In Interspeech 2016

  28. [36]

    Kristina Toutanova, Dan Klein, Christopher D Manning, and Yoram Singer. 2003. Feature-rich part-of-speech tagging with a cyclic de- pendency network. In Proceedings of the 2003 conference of the North American chapter of the association for computational linguistics on human l...

  29. [1999]

    Journal of Epi- demiology & Community Health 53, 2 (1999), 105–111

    DISCERN: an instrument for judging the quality of written consumer health information on treatment choices. Journal of Epi- demiology & Community Health 53, 2 (1999), 105–111

  30. [2011]

    Journal of Machine Learning Research 12 (2011), 2825–2830

    Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research 12 (2011), 2825–2830

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.