REVIEW 42 references
Public Health Advocacy Dataset: A Dataset of Tobacco Usage Videos from Social Media
T0 review · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The PHAD dataset of 5,730 tobacco videos with rich metadata is new, but its classification results likely rely on text labels rather than visual understanding.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
The paper also benchmarks a classifier: it extracts visual features from frames with a ResNet-50 network, then a second stage combines them with text features (descriptions and hashtags) using a vision-language encoder. The authors report 73.2% accuracy on an 8-way task (cigarettes, e-cigarettes, vaping devices, smokeless tobacco, etc.), beating several baselines.
There is a serious problem. The text features can include the hashtag and description that name the very product the classifier is supposed to identify. For example, a video found with the hashtag '#vape' and a description saying 'reviewing my new vape' carries the answer in the input. The ablation study shows that adding text boosts F1 from 62.4% to 71.8%, exactly the kind of jump you would expect from label leakage. The paper also has internal inconsistencies: the 'balanced sampling' description contradicts itself, and the stated frame count (4.3 million at 30fps) does not match the average video length (120 seconds), which would imply about 20 million frames.
If the dataset itself is cleaned and the text leakage is removed, it could still be useful for studying how tobacco products are promoted online. But as written, the classification results do not demonstrate that the model 'understands' the videos.
Extended reading notes
Core claim
The abstract claims: "This is the first dataset with these features providing a valuable resource for analyzing tobacco-related content and its impact," and that the two-stage Vision-Language classifier shows "superior performance" on categorizing tobacco products. If correct, PHAD provides a new 5,730-video public dataset with engagement metadata, and the VL approach offers a way to monitor tobacco promotion on social media.
Load-bearing premise
The assumption that the textual descriptors used in Eq. (6) (search hashtags, video descriptions) do not encode the target tobacco-product class. If hashtags like '#vape' or annotator-written summaries naming the product are fed as inputs, the reported classification accuracy is an artifact of label leakage, invalidating the main experimental claim. This premise enters in Section 6.2 Eq. (6) and in Appendix C, where the data format lists searchHashtag/name and Desc. as model-available fields.
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
assumptions (4)
- domain assumption The videos scraped from YouTube and TikTok via Apify constitute a representative sample of tobacco-related content on those platforms.
- domain assumption The manual annotations of tobacco product types are accurate and reliable enough to serve as ground truth.
- ad hoc to paper The textual inputs to the VL encoder (search keywords and descriptions) do not contain the target class labels.
- domain assumption Publicly available social media videos may be redistributed under CC BY-NC-SA without explicit creator consent.
Cite this review
Pith. "Pith review of Public Health Advocacy Dataset: A Dataset of Tobacco Usage Videos from Social Media." pith.science (2026). https://pith.science/paper/PF4YMDCE
@misc{pith2026241113572,
author = {Pith},
title = {Pith review of: Public Health Advocacy Dataset: A Dataset of Tobacco Usage Videos from Social Media},
year = {2026},
howpublished = {\url{https://pith.science/paper/PF4YMDCE}},
note = {Machine review of arXiv:2411.13572}
}
read the original abstract
The Public Health Advocacy Dataset (PHAD) is a comprehensive collection of 5,730 videos related to tobacco products sourced from social media platforms like TikTok and YouTube. This dataset encompasses 4.3 million frames and includes detailed metadata such as user engagement metrics, video descriptions, and search keywords. This is the first dataset with these features providing a valuable resource for analyzing tobacco-related content and its impact. Our research employs a two-stage classification approach, incorporating a Vision-Language (VL) Encoder, demonstrating superior performance in accurately categorizing various types of tobacco products and usage scenarios. The analysis reveals significant user engagement trends, particularly with vaping and e-cigarette content, highlighting areas for targeted public health interventions. The PHAD addresses the need for multi-modal data in public health research, offering insights that can inform regulatory policies and public health strategies. This dataset is a crucial step towards understanding and mitigating the impact of tobacco usage, ensuring that public health efforts are more inclusive and effective.
Reference graph
Works this paper leans on
-
[1]
Advances in neural infor- mation processing systems36 (2024) 25
Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. Advances in neural infor- mation processing systems36 (2024) 25
work page 2024
-
[2]
In: The Twelfth Interna- tional Conference on Learning Representa- tions (2024)
Zhu, D., Chen, J., Shen, X., Li, X., Elho- seiny, M.: MiniGPT-4: Enhancing vision- language understanding with advanced large language models. In: The Twelfth Interna- tional Conference on Learning Representa- tions (2024). https://openreview.net/forum? id=1tZbq88f27
work page 2024
-
[3]
In: International Conference on Machine Learning, pp
Li, J., Li, D., Xiong, C., Hoi, S.: Blip: Boot- strapping language-image pre-training for unified vision-language understanding and generation. In: International Conference on Machine Learning, pp. 12888–12900 (2022). PMLR
2022
-
[4]
In: International Conference on Machine Learning, pp
Li, J., Li, D., Savarese, S., Hoi, S.: Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large lan- guage models. In: International Conference on Machine Learning, pp. 19730–19742 (2023). PMLR
work page 2023
-
[5]
Advances in Neural Information Processing Systems 36 (2024)
Dai,W.,Li,J.,Li,D.,Tiong,A.M.H.,Zhao,J., Wang, W., Li, B., Fung, P.N., Hoi, S.: Instruct- blip: Towards general-purpose vision-language models with instruction tuning. Advances in Neural Information Processing Systems 36 (2024)
work page 2024
-
[6]
Current Addiction Reports10(1), 29– 37 (2023)
Lee, J., Suttiratana, S.C., Sen, I., Kong, G.: E- cigarette marketing on social media: a scoping review. Current Addiction Reports10(1), 29– 37 (2023)
work page 2023
-
[7]
JMIR public health and surveillance 6(1), 13673 (2020)
Kwon, M., Park, E.,et al.: Perceptions and sentiments about electronic cigarettes on social media platforms: systematic review. JMIR public health and surveillance 6(1), 13673 (2020)
work page 2020
-
[8]
International Journal of Envi- ronmental Research and Public Health20(10), 5761 (2023)
Jancey, J., Leaver, T., Wolf, K., Freeman, B., Chai, K., Bialous, S., Bromberg, M., Adams, P., Mcleod, M., Carey, R.N.,et al.: Promo- tion of e-cigarettes on tiktok and regulatory considerations. International Journal of Envi- ronmental Research and Public Health20(10), 5761 (2023)
work page 2023
Show all 42 references
-
[9]
Tobacco Control32(2), 251–254 (2023)
Sun, T., Lim, C.C., Chung, J., Cheng, B., Davidson, L., Tisdale, C., Leung, J., Gartner, C.E., Connor, J., Hall, W.D., et al.: Vap- ing on tiktok: a systematic thematic analysis. Tobacco Control32(2), 251–254 (2023)
2023
-
[10]
Nicotine & Tobacco Research 26(Supplement_1), 36–42 (2024) https://doi.org/10.1093/ntr/ntad184 https://academic.oup.com/ntr/article- pdf/26/Supplement_1/S36/56684002/ntad184.pdf
Murthy, D., Ouellette, R.R., Anand, T., Radhakrishnan, S., Mohan, N.C., Lee, J., Kong, G.: Using Computer Vision to Detect E-cigarette Content in TikTok Videos. Nicotine & Tobacco Research 26(Supplement_1), 36–42 (2024) https://doi.org/10.1093/ntr/ntad184 https://academic.oup....
2024 doi
-
[11]
Nicotine & Tobacco Research26(5), 552–560 (2023) https://doi.org/10.1093/ntr/ntad224 https://academic.oup.com/ntr/article- pdf/26/5/552/57181436/ntad224.pdf
Vassey, J., Kennedy, C.J., Herbert Chang, H.-C., Smith, A.S., Unger, J.B.: Scalable Surveillance of E-Cigarette Products on Insta- gram and TikTok Using Computer Vision. Nicotine & Tobacco Research26(5), 552–560 (2023) https://doi.org/10.1093/ntr/ntad224 https://academic.oup.c...
2023 doi
-
[12]
arXiv preprint arXiv:2109.08472 (2021)
Wang, M., Xing, J., Liu, Y.: Actionclip: A new paradigm for video action recognition. arXiv preprint arXiv:2109.08472 (2021)
2021 arXiv
-
[13]
arXiv preprint arXiv:2305.06310 (2023)
Chappa, N.V., Nguyen, P., Nelson, A.H., Seo, H.-S., Li, X., Dobbs, P.D., Luu, K.: Sogar: Self- supervised spatiotemporal attention-based social group activity recognition. arXiv preprint arXiv:2305.06310 (2023)
2023 arXiv
-
[14]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp
Chappa, N.V., Nguyen, P., Nelson, A.H., Seo, H.-S., Li, X., Dobbs, P.D., Luu, K.: Spartan: Self-supervised spatiotemporal transformers approach to group activity recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5157–5167 (2023)
2023
-
[15]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp
Tang, K., Zhang, H., Wu, B., Luo, W., Liu, W.: Learning to compose dynamic tree struc- tures for visual contexts. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6619– 6628 (2019)
2019
-
[16]
Advances in Neural Information 26 Processing Systems36 (2024)
Zhong, H., Mishra, S., Kim, D., Jin, S., Panda, R., Kuehne, H., Karlinsky, L., Saligrama, V., Oliva, A., Feris, R.: Learning human action recognition representations without real humans. Advances in Neural Information 26 Processing Systems36 (2024)
2024
-
[17]
In: Computer Vision–ECCV 2020: 16th Euro- pean Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16, pp
Zareian, A., Karaman, S., Chang, S.-F.: Bridg- ingknowledge graphstogenerate scenegraphs. In: Computer Vision–ECCV 2020: 16th Euro- pean Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16, pp. 606–623 (2020). Springer
2020
-
[18]
In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16, pp
Zareian, A., Wang, Z., You, H., Chang, S.-F.: Learning visual commonsense for robust scene graph generation. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16, pp. 642–657 (2020). Springer
2020
-
[19]
In: Proceedings of the European Conference on Computer Vision (ECCV), pp
Lu, C., Krishna, R., Bernstein, M., Fei-Fei, L.: Visual relationship detection with lan- guage priors. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 852–869 (2016)
2016
-
[20]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp
Zhong, Y., Shi, J., Yang, J., Xu, C., Li, Y.: Learning to generate scene graph from natu- ral language supervision. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1823–1834 (2021)
2021
-
[21]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp
Ye, K., Kovashka, A.: Linguistic structures as weak supervision for visual scene graph generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8289–8299 (2021)
2021
-
[22]
Advances in Neural Information Processing Systems36 (2024)
Nguyen, P., Quach, K.G., Kitani, K., Luu, K.: Type-to-track: Retrieve any object via prompt-based tracking. Advances in Neural Information Processing Systems36 (2024)
2024
-
[23]
Authorea Preprints (2024)
Chappa, N.V.S.R., Dobbs, P.D., Raj, B., Luu, K.: Flaash: Flow-attention adaptive semantic hierarchical fusion for multi-modal tobacco content analysis. Authorea Preprints (2024)
2024
-
[24]
In: 2022 26th International Conference on Pattern Recogni- tion (ICPR), pp
Truong, T.-D., Chappa, R.T.N., Nguyen, X.- B., Le, N., Dowling, A.P., Luu, K.: Otadapt: Optimal transport-based approach for unsu- pervised domain adaptation. In: 2022 26th International Conference on Pattern Recogni- tion (ICPR), pp. 2850–2856 (2022). IEEE
2022
-
[25]
IEEE Access 10, 93203– 93211 (2022)
Jalata, I., Chappa, N.V.S.R., Truong, T.-D., Helton, P., Rainwater, C., Luu, K.: Eqadap: Equipollent domain adaptation approach to image deblurring. IEEE Access 10, 93203– 93211 (2022)
2022
-
[26]
In: 2020 10th Annual Computing and Communication Workshop and Conference (CCWC), pp
Chappa, R.T.N., El-Sharkawy, M.: Squeeze- and-excitation squeezenext: An efficient dnn for hardware deployment. In: 2020 10th Annual Computing and Communication Workshop and Conference (CCWC), pp. 0691– 0697 (2020). IEEE
2020
-
[27]
In: 2024 IEEE Green Technologies Conference (GreenTech), pp
Chappa, N.V.R., McCormick, C., Gongora, S.R., Dobbs, P.D., Luu, K.: Advanced deep learning techniques for tobacco usage assess- ment in tiktok videos. In: 2024 IEEE Green Technologies Conference (GreenTech), pp. 162–163 (2024). IEEE
2024
-
[28]
Chappa, N.V.S.R., Dobbs, P.D., Luu, K.: Pub- lic health advocacy dataset: A dataset of tobaccousagevideosfromsocialmedia.Public Health 20, 4 (2024)
2024
-
[29]
Machine Vision and Appli- cations 35(4), 102 (2024)
Chappa, N.V.R., Nguyen, P., Dobbs, P.D., Luu, K.: React: Recognize every action every- where all at once. Machine Vision and Appli- cations 35(4), 102 (2024)
2024
-
[30]
Sensors 24(11), 3372 (2024)
Chappa, N.V.S.R., Nguyen, P., Le, T.H.N., Dobbs, P.D., Luu, K.: Hatt-flow: Hierarchical attention-flow mechanism for group-activity scene graph generation in videos. Sensors 24(11), 3372 (2024)
2024
-
[31]
Tobacco control 32(6), 739–746 (2023)
Kong, G., Schott, A.S., Lee, J., Dashtian, H., Murthy, D.: Understanding e-cigarette content and promotion on youtube through machine learning. Tobacco control 32(6), 739–746 (2023)
2023
-
[32]
JMIR infodemiology3(1), 42218 (2023)
Murthy, D., Lee, J., Dashtian, H., Kong, G., et al.: Influence of user profile attributes on e-cigarette–related searches on youtube: Machine learning clustering and classification. JMIR infodemiology3(1), 42218 (2023)
2023
-
[33]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 27 pp
Wang, C.-Y., Bochkovskiy, A., Liao, H.-Y.M.: Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 27 pp. 7464–7475 (2023)
2023
-
[34]
https://www.youtube.com/
YouTube: YouTube, Copyright @2024 (2024). https://www.youtube.com/
2024
-
[35]
https://www.tiktok.com/
TikTok: TikTok, Copyright @2024 (2024). https://www.tiktok.com/
2024
-
[36]
http://apify.com Accessed 2024-05-13
Apify: Apify Technologies, Copyright @2024 (2024). http://apify.com Accessed 2024-05-13
2024
-
[37]
In: Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition, pp
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition, pp. 770–778 (2016)
2016
-
[38]
Advances in neural information processing systems 30 (2017)
Vaswani, A., Shazeer, N., Parmar, N., Uszko- reit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. Advances in neural information processing systems 30 (2017)
2017
-
[39]
In: International Conference on Machine Learning, pp
Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., Duerig, T.: Scaling up visual and vision- language representation learning with noisy text supervision. In: International Conference on Machine Learning, pp. 4904–4916 (2021). PMLR
2021
-
[40]
Neural computation9(8), 1735– 1780 (1997)
Hochreiter, S., Schmidhuber, J.: Long short- term memory. Neural computation9(8), 1735– 1780 (1997)
1997
-
[41]
In: Proceedings of the IEEE Con- ference on Computer Vision and Pattern Recognition, pp
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: Proceedings of the IEEE Con- ference on Computer Vision and Pattern Recognition, pp. 6299–6308 (2017)
2017
-
[42]
https://arxiv.org/abs/2303.08774 29
OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad,L.,Akkaya,I.,Aleman,F.L.,Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Bal- com, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett- Shapiro, ...
2024 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.