Pith. sign in

REVIEW 42 references

Public Health Advocacy Dataset: A Dataset of Tobacco Usage Videos from Social Media

T0 review · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The PHAD dataset of 5,730 tobacco videos with rich metadata is new, but its classification results likely rely on text labels rather than visual understanding.

arxiv 2411.13572 v1 pith:PF4YMDCE submitted 2024-11-12 cs.CV

classification cs.CV
keywords datasethealthpublictobaccousageadvocacycontentengagement
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

PHAD is a new dataset of 5,730 short videos about tobacco products collected from TikTok and YouTube. Each video comes with metadata: number of views, likes, comments, shares, the hashtag used to find it, and a written description. The creators say this is the first such collection with all these extra fields.

The paper also benchmarks a classifier: it extracts visual features from frames with a ResNet-50 network, then a second stage combines them with text features (descriptions and hashtags) using a vision-language encoder. The authors report 73.2% accuracy on an 8-way task (cigarettes, e-cigarettes, vaping devices, smokeless tobacco, etc.), beating several baselines.

There is a serious problem. The text features can include the hashtag and description that name the very product the classifier is supposed to identify. For example, a video found with the hashtag '#vape' and a description saying 'reviewing my new vape' carries the answer in the input. The ablation study shows that adding text boosts F1 from 62.4% to 71.8%, exactly the kind of jump you would expect from label leakage. The paper also has internal inconsistencies: the 'balanced sampling' description contradicts itself, and the stated frame count (4.3 million at 30fps) does not match the average video length (120 seconds), which would imply about 20 million frames.

If the dataset itself is cleaned and the text leakage is removed, it could still be useful for studying how tobacco products are promoted online. But as written, the classification results do not demonstrate that the model 'understands' the videos.

Extended reading notes

Core claim

The abstract claims: "This is the first dataset with these features providing a valuable resource for analyzing tobacco-related content and its impact," and that the two-stage Vision-Language classifier shows "superior performance" on categorizing tobacco products. If correct, PHAD provides a new 5,730-video public dataset with engagement metadata, and the VL approach offers a way to monitor tobacco promotion on social media.

Load-bearing premise

The assumption that the textual descriptors used in Eq. (6) (search hashtags, video descriptions) do not encode the target tobacco-product class. If hashtags like '#vape' or annotator-written summaries naming the product are fed as inputs, the reported classification accuracy is an artifact of label leakage, invalidating the main experimental claim. This premise enters in Section 6.2 Eq. (6) and in Appendix C, where the data format lists searchHashtag/name and Desc. as model-available fields.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

This is an empirical dataset paper with no fitted physical constants. The central claims rest on dataset curation choices and on the assumption that text features do not leak labels; the free-parameter list is empty because the model weights are ordinary trained parameters, not hand-fitted constants.

assumptions (4)
  • domain assumption The videos scraped from YouTube and TikTok via Apify constitute a representative sample of tobacco-related content on those platforms.
    Section 3.3 says videos were selected by automated scraping and manual curation, but the datasheet admits representativeness is not validated.
  • domain assumption The manual annotations of tobacco product types are accurate and reliable enough to serve as ground truth.
    Appendix B.6 mentions inter-annotator agreement but reports no values, and the evaluation treats the labels as correct.
  • ad hoc to paper The textual inputs to the VL encoder (search keywords and descriptions) do not contain the target class labels.
    Eq. (6) feeds z_t into the classifier; Appendix C lists searchHashtag/name and Desc. as available fields, and no evidence is provided that these are written independently of the product type.
  • domain assumption Publicly available social media videos may be redistributed under CC BY-NC-SA without explicit creator consent.
    Section 3.4 states consent was not obtained and relies on platform terms of service, which is a legal assumption the paper does not verify.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Public Health Advocacy Dataset: A Dataset of Tobacco Usage Videos from Social Media." pith.science (2026). https://pith.science/paper/PF4YMDCE

@misc{pith2026241113572,
  author       = {Pith},
  title        = {Pith review of: Public Health Advocacy Dataset: A Dataset of Tobacco Usage Videos from Social Media},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PF4YMDCE}},
  note         = {Machine review of arXiv:2411.13572}
}
read the original abstract

The Public Health Advocacy Dataset (PHAD) is a comprehensive collection of 5,730 videos related to tobacco products sourced from social media platforms like TikTok and YouTube. This dataset encompasses 4.3 million frames and includes detailed metadata such as user engagement metrics, video descriptions, and search keywords. This is the first dataset with these features providing a valuable resource for analyzing tobacco-related content and its impact. Our research employs a two-stage classification approach, incorporating a Vision-Language (VL) Encoder, demonstrating superior performance in accurately categorizing various types of tobacco products and usage scenarios. The analysis reveals significant user engagement trends, particularly with vaping and e-cigarette content, highlighting areas for targeted public health interventions. The PHAD addresses the need for multi-modal data in public health research, offering insights that can inform regulatory policies and public health strategies. This dataset is a crucial step towards understanding and mitigating the impact of tobacco usage, ensuring that public health efforts are more inclusive and effective.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

42 extracted references · 37 canonical work pages

  1. [1]

    Advances in neural infor- mation processing systems36 (2024) 25

    Liu, H., Li, C., Wu, Q., Lee, Y.J.: Visual instruction tuning. Advances in neural infor- mation processing systems36 (2024) 25

  2. [2]

    In: The Twelfth Interna- tional Conference on Learning Representa- tions (2024)

    Zhu, D., Chen, J., Shen, X., Li, X., Elho- seiny, M.: MiniGPT-4: Enhancing vision- language understanding with advanced large language models. In: The Twelfth Interna- tional Conference on Learning Representa- tions (2024). https://openreview.net/forum? id=1tZbq88f27

  3. [3]

    In: International Conference on Machine Learning, pp

    Li, J., Li, D., Xiong, C., Hoi, S.: Blip: Boot- strapping language-image pre-training for unified vision-language understanding and generation. In: International Conference on Machine Learning, pp. 12888–12900 (2022). PMLR

  4. [4]

    In: International Conference on Machine Learning, pp

    Li, J., Li, D., Savarese, S., Hoi, S.: Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large lan- guage models. In: International Conference on Machine Learning, pp. 19730–19742 (2023). PMLR

  5. [5]

    Advances in Neural Information Processing Systems 36 (2024)

    Dai,W.,Li,J.,Li,D.,Tiong,A.M.H.,Zhao,J., Wang, W., Li, B., Fung, P.N., Hoi, S.: Instruct- blip: Towards general-purpose vision-language models with instruction tuning. Advances in Neural Information Processing Systems 36 (2024)

  6. [6]

    Current Addiction Reports10(1), 29– 37 (2023)

    Lee, J., Suttiratana, S.C., Sen, I., Kong, G.: E- cigarette marketing on social media: a scoping review. Current Addiction Reports10(1), 29– 37 (2023)

  7. [7]

    JMIR public health and surveillance 6(1), 13673 (2020)

    Kwon, M., Park, E.,et al.: Perceptions and sentiments about electronic cigarettes on social media platforms: systematic review. JMIR public health and surveillance 6(1), 13673 (2020)

  8. [8]

    International Journal of Envi- ronmental Research and Public Health20(10), 5761 (2023)

    Jancey, J., Leaver, T., Wolf, K., Freeman, B., Chai, K., Bialous, S., Bromberg, M., Adams, P., Mcleod, M., Carey, R.N.,et al.: Promo- tion of e-cigarettes on tiktok and regulatory considerations. International Journal of Envi- ronmental Research and Public Health20(10), 5761 (2023)

Show all 42 references
  1. [9]

    Tobacco Control32(2), 251–254 (2023)

    Sun, T., Lim, C.C., Chung, J., Cheng, B., Davidson, L., Tisdale, C., Leung, J., Gartner, C.E., Connor, J., Hall, W.D., et al.: Vap- ing on tiktok: a systematic thematic analysis. Tobacco Control32(2), 251–254 (2023)

  2. [10]

    Nicotine & Tobacco Research 26(Supplement_1), 36–42 (2024) https://doi.org/10.1093/ntr/ntad184 https://academic.oup.com/ntr/article- pdf/26/Supplement_1/S36/56684002/ntad184.pdf

    Murthy, D., Ouellette, R.R., Anand, T., Radhakrishnan, S., Mohan, N.C., Lee, J., Kong, G.: Using Computer Vision to Detect E-cigarette Content in TikTok Videos. Nicotine & Tobacco Research 26(Supplement_1), 36–42 (2024) https://doi.org/10.1093/ntr/ntad184 https://academic.oup....

  3. [11]

    Nicotine & Tobacco Research26(5), 552–560 (2023) https://doi.org/10.1093/ntr/ntad224 https://academic.oup.com/ntr/article- pdf/26/5/552/57181436/ntad224.pdf

    Vassey, J., Kennedy, C.J., Herbert Chang, H.-C., Smith, A.S., Unger, J.B.: Scalable Surveillance of E-Cigarette Products on Insta- gram and TikTok Using Computer Vision. Nicotine & Tobacco Research26(5), 552–560 (2023) https://doi.org/10.1093/ntr/ntad224 https://academic.oup.c...

  4. [12]

    arXiv preprint arXiv:2109.08472 (2021)

    Wang, M., Xing, J., Liu, Y.: Actionclip: A new paradigm for video action recognition. arXiv preprint arXiv:2109.08472 (2021)

  5. [13]

    arXiv preprint arXiv:2305.06310 (2023)

    Chappa, N.V., Nguyen, P., Nelson, A.H., Seo, H.-S., Li, X., Dobbs, P.D., Luu, K.: Sogar: Self- supervised spatiotemporal attention-based social group activity recognition. arXiv preprint arXiv:2305.06310 (2023)

  6. [14]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Chappa, N.V., Nguyen, P., Nelson, A.H., Seo, H.-S., Li, X., Dobbs, P.D., Luu, K.: Spartan: Self-supervised spatiotemporal transformers approach to group activity recognition. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5157–5167 (2023)

  7. [15]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Tang, K., Zhang, H., Wu, B., Luo, W., Liu, W.: Learning to compose dynamic tree struc- tures for visual contexts. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6619– 6628 (2019)

  8. [16]

    Advances in Neural Information 26 Processing Systems36 (2024)

    Zhong, H., Mishra, S., Kim, D., Jin, S., Panda, R., Kuehne, H., Karlinsky, L., Saligrama, V., Oliva, A., Feris, R.: Learning human action recognition representations without real humans. Advances in Neural Information 26 Processing Systems36 (2024)

  9. [17]

    In: Computer Vision–ECCV 2020: 16th Euro- pean Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16, pp

    Zareian, A., Karaman, S., Chang, S.-F.: Bridg- ingknowledge graphstogenerate scenegraphs. In: Computer Vision–ECCV 2020: 16th Euro- pean Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16, pp. 606–623 (2020). Springer

  10. [18]

    In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16, pp

    Zareian, A., Wang, Z., You, H., Chang, S.-F.: Learning visual commonsense for robust scene graph generation. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XXIII 16, pp. 642–657 (2020). Springer

  11. [19]

    In: Proceedings of the European Conference on Computer Vision (ECCV), pp

    Lu, C., Krishna, R., Bernstein, M., Fei-Fei, L.: Visual relationship detection with lan- guage priors. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 852–869 (2016)

  12. [20]

    In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp

    Zhong, Y., Shi, J., Yang, J., Xu, C., Li, Y.: Learning to generate scene graph from natu- ral language supervision. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 1823–1834 (2021)

  13. [21]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp

    Ye, K., Kovashka, A.: Linguistic structures as weak supervision for visual scene graph generation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8289–8299 (2021)

  14. [22]

    Advances in Neural Information Processing Systems36 (2024)

    Nguyen, P., Quach, K.G., Kitani, K., Luu, K.: Type-to-track: Retrieve any object via prompt-based tracking. Advances in Neural Information Processing Systems36 (2024)

  15. [23]

    Authorea Preprints (2024)

    Chappa, N.V.S.R., Dobbs, P.D., Raj, B., Luu, K.: Flaash: Flow-attention adaptive semantic hierarchical fusion for multi-modal tobacco content analysis. Authorea Preprints (2024)

  16. [24]

    In: 2022 26th International Conference on Pattern Recogni- tion (ICPR), pp

    Truong, T.-D., Chappa, R.T.N., Nguyen, X.- B., Le, N., Dowling, A.P., Luu, K.: Otadapt: Optimal transport-based approach for unsu- pervised domain adaptation. In: 2022 26th International Conference on Pattern Recogni- tion (ICPR), pp. 2850–2856 (2022). IEEE

  17. [25]

    IEEE Access 10, 93203– 93211 (2022)

    Jalata, I., Chappa, N.V.S.R., Truong, T.-D., Helton, P., Rainwater, C., Luu, K.: Eqadap: Equipollent domain adaptation approach to image deblurring. IEEE Access 10, 93203– 93211 (2022)

  18. [26]

    In: 2020 10th Annual Computing and Communication Workshop and Conference (CCWC), pp

    Chappa, R.T.N., El-Sharkawy, M.: Squeeze- and-excitation squeezenext: An efficient dnn for hardware deployment. In: 2020 10th Annual Computing and Communication Workshop and Conference (CCWC), pp. 0691– 0697 (2020). IEEE

  19. [27]

    In: 2024 IEEE Green Technologies Conference (GreenTech), pp

    Chappa, N.V.R., McCormick, C., Gongora, S.R., Dobbs, P.D., Luu, K.: Advanced deep learning techniques for tobacco usage assess- ment in tiktok videos. In: 2024 IEEE Green Technologies Conference (GreenTech), pp. 162–163 (2024). IEEE

  20. [28]

    Chappa, N.V.S.R., Dobbs, P.D., Luu, K.: Pub- lic health advocacy dataset: A dataset of tobaccousagevideosfromsocialmedia.Public Health 20, 4 (2024)

  21. [29]

    Machine Vision and Appli- cations 35(4), 102 (2024)

    Chappa, N.V.R., Nguyen, P., Dobbs, P.D., Luu, K.: React: Recognize every action every- where all at once. Machine Vision and Appli- cations 35(4), 102 (2024)

  22. [30]

    Sensors 24(11), 3372 (2024)

    Chappa, N.V.S.R., Nguyen, P., Le, T.H.N., Dobbs, P.D., Luu, K.: Hatt-flow: Hierarchical attention-flow mechanism for group-activity scene graph generation in videos. Sensors 24(11), 3372 (2024)

  23. [31]

    Tobacco control 32(6), 739–746 (2023)

    Kong, G., Schott, A.S., Lee, J., Dashtian, H., Murthy, D.: Understanding e-cigarette content and promotion on youtube through machine learning. Tobacco control 32(6), 739–746 (2023)

  24. [32]

    JMIR infodemiology3(1), 42218 (2023)

    Murthy, D., Lee, J., Dashtian, H., Kong, G., et al.: Influence of user profile attributes on e-cigarette–related searches on youtube: Machine learning clustering and classification. JMIR infodemiology3(1), 42218 (2023)

  25. [33]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 27 pp

    Wang, C.-Y., Bochkovskiy, A., Liao, H.-Y.M.: Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 27 pp. 7464–7475 (2023)

  26. [34]

    https://www.youtube.com/

    YouTube: YouTube, Copyright @2024 (2024). https://www.youtube.com/

  27. [35]

    https://www.tiktok.com/

    TikTok: TikTok, Copyright @2024 (2024). https://www.tiktok.com/

  28. [36]

    http://apify.com Accessed 2024-05-13

    Apify: Apify Technologies, Copyright @2024 (2024). http://apify.com Accessed 2024-05-13

  29. [37]

    In: Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition, pp

    He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Com- puter Vision and Pattern Recognition, pp. 770–778 (2016)

  30. [38]

    Advances in neural information processing systems 30 (2017)

    Vaswani, A., Shazeer, N., Parmar, N., Uszko- reit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. Advances in neural information processing systems 30 (2017)

  31. [39]

    In: International Conference on Machine Learning, pp

    Jia, C., Yang, Y., Xia, Y., Chen, Y.-T., Parekh, Z., Pham, H., Le, Q., Sung, Y.-H., Li, Z., Duerig, T.: Scaling up visual and vision- language representation learning with noisy text supervision. In: International Conference on Machine Learning, pp. 4904–4916 (2021). PMLR

  32. [40]

    Neural computation9(8), 1735– 1780 (1997)

    Hochreiter, S., Schmidhuber, J.: Long short- term memory. Neural computation9(8), 1735– 1780 (1997)

  33. [41]

    In: Proceedings of the IEEE Con- ference on Computer Vision and Pattern Recognition, pp

    Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: Proceedings of the IEEE Con- ference on Computer Vision and Pattern Recognition, pp. 6299–6308 (2017)

  34. [42]

    https://arxiv.org/abs/2303.08774 29

    OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad,L.,Akkaya,I.,Aleman,F.L.,Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Bal- com, V., Baltescu, P., Bao, H., Bavarian, M., Belgum, J., Bello, I., Berdine, J., Bernadett- Shapiro, ...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.