{"id":"31b482eb-5ccb-4639-81e9-1c7d60b294b4","arxiv_id":"2501.13020","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A mixed-method study of 373 ADHD videos finds YouTube content skews professional and TikTok personal, with viewers and creators collectively policing quality and accessibility, though these efforts remain incomplete.","lead":"Researchers analyzed 373 ADHD-related videos and their comment sections on YouTube and TikTok, showing that creators and viewers jointly check quality through references, disclaimers, and timestamped summaries, while accessibility problems like length, pace, and missing captions remain.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The unweighted chi-square comparisons in §4.1 rest on a per-category top-five-plus-five-random sample with different search frames per platform; the reported platform differences may be artifacts of that sampling design.","rationale":"The reader's weakest assumption is that the critical-case sample may not generalize to the broader population of ADHD-relevant videos. My concern is related but sharper: even as an internal comparison, the chi-square tests in §4.1 are not valid estimates of platform differences because the sample is a non-probability stratified sample with equal allocation per category and different category constructions per platform. The top-five-plus-five-random selection intentionally balances categories, so marginal distributions are design-dependent, and the absence of weights or clustering means the reported p-values cannot be interpreted as population inferences. This affects the quantitative 'distinct profiles' part of the central claim, while leaving the qualitative findings about collective quality control largely intact. Because the paper already reads as a descriptive mixed-methods contribution and the manuscript's limitations partially acknowledge other threats, the appropriate verdict remains conditional; my read does not move the verdict, but it does sharpen the reason the quantitative claims should be treated cautiously. The proposed random-sample replication would settle whether the platform differences are real or artifacts of the sampling scheme.","tokens_in":24534,"tokens_out":7375,"duration_ms":78602,"concrete_test":"Re-run the quantitative comparison in §4.1 on a probability sample drawn from the same post-filter corpora: take a simple random sample of 373 videos per platform from the 26,318 TikTok and 2,281 YouTube videos with at least 20 comments, apply the same 61-code codebook to code creator type, content type, and video form, and recompute the chi-square tests with standard errors clustered by creator. If the significant differences reported in §4.1 (e.g., creator-type proportions, awareness/advocacy content, reference use) do not replicate, the per-category critical-case sampling is the source of the claimed platform contrasts. Alternatively, compute design-based weights from the selection probabilities within each category and re-run the tests; if the effect sizes collapse, the current analysis describes the sampling strata rather than the underlying video populations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.2.3 builds the final dataset by selecting, per topic category and platform, the five most-viewed videos plus five randomly sampled videos, with TikTok categories assigned from hashtags and YouTube categories assigned manually from keyword searches (ADHD plus category keyword). Section 4.1 then reports unweighted chi-square tests on these 373 videos to support the headline claim that TikTok and YouTube have distinct creator, content, and format profiles. This inference is not supported by the design. Equal allocation of up to ten videos per category force-represents rare categories on each platform: a topic that is nearly absent from one platform's natural distribution still contributes its full quota, inflating its apparent share. The top-five component deliberately overweights already-influential videos, and the two platforms' sampling frames differ (hashtag snowball for TikTok versus keyword search for YouTube), so the same category label does not denote the same population on both platforms. Without design weights, cluster-robust errors, or a random-sample replication, the p-values in §4.1 estimate differences among the sampled strata, not differences between the ADHD-video populations. The qualitative themes in §4.2-4.3 are less exposed because they are descriptive, but the quantitative 'distinct profiles' component of the central claim is directly affected. The paper's limitation section does not mention sampling representativeness; it only addresses commenter identity and the absence of health-expert input.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a mixed-method content analysis of 373 ADHD-relevant videos (189 from YouTube, 184 from TikTok) and their top comments. It characterizes creator types, content categories, and presentation formats, and develops qualitative themes around authority building, collective quality checking, accessibility practices, and tensions around self-diagnosis. Based on these findings, the authors propose design implications for making video-sharing platforms more reliable and ADHD-friendly. The central claim is that TikTok and YouTube have distinct creator and content profiles and that video quality is maintained through visible but imperfect collective efforts by creators and viewers.","tokens_in":24726,"tokens_out":3856,"duration_ms":43905,"significance":"If the quantitative parts are read descriptively, this is a useful empirical contribution to the accessibility and social computing literature: it offers a rich, systematically collected corpus, detailed qualitative coding, and concrete design directions grounded in community practices. The strength of the paper is its qualitative analysis, which is extensively evidenced with quotes and descriptions of creator and viewer behavior. The paper also makes a good-faith effort to address ethical considerations, and it cites prior work on ADHD content quality appropriately. However, the quantitative cross-platform comparison in Section 4.1 is not supported by the sampling design, and the coding process would benefit from explicit reliability reporting. The qualitative themes in Sections 4.2 and 4.3 are less exposed to these issues and are the more convincing contribution.","major_comments":[{"comment":"The unweighted chi-square comparisons in Section 4.1 are not supported by the sampling design. The final dataset is constructed by selecting, per topic category and platform, the five most-viewed videos plus five randomly sampled videos, with TikTok categories assigned from hashtags and YouTube categories assigned manually from keyword searches. Equal allocation of up to ten videos per category force-represents rare categories on each platform, the top-five component deliberately overweights already-influential videos, and the two platforms use different sampling frames. Consequently, the p-values in Section 4.1 estimate differences among the sampled strata, not differences between the broader populations of ADHD-relevant videos on YouTube and TikTok. The limitation section (5.3) acknowledges unverified commenter identities and the lack of health-expert input but does not mention sampling representativeness. Please either reframe Section 4.1 as descriptive statistics for the analytic sample and remove inferential claims, or redesign the sampling with design weights or a clearly defined target population (e.g., the five most-viewed videos per category) and restate the claims accordingly. The qualitative themes in Sections 4.2 and 4.3 are less affected by this issue.","section":"Section 3.3.2"},{"comment":"The coding process is described in detail, but no inter-rater reliability metric is reported for the 61-code video codebook or the 90-code comment codebook. The text states that two researchers independently coded 40 videos, then three researchers divided the remaining videos, with weekly checks and a fourth researcher overseeing the process. Without agreement statistics (e.g., Cohen's kappa or Krippendorff's alpha) on a subsample, the reproducibility of the codes that underlie both the quantitative counts and the thematic percentages cannot be independently assessed. Please report reliability on a subsample and describe how disagreements were resolved.","section":"Section 4.1"},{"comment":"Several chi-square tests in Section 4.1 involve very small observed counts, making the asymptotic approximation unreliable. For example, ADHD demographics are reported as 9.0% on YouTube versus 0.5% on TikTok, and institutions/organizations as 14.2% versus 3.2%, with only 373 videos and 252 creators in the corresponding tests. Expected cell counts below five are likely in these and other comparisons. The paper should report exact raw counts, use Fisher's exact test where appropriate, and avoid treating corrected p-values near 1.0 (e.g., p=1.00 for life beyond ADHD) as meaningful evidence of similarity. This is a statistical correctness issue in the current presentation, though it would be resolved if Section 4.1 were reframed descriptively.","section":"Section 4.1"}],"minor_comments":[{"comment":"The phrase 'balanced video sampling across all ADHD topics' is misleading: equal allocation per category does not balance the sample with respect to population shares, and it is better described as stratified quota sampling.","section":"Section 3.2.3"},{"comment":"The mapping between the 55 TikTok hashtags and the 26 YouTube categories is not fully specified; please clarify how hashtags were grouped into categories and whether the YouTube keyword categories were intended to match the TikTok hashtag groups exactly.","section":"Table 1"},{"comment":"The bar charts in Figure 1 show percentages without raw counts; please add counts or sample sizes so that readers can assess the precision of each proportion.","section":"Figure 1"},{"comment":"The creator-type chi-square test uses n=252 creators while the content and form tests use n=373 videos; please state the unit of analysis for each test explicitly.","section":"Section 4.1.1"},{"comment":"The number of comparisons controlled by the Bonferroni correction is not stated; please report how many tests were performed and whether the correction was applied per family or across all comparisons.","section":"Section 3.3.1"},{"comment":"The mean video lengths are reported with standard deviations larger than the means (e.g., 13.4 ± 22.3 minutes); consider reporting median and interquartile range as well.","section":"Section 4.3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the ASSETS audience and makes a real qualitative contribution. The main problem is the mismatch between the sampling design and the inferential statistics in Section 4.1; this is fixable by reframing those results as descriptive and adding coding reliability metrics. I did not find concerns about novelty or attribution. If the authors revise the quantitative claims accordingly, I would support publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here is my take. The paper does something genuinely useful: it maps the ADHD video landscape on YouTube and TikTok from a collective quality control angle, and the qualitative analysis is well-evidenced with quotes and systematic coding. The themes around authority building, collective quality checking, and accessibility improvement are credible and new relative to the clinical scoring studies (Yeung, Thapa) and the community-building work (Eagle & Ringland). The design implications are concrete and grounded in the data.\n\nThe soft spot is exactly where the stress-test lands. Section 4.1's chi-square tests compare YouTube and TikTok using a per-category sample of five most-viewed plus five random videos, with different sampling frames (hashtag-based for TikTok, keyword-based for YouTube). That design does not support unweighted population-level comparisons. Rare categories are force-represented by the equal allocation, and the top-five component overweights influential videos. The p-values are better read as descriptive of the sample, not as evidence of platform population differences. The limitation section is honest about commenter identity and the lack of health-expert input, but it does not mention sampling representativeness, which is a real gap. I would ask the authors to either soften the quantitative claims or re-analyze with design weights, cluster-robust errors, or a random-sample replication.\n\nAlso missing: no inter-rater reliability metric for the coding, and several percentages rest on small cell counts, making some specific comparisons fragile. No code or data is shipped, which is a pity for a descriptive study.\n\nThat said, the qualitative findings in Sections 4.2 and 4.3 are not exposed to the sampling critique in the same way; the themes are inductively derived and well-illustrated with quotes. This is a solid contribution to accessibility and health HCI, and I would expect it to be cited. I would send it to peer review with a request for revision rather than desk reject.","headline":"Useful qualitative map of ADHD video ecosystems, but the quantitative platform comparisons are overclaimed given the sampling design.","tokens_in":25298,"tokens_out":2100,"would_cite":true,"duration_ms":21201,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ADHD videos on YouTube and TikTok follow distinct creator and format profiles, with quality sustained by collective but imperfect viewer-creator efforts.","keywords":["ADHD","video-sharing platforms","YouTube","TikTok","content quality","accessibility","collective quality control","online health information"],"falsifier":"A replication that draws a random sample from the full population of ADHD-tagged videos posted in a fixed period—rather than the top five most viewed plus five random per category—and finds that health professionals and organizations are just as prevalent on TikTok as on YouTube would falsify the claimed creator-type difference.","tokens_in":177,"feed_emoji":"🎥","tokens_out":3793,"duration_ms":93970,"temperature":0.7,"pith_summary":"This paper tries to establish that ADHD-relevant videos on YouTube and TikTok differ systematically in who makes them, how they are presented, and what quality problems they carry, and that content quality is maintained through collective, visible, yet imperfect efforts by creators, viewers, and platforms. The authors analyzed 373 videos with associated comments using mixed methods. They argue that TikTok is dominated by individuals with ADHD sharing personal experiences in short, creative formats, while YouTube hosts more health professionals and organizations producing longer, more formal educational content. The paper also claims that quality control happens through authority building, references, disclaimers, and comment-based checking, and that accessibility issues such as video length, pace, distracting elements, and caption quality are central to whether ADHD viewers can use the content. The stakes matter because many people with ADHD rely on these videos for symptom recognition, help-seeking, and support.","feed_headline":"373 ADHD videos reveal distinct platform cultures and collective quality checks","feed_subtitle":"Viewers and creators jointly police accuracy and accessibility, but quality signals stay scattered and easy to miss.","key_machinery":"The analytical machinery is a mixed-method corpus study: 373 videos (189 YouTube, 184 TikTok) and the top 20 comments under each, collected through hashtag- and keyword-based searches and sampled by critical case sampling. The authors coded videos and comments through iterative thematic analysis, producing codebooks of 61 video codes and over 90 comment codes, then cross-referenced them into themes. This design lets the paper connect quantitative distributions of creator types, content types, and presentation forms with qualitative evidence of how quality and accessibility are negotiated in practice. The conceptual mechanism that carries the argument is the framing of quality control as a collective, multi-stakeholder effort—platform, creator, and viewer—rather than a property of individual videos.","core_discovery":"The central claim is that ADHD-relevant content on video-sharing platforms is not a single genre: TikTok and YouTube serve different purposes and present different quality and accessibility challenges. On TikTok, 69.6% of creators were self-disclosed individuals with ADHD, and content leaned toward role-play, memes, and first-person POV, fostering intimate community support; on YouTube, institutions and organizations made up 14.2% of creators, health professionals were more common, and formal forms such as talks, interviews, news, and documentaries were significantly more prevalent. The paper further claims that quality control is a collective process: creators establish authorship through identity disclosure and platform recognition, add references, and post disclaimers; viewers assess, challenge, supplement, and summarize content in comments. These efforts have clear failure modes—vague credentials, irrelevant references, low-visibility disclaimers, and a gap between encouraging clinical diagnosis and providing resource pointers. Accessibility, including video length, slow pace, distracting sounds and visuals, and caption quality, is treated as an integral part of video quality for ADHD audiences, with viewers and creators improvising workarounds like timestamped breakdowns and speed adjustments.","pith_inferences":["Because the sample overrepresents top-viewed videos, the exact percentages are less portable than the qualitative mechanisms; a broader random sample could shift the numbers without undermining the collective-effort finding.","The results imply that for ADHD viewers, content quality and accessibility are inseparable, so any quality-assessment tool that ignores captions, pacing, and length will systematically underrate videos that are actually useful.","The same three-part mechanism—authority building, collective checking, and accessibility improvement—likely appears in other neurodivergent health topics such as autism, but with different authority markers and different self-diagnosis stakes.","A testable design consequence is that consolidating quality signals into a single visible, low-distraction interface element would improve trust without requiring viewers to read descriptions or profiles."],"forward_implications":["Platform designers cannot apply one quality standard across both platforms: TikTok's personal, identity-driven content and YouTube's professional, institutional content serve different help-seeking and community-building needs.","Quality-control signals such as disclaimers, references, and creator credentials need to be placed where ADHD viewers will actually see them, since these signals are currently scattered across descriptions, profiles, and comments and are often missed.","Comment-based quality checking is already functioning and could be amplified by platform features that surface challenges, corrections, and supplementary experiences, while guarding against misinformation that also appears in comments.","Accessibility features—chapters, summaries, pacing controls, caption quality, and distraction reduction—are not secondary polish but core determinants of whether ADHD viewers can benefit from health content.","Any quality-control intervention must respect the community's tension around self-diagnosis: encouraging clinical help should not stigmatize members who self-diagnose after careful research because professional evaluation is inaccessible to them.","Long, expert-produced YouTube videos, which are most likely to contain authoritative medical information, are often the least accessible to ADHD viewers due to length and pace, creating a paradox where the most reliable content is hardest to consume."],"supporting_citations":[{"why":"Establishes TikTok and related platforms as sources of shared expertise and diagnosis tensions in ADHD communities, grounding the community context.","marker":"[18]"},{"why":"Supplies the earlier finding that 38.4% of ADHD-related YouTube videos were misleading, the quality baseline this study extends.","marker":"[75]"},{"why":"Supplies the earlier finding that 52% of top ADHD TikTok videos were rated misleading, motivating the TikTok quality analysis.","marker":"[86]"},{"why":"Provides prior interview evidence on video accessibility challenges for people with ADHD, which this study corroborates through comment analysis.","marker":"[31]"},{"why":"Documents neurodiverse TikTok creators' captioning practices and infrastructure needs, supporting the accessibility findings.","marker":"[69]"},{"why":"Provides the clinical-versus-pragmatic content taxonomy for mental health videos on TikTok that the paper extends with ADHD-specific subcategories.","marker":"[92]"},{"why":"Frames the credibility, trust, and safety concerns specific to video-sharing platforms that motivate the quality-control analysis.","marker":"[55]"},{"why":"Supplies the thematic analysis method used to code videos and comments.","marker":"[14]"}],"fun_headline_variants":["ADHD video quality is a group effort on TikTok and YouTube","Creators and viewers team up to police ADHD video accuracy","TikTok and YouTube diverge in how ADHD content gets vetted","Collective checks, scattered signals: ADHD video quality control","373 ADHD videos show platform splits and joint quality policing"],"cache_read_input_tokens":27392,"weakest_assumption_plain":"The 373-video sample, built from platform-default search rankings and the five most-viewed plus five random videos per topic category, is treated as representative enough to characterize ADHD content on each platform.","fun_headline_variants_meta":{"raw":{"variants":["ADHD video quality is a group effort on TikTok and YouTube","Creators and viewers team up to police ADHD video accuracy","TikTok and YouTube diverge in how ADHD content gets vetted","Collective checks, scattered signals: ADHD video quality control","373 ADHD videos show platform splits and joint quality policing"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000506,"raw_usage":{"total_tokens":2462,"prompt_tokens":935,"completion_tokens":1527,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":1443}},"tokens_in":551,"tokens_out":1527,"duration_ms":10476,"temperature":1.0,"reasoning_tokens":1443,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T16:31:11.511592+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication that draws a random sample from the full population of ADHD-tagged videos posted in a fixed period—rather than the top five most viewed plus five random per category—and finds that health professionals and organizations are just as prevalent on TikTok as on YouTube would falsify the claimed creator-type difference.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the earlier finding that 38.4% of ADHD-related YouTube videos were misleading, the quality baseline this study extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the earlier finding that 52% of top ADHD TikTok videos were rated misleading, motivating the TikTok quality analysis."},{"cited_title":"Hey, Can You Add Captions?","cited_arxiv_id":null,"evidence_quote":"Documents neurodiverse TikTok creators' captioning practices and infrastructure needs, supporting the accessibility findings."}],"review_version":1}