{"id":"270347b4-4239-49ac-b8ed-68214b3713b0","arxiv_id":"2506.10932","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"People with schizophrenia use two video structures (talk-to-camera and in-the-moment), two verbal strategies (direct expression and storytelling), and distinct visual styles to disclose emotions on YouTube, with in-the-moment videos more common among joy-focused vlogs.","lead":"Researchers analyzed 200 YouTube vlogs by people with schizophrenia and mapped how they disclose fear, sadness, and joy through words, setting, and visual style. The study offers an early framework for how video aesthetics and structure shape emotional disclosure and viewer support on social media, with implications for designing safer mental-health platforms.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim about visual elements fostering supportive viewer responses rests on anecdotal comments, not systematic engagement data.","rationale":"The reader's weakest assumption about uploader diagnosis is legitimate, but it is not the most load-bearing concern for the central claim. Even if every uploader were verified as having schizophrenia, the headline assertion about visual construction fostering supportive viewer responses would still rest on a handful of illustrative comments and an informal observation about view counts. I therefore focus on the evidential gap between the presented qualitative data and the causal-sounding claim. This does not change the overall conditional verdict: the paper is transparently exploratory, its visual analysis framework is a useful contribution, and the two-structure emotion finding is quantified. However, the abstract's 'We found' should be softened unless systematic viewer-response data are added. The conditional verdict already captures the need for revision, so no change to the reader's recommendation is warranted.","tokens_in":14933,"tokens_out":4619,"duration_ms":56277,"concrete_test":"Re-analyze the 200-video corpus: for each video, code the visual dimensions from the paper's framework (stage richness, color palette, anonymity, video structure) and collect engagement outcomes (views, likes, and a sentiment-coded measure of supportive comments, e.g., the first 50 comments per video). Fit a mixed-effects regression predicting supportive engagement from the visual dimensions with random intercepts for video and controls for major emotion, video length, and channel subscriber count. If the visual dimensions do not jointly and significantly predict supportive engagement, the abstract's claim should be reframed as a hypothesis for future work rather than a finding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim, that deliberate visual construction 'appears to foster more supportive and engaged viewer responses,' is not supported by the reported analysis. In the Findings (Visual Frame, Stage), the only direct evidence is the impressionistic statement 'It seems that videos with a carefully crafted stage created opportunities for viewer interaction,' followed by two illustrative comments, plus a note that 17 of 19 anonymous videos have fewer than 50 views. No systematic coding of viewer comments, no engagement metrics linked to visual features, and no comparison across visual-style categories are presented. The Discussion itself uses pervasively hedged language ('may shape,' 'might influence,' 'possible relationships'), indicating the authors recognize the evidence is preliminary. The Limitations section acknowledges emotion-detection error but does not flag this gap. Consequently, the abstract's 'We found' overstates what the data can establish: the visual-to-response relation is an observed correlation in a handful of examples, not a demonstrated effect. Since this is the paper's headline contribution, this is the most load-bearing weakness; the only quantitative result, the structure-emotion chi-square, concerns video structure and emotion, not viewer responses.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a qualitative content analysis of 200 English-language YouTube vlogs retrieved using schizophrenia-related keywords, aiming to characterize how people with schizophrenia disclose fear, sadness, and joy through video. The authors identify two video structures (talk-to-camera and in-the-moment), two verbal strategies (direct emotional expression and emotional storytelling), and a three-part visual framework (vlogger, stage, style). They use an unspecified emotion-detection model on transcripts to assign a major emotion to each video, then report pairwise chi-square tests linking video structure to emotion. The abstract's headline claim is that deliberately constructed visual elements appear to foster more supportive and engaged viewer responses, but the supporting evidence in the findings is anecdotal and the Discussion is explicitly hedged.","tokens_in":15103,"tokens_out":4031,"duration_ms":45923,"significance":"If treated as a preliminary qualitative study, the paper is a useful contribution to video-mediated emotion disclosure research: it extends prior text-based disclosure work to multimodal vlogs, proposes a reusable visual analysis framework, uses a relatively large corpus for a qualitative study, and includes careful ethical protections (blurring faces, paraphrasing quotes) and a clear generative-AI disclosure. The two-part structural finding is a concrete, checkable quantitative result. However, the headline claim about visual elements causing supportive viewer responses is not supported by the reported data, and the quantitative emotion classification rests on an unvalidated model. These problems are fixable through reframing and additional analysis, so the manuscript warrants major revision rather than rejection.","major_comments":[{"comment":"The central claim that deliberate visual construction \"fosters more supportive and engaged viewer responses\" is not established by the evidence. In the Stage subsection, support consists of the sentence \"It seems that videos with a carefully crafted stage created opportunities for viewer interaction\" followed by two illustrative viewer comments; the Anonymity subsection notes that 17 of 19 anonymous videos have fewer than 50 views and 12 lack comments, but no systematic coding of viewer comments, no engagement metrics linked to visual features, and no comparison across visual-style categories are provided. The Discussion itself uses hedged wording (\"possible relationships,\" \"might influence,\" \"we speculate\"), indicating preliminary observation rather than a demonstrated effect. Because this is the manuscript's headline contribution, the abstract should be revised to state the claim as an anecdotal observation, or the analysis should be extended with systematic viewer-response data.","section":"Abstract; Findings, Visual Frame, Stage; Discussion"},{"comment":"The emotion labels used for the structure-emotion chi-square tests come from an \"emotion detection model\" that is neither identified nor validated: no model name, version, training data, or performance metrics are reported, and the footnote marker in the text has no corresponding footnote in the submitted manuscript. Because Table 1 and the pairwise p-values (p=.0027, p=.0003, p=.7702) depend on these labels, the quantitative result is not independently verifiable. Please specify the model, report validation on a held-out or manually labeled sample, and ideally provide inter-coder reliability for the human emotion coding.","section":"Method, Data Analysis (Verbal narrative); Limitations"},{"comment":"The sample is described as 200 YouTube videos \"created by individuals with schizophrenia,\" but the retrieval procedure (keyword search combining schizophrenia-related terms with \"vlog,\" \"vlogging,\" and \"story,\" followed by removal of incomplete and institution-created videos) does not verify that uploaders self-identify as having schizophrenia, nor does it screen out family members, advocates, or others discussing the illness. Since the study's central object is the emotion disclosure of people with schizophrenia, the sampling criteria should either demonstrate how creator status was verified (e.g., self-identification in the video or channel description) or explicitly scope the claims to \"videos discussing schizophrenia,\" with the creator population treated as a plausible but unverified characteristic.","section":"Method, Data Collection"},{"comment":"The pairwise chi-square tests are reported as p-values only, without the underlying test statistics, effect sizes, or confidence intervals, and without correction for multiple comparisons. Given that these tests are the only quantitative results in the paper and one of them drives the secondary claim about joy videos, please report the full statistical details and either justify the lack of adjustment or apply a correction such as Bonferroni or false-discovery-rate control.","section":"Findings, Relationship between Video Structure and Emotions"}],"minor_comments":[{"comment":"The In-the-Moment percentage for the joy group is written as \"37.5\" while the other cells include percent signs; please add \"%\" for consistency.","section":"Table 1"},{"comment":"The text refers to \"Hoffman's theoretical framework of self-presentation (Hoffman et al., 2019),\" but the framework used throughout the paper is Goffman's dramaturgical self-presentation (Goffman, 1959), and the cited Hoffman et al. (2019) paper is about explainable AI metrics rather than self-presentation. Please correct the citation and the name.","section":"Discussion, Developing Visual Analysis Frameworks"},{"comment":"The sentence \"17 out of the 19 videos analyzed have fewer than 50 views\" is ambiguous about whether \"analyzed\" refers to all 200 videos or to the anonymous subset; please clarify the base for this observation.","section":"Findings, Visual Frame, Vlogger"},{"comment":"The superscript \"3\" referring to the emotion detection model has no matching footnote in the submitted text; please provide the full model reference or a footnote explaining its provenance.","section":"Method, Data Analysis (Verbal narrative)"},{"comment":"The statement that \"low-key lighting... frequently accompanies more somber narratives\" and the related color interpretations are presented as observations without counts or systematic coding across the 200 videos; adding a sentence on how often such patterns occurred and how they were coded would strengthen the claim.","section":"Findings, Visual Frame, Style"}],"recommendation":"major_revision","confidential_remarks":"This is a promising qualitative HCI/CSCW contribution that is currently oversold in its abstract and overly hedged in its discussion. The central visual-to-viewer-response claim needs either systematic evidence or an explicit downgrade to an exploratory observation; the emotion-detection model needs full specification; and the sampling claims need to be scoped honestly. All of these are achievable within the manuscript's scope, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should read this as a decent exploratory qualitative study with an over-reach in the abstract. What's actually new: the two video structures (talk-to-camera vs. in-the-moment), the direct-expression vs. storytelling verbal strategies, and the vlogger/stage/style visual frame categories, all applied to a population we don't see much of in HCI. The chi-square result—joy videos more likely to use in-the-moment structure than fear or sadness—is a real, if minor, quantitative complement. The visual analysis framework is modular and could be reused by computational studies, which the authors explicitly position it for.\n\nThe paper does some things well. The coding process is described with enough detail to be credible, and the limitations section is candid about the emotion-detection issue. The ethical care (blurring faces, paraphrasing quotes) is appropriate for this sensitive content.\n\nNow the soft spots, in order of weight. The stress-test note is right: the abstract claim that deliberate visual construction 'appears to foster more supportive and engaged viewer responses' rests on two illustrative comments and an 'it seems' in the findings. There's no systematic coding of viewer comments, no engagement metrics tied to visual features, and no comparison across visual categories. The Discussion itself is properly hedged, so the authors know the evidence is preliminary; the abstract just got ahead of them. The sampling also doesn't verify that uploaders actually have schizophrenia, as opposed to advocates or family members. And the emotion detection model is not specified or validated, which adds noise to the emotion-group comparisons. These are real issues, but for a qualitative conference paper they're reparable rather than fatal. The framework and descriptive findings stand.\n\nWho is this for? Researchers in HCI/CSCW studying health vlogging, self-presentation theory, or platform design for mental health communities. It deserves a serious referee, but with a requirement to soften the causal claim and either add systematic engagement analysis or explicitly frame the visual-response link as preliminary. I'd send it to review, not desk-reject, with major revisions.","headline":"Solid descriptive framework and one clean structure-emotion association, but the headline claim about visual staging driving supportive viewer responses is anecdotal and needs reining in before publication.","tokens_in":15586,"tokens_out":1420,"would_cite":true,"duration_ms":19486,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Vloggers with schizophrenia disclose emotion through visual staging as much as through words, and deliberate visual construction appears to draw more supportive viewer responses.","keywords":["schizophrenia","YouTube vlogs","emotion disclosure","video-mediated communication","self-presentation","visual analysis framework","mental health","viewer engagement"],"falsifier":"A concrete check is to read each channel's uploads for first-person self-identification as someone with schizophrenia, such as the apparent creator saying \"my schizophrenia\" or \"my diagnosis.\" If a substantial share of the 200 videos lacks any such first-person marker, then the study's object—emotion disclosure by people with schizophrenia—would not be established, and the visual-engagement patterns would need to be re-attributed.","tokens_in":14744,"feed_emoji":"🎥","tokens_out":8352,"duration_ms":83842,"temperature":0.7,"pith_summary":"This paper asks how people with schizophrenia use YouTube vlogs to disclose fear, sadness, and joy through both speech and image. The authors analyzed 200 middle-length vlogs uploaded in 2023, coding video structure, verbal narrative, and frame-level visual features such as setting, color, and staging. Their central claim is that the deliberate construction of visual elements—environmental settings and aesthetic choices—appears to foster more supportive and engaged viewer responses. That claim matters because if visual staging shapes how viewers react, video platforms could be designed to help vulnerable creators disclose emotions in ways that invite support rather than stigma.","feed_headline":"Visual staging in schizophrenia vlogs draws engaged comments","feed_subtitle":"In 200 YouTube vlogs, deliberate settings and color appear to draw more engaged viewer responses","key_machinery":"The central object is the paper's visual analysis framework: a coding scheme that examines each vlog along three dimensions—vlogger (identity and anonymity), stage (setting, activity, and other people), and style (color, lighting, and aesthetics)—alongside a two-way coding of verbal narrative as direct emotion expression or emotional storytelling. This framework does the work of converting raw video frames into comparable categories, allowing the authors to connect visual self-presentation to emotional state and to viewer engagement. A second mechanism is the emotion detection model applied to transcripts, which assigns each video a major emotion and feeds the chi-square tests that link joy to in-the-moment structure.","core_discovery":"The paper's central claim is that emotion disclosure in schizophrenia vlogs is multimodal: verbal expression is only half of the story, and the visual frame—the room, the lighting, the color palette, whether the creator faces the camera or films an activity—carries emotional meaning and affects how viewers respond. On the structural side, the paper reports two video formats: talk-to-camera diary entries and in-the-moment footage, with in-the-moment appearing in 37.5% of joy videos versus 10.4% of fear videos and 11.9% of sadness videos. On the visual side, it observes that videos with carefully composed stages and aesthetic choices drew comments that engaged with the setting and appreciated the experience, while anonymous framing was associated with very low view counts: 17 of 19 such videos had fewer than 50 views and 12 had no comments. The paper presents these as observed patterns that lay groundwork for large-scale quantitative testing, not as proven causal effects.","pith_inferences":["Beyond the paper: if the visual-construction pattern is causal, platform templates that suggest calm staging or simplified editing during distress states could reduce cognitive load, but a causal test would need randomized comparisons.","Beyond the paper: the anonymity finding suggests a protection-versus-support dilemma: the same visual hiding that shields creators from stigma may also cut them off from the supportive community the paper documents, and platform design should address both sides.","Beyond the paper: the automated emotion detection draws only on transcripts, so word-frequency imbalances could drive the structure-emotion association; re-coding with human raters or visual features would test this.","Beyond the paper: the visual analysis framework is transportable to other illness vlogs and to computational studies, but only after inter-rater reliability is measured on its three dimensions."],"forward_implications":["Joy-centered vlogs are statistically more likely to use in-the-moment structure than fear- or sadness-centered vlogs, indicating that emotional state is tied to video format choice.","Because vloggers combine direct emotion statements with emotional storytelling, transcript-only analysis undercounts disclosure; multimodal reading is necessary to capture how emotions are narrated.","Visual anonymity appears to carry a visibility cost: 17 of 19 anonymous videos had fewer than 50 views and 12 had no comments, suggesting a tradeoff between protection and engagement.","Deliberate staging, such as detailed backgrounds and polished color choices, appears to invite viewer comments that engage with the creator's world, supporting the idea that visual construction shapes reception.","The observed link between aesthetics and engagement implies a possible visibility hierarchy on video platforms, where certain mental-health narratives may be amplified over others by algorithms and audience taste."],"supporting_citations":[{"why":"Supplies the self-presentation lens used to analyze vlogs as performed identity work.","marker":"(Goffman, 1959)"},{"why":"Establishes health vlogs as social support and informs the video-structure coding and engagement observations.","marker":"(Huh et al., 2014)"},{"why":"Documents self-disclosure formats in health vlogs, guiding the verbal-disclosure analysis.","marker":"(Misoch, 2014)"},{"why":"Provides semiotic theory used to read visual signs and personal spaces in sampled frames.","marker":"(Barthes, 1968)"},{"why":"Basis for the claim that visual attributes of mental-health disclosures carry distinct self-disclosure meaning.","marker":"(Manikonda & De Choudhury, 2017)"},{"why":"Shows how crying and negative affect on YouTube expand public discourse, backing the vulnerability-on-camera observations.","marker":"(Berryman & Kavka, 2018)"},{"why":"Broaden-and-build theory is used to explain why distress videos favor the minimal talk-to-camera format.","marker":"(Fredrickson, 2004)"},{"why":"Supplies the thematic analysis method used to code verbal narratives and visual themes.","marker":"(Braun & Clarke, 2006)"}],"fun_headline_variants":["Schizophrenia vlogs: visual staging boosts viewer engagement","Why schizophrenia vloggers' settings matter more than words","In vlogs, schizophrenia emotion shows in the frame","YouTube schizophrenia vlogs: visual choices drive comments","Fear, sadness, joy: schizophrenia vlog staging shapes response"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the sampled videos were actually made by people who have schizophrenia; the sample was gathered by keyword search and by excluding institution-created videos, without verifying that uploaders self-identify as having the condition.","fun_headline_variants_meta":{"raw":{"variants":["Schizophrenia vlogs: visual staging boosts viewer engagement","Why schizophrenia vloggers' settings matter more than words","In vlogs, schizophrenia emotion shows in the frame","YouTube schizophrenia vlogs: visual choices drive comments","Fear, sadness, joy: schizophrenia vlog staging shapes response"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000466,"raw_usage":{"total_tokens":2303,"prompt_tokens":898,"completion_tokens":1405,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":1327}},"tokens_in":514,"tokens_out":1405,"duration_ms":10709,"temperature":1.0,"reasoning_tokens":1327,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:12:46.920529+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check is to read each channel's uploads for first-person self-identification as someone with schizophrenia, such as the apparent creator saying \"my schizophrenia\" or \"my diagnosis.\" If a substantial share of the 200 videos lacks any such first-person marker, then the study's object—emotion disclosure by people with schizophrenia—would not be established, and the visual-engagement patterns would need to be re-attributed.","supporting_citations":[],"review_version":1}