{"id":"50f4fde1-23a2-4499-90f4-aaf0acef2199","arxiv_id":"2411.19169","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"ComViewer adds zoomable topic circles, support-type filters, and LLM-based note-taking and questioning to online mental health communities, and a within-subjects study shows more informational support and engagement versus a baseline.","lead":"This paper introduces ComViewer, a web tool that helps lurkers in online mental health communities find and understand helpful posts and comments. A 20-person study found that people using it recorded more useful advice and reported a more engaging experience than with a standard forum interface.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The RQ1 outcome measure may be inflated by ComViewer's auto-generated summaries; the significant increase in coded informational points could reflect LLM output rather than improved support-seeking.","rationale":"The reader's weakest assumption is the whole-package baseline and lack of ablation; I agree that is real, but I see a more acute version of the same causal-attribution problem. The RQ1 dependent variable is a document of key points, and in the ComViewer condition that document can be populated by the system's own LLM-generated summary and answers. The Appendix A prompt returns multiple subtitle/content units, so one highlighted passage can yield several countable 'meaningful points.' Thus the significant p<0.001 difference in informational points may not show that users extracted more support; it may show that the tool generated more text for them to transcribe. This is checkable from the persisted JSON logs and the written documents, so it is a concrete measurement-validity concern rather than a vague novelty worry. I do not recommend rejection. The paper is candid about limitations (Section 7.3.2), and the subjective measures (satisfaction, confidence, engagement, usefulness) are less affected by the summary-generation confound, though they are still subject to novelty and demand characteristics. The appropriate verdict remains conditional: the RQ1 outcome should be re-analyzed with source attribution, and the claims should be adjusted if the re-analysis shows the effect is primarily LLM-generated. This is why I set verdict_should_be to UNCHANGED rather than moving to reject.","tokens_in":29095,"tokens_out":5132,"duration_ms":45976,"concrete_test":"Re-code the 20 key-point documents with a source-attribution protocol: for each coded informational point in the ComViewer condition, determine by string overlap (verbatim n-gram or longest common subsequence) whether it appears in the participant's logged highlighted content, the auto-generated Note-taking summary, or the Questioning panel answers. Compute the proportion of coded points that are near-verbatim from generated outputs, and recompute the paired comparison after excluding those points. If the exclusion removes the significant difference (p>=0.05) or if the generated-source proportion is large (e.g., >50%), the RQ1 outcome is an artifact of LLM generation rather than improved support-seeking. Also record whether the two coders were blind to condition and report inter-rater reliability after the source-attribution coding.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the RQ1 count of 'meaningful points about informational support' measures what participants gained from the OMHC, rather than what ComViewer's LLM wrote for them. In ComViewer, the Note-taking panel automatically generates a structured summary of all highlighted content (Section 4.2.2, Summary page), and the Questioning panel generates answers to user questions. Participants were asked to summarize useful suggestions in a document while exploring (Section 5.3), and the procedure has them write key points after each task (Section 5.4). The Appendix A summarization prompt produces multiple 'subtitle/content' units, which map directly onto the coding unit of 'meaningful points.' A participant can produce many coded points by copying or lightly editing the generated summary, as P5 describes: 'I can highlight... and get an automatically generated summary... This helped me easily summarize the take-away advice' (Section 6.1). The baseline condition has no such generator, so the significant increase (5.25 vs 3.30, p<0.001) conflates support-seeking outcome with text-generation capability. The paper does not report whether key-point documents were compared against ComViewer's stored summaries/answers (which are persisted as JSON per Section 4.3.3), nor whether coders were blind to condition. This is a measurement-validity threat to the central quantitative claim, distinct from the acknowledged novelty/ablation limitation in Section 7.3.2.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"ComViewer is an interactive visual tool designed to help 'viewers' of online mental health communities (OMHCs) find and make sense of posts and comments that provide social support. The paper reports a formative study (N=10) that identifies five challenges in viewing OMHC content, derives four design requirements, and motivates three main panels: a Zoomable Posts panel for topic- and support-filtered exploration, a Note-taking panel with color-coded highlighting, automatic summarization, and mind-map generation, and a Questioning panel that uses an LLM to recommend questions and generate answers. The authors then describe a within-subjects user study (N=20) comparing ComViewer to a baseline OMHC-style interface without these panels. They report that ComViewer significantly increased the number of recorded meaningful informational-support points (5.25 vs. 3.30, p<0.001), improved satisfaction and confidence, raised engagement ratings, reduced perceived cognitive load, and led to more favorable usefulness and intention-to-use ratings. Qualitative interview data are used to explain these results and to derive three design considerations for future OMHC tools.","tokens_in":29390,"tokens_out":5220,"duration_ms":49572,"significance":"The paper addresses a real and underserved population—viewers who consume but do not post in OMHCs—and combines two currently relevant threads: information visualization for online communities and LLM-based sensemaking. The formative study is well conducted, the design requirements are concretely tied to the interface, the tool is substantial and the authors state the code will be open-sourced, and the statistical analyses are appropriate for the within-subjects design (paired t-test for the normally distributed informational-support count, Wilcoxon tests for ordinal ratings, and Latin-square counterbalancing). The authors are also transparent about limitations, including the novelty effect and the absence of an ablation study. However, as detailed in the major comments, the central quantitative claim—that ComViewer improves the outcome of support-seeking—is currently threatened by a measurement-validity problem: the outcome measure may count LLM-generated text rather than what participants actually gained from the community.","major_comments":[{"comment":"The RQ1 outcome measure is not shown to measure support-seeking rather than LLM output. In ComViewer, the Note-taking panel automatically generates a summary of all highlighted content (Section 4.2.2, Summary page), and the Appendix A.1 summarization prompt produces structured 'subtitle/content' units that map directly onto the coding unit of 'meaningful points' counted in Section 5.3. Participants were asked to write key points after each task (Section 5.4), and P5's quote in Section 6.1 ('I can highlight the content that may be helpful and get an automatically generated summary on all the highlighted content. This helped me easily summarize the take-away advice learned from the community.') shows that participants could produce many coded points by copying or lightly editing the generated summary. The baseline condition has no such generator. The paper does not report whether the key-point documents were compared with the persisted JSON summaries/answers stored per Section 4.3.3, nor whether the two coders were blind to condition. Therefore the significant increase in informational-support points (5.25 vs. 3.30, p<0.001) can be explained by the presence of an LLM text generator rather than by improved support-seeking. Please add a validation analysis (e.g., compute overlap between written points and stored JSON outputs; recode with condition-blinded coders; or run an ablation without automatic summaries) or explicitly soften the RQ1 claim to an analysis of what participants recorded, not what they learned.","section":"§5.3, §5.4, §4.2.2, Appendix A.1"},{"comment":"The comparison is whole-package, not a test of the three panels. The baseline interface removes not only the three named panels but also highlighting, questioning, automatic summarization, support filtering, similar-post highlighting, and the mind-map view, and it also omits crowd-contributed tags from both arms (Section 5.1). The authors acknowledge in Section 7.3.2 that no ablation study was conducted. Consequently the significant differences in engagement, cognitive load, and perceived usefulness cannot be attributed to any specific panel, and the qualitative per-panel attributions in Section 6.3 are self-reports rather than controlled evidence. This is not a fatal flaw if the central claim is read as whole-tool effectiveness, but the contributions section and Sections 6.2 and 6.3 currently present stronger panel-level conclusions than the design supports. The manuscript should either add an ablation condition or clearly limit the claims to the combined system and describe the per-panel results as exploratory.","section":"§5.1, §7.3.2"}],"minor_comments":[{"comment":"In the RQ1 results text, the sentence 'Overall, participants with ComViewer (Mean = 5.25, SD = 2.25) are significantly more satisfied with the social support they have received' uses the informational-support point count rather than the satisfaction mean reported in Table 2 (ComViewer: 5.60, SD = 1.10; baseline: 4.40, SD = 1.54). Please correct this mismatch.","section":"§6.1"},{"comment":"The coding procedure states that two authors coded twenty documents, met to discuss disagreements, and then applied the scheme to the rest; it would be helpful to report an inter-coder reliability statistic (e.g., Cohen's kappa or Krippendorff's alpha) for the final coding of the meaningful points.","section":"§5.3"},{"comment":"The support classifiers were trained on r/depression data from [42,44] and applied to r/Anxiety without reporting any validation of that cross-community transfer; a brief justification or reference to prior cross-community evaluation would strengthen the trustworthiness of the support-filtering and support-distribution features.","section":"§4.3.1"},{"comment":"The phrase 'Our contributions to CSCW communities are three folds' should be 'threefold' (or 'three-fold'), and the three contributions are listed as full sentences in a way that would be clearer as parallel noun phrases.","section":"§1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid systems-and-study contribution that fits CSCW, and the authors have been transparent about many limitations. My major concern is the RQ1 outcome measure: the reported effect could largely reflect direct copying or light editing of LLM-generated summaries, which is a measurement-validity issue rather than a mere presentation problem. If the authors can retrieve the persisted JSON logs and compare them with the participants' key-point documents, or otherwise demonstrate that the written points were not substantially LLM-derived, the paper could become acceptable after revision. If no such validation is possible, the authors should substantially weaken the central 'fruitful support-seeking outcome' claim. I would not reject at this stage, but I would not accept without this point being addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"ComViewer is a well-built tool for a real gap—OMHC viewers who don't post—and the study is thoughtfully designed. But the headline quantitative result, that ComViewer significantly increases informational support points (5.25 vs 3.30, p<0.001), is likely inflated by the tool's own text generation. In ComViewer, the Note-taking panel automatically produces a structured summary of highlighted content (Appendix A.1). Participants then wrote their key points after each task (Section 5.4). If they copied those auto-generated 'subtitle/content' units into their documents, the count of meaningful points goes up without meaningful support being gained. The paper does not report whether the documents were checked against the stored summaries, nor whether coders were blind to condition. That's a measurement-validity threat, distinct from the novelty/ablation limits the authors do acknowledge.\n\nWhat's genuinely new: prior visualization work for health communities targets moderators or posters; ComViewer addresses viewers, and the three-panel combination (zoomable circle packing with support filters, LLM note-taking, and questioning) is a sensible integration. The formative study (N=10) is solid and the design requirements follow from it. The user study is properly within-subjects and counterbalanced, and the qualitative interview data gives a concrete sense of how the tool helps. The authors are also honest about the lack of an ablation study and about the novelty effect.\n\nThe other soft spots are minor. The baseline omits all three panels plus crowd tags, so the comparison is whole-package; that is acknowledged. The support labels come from the authors' prior classifiers, which is fine but means the filtering feature depends on models not independently validated here.\n\nWho this is for: HCI/CSCW researchers working on online health communities, visualization, or LLM sensemaking tools. The design insights are useful; the empirical claim needs to be re-examined. I recommend it goes to peer review with a major-revision request. The RQ1 outcome should be re-analyzed with a check for copied summary text, or reframed as a product effectiveness measure. The authors are clearly capable of doing that.","headline":"ComViewer is a well-built tool for a real gap—OMHC viewers who don't post—but the study's central quantitative outcome is probably inflated by the tool's own auto-generated summaries.","tokens_in":29906,"tokens_out":4399,"would_cite":false,"duration_ms":39373,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ComViewer shows that an interactive, LLM-assisted interface for browsing mental-health forums helps viewers extract more informational support and feel more engaged than the standard list view.","keywords":["online mental health communities","social support","information seeking","sensemaking","information visualization","large language models","circle packing","within-subjects study"],"falsifier":"An ablation study that removes one of the three panels at a time—for example, replacing the zoomable circle packing with a plain list while keeping highlighting and questioning—would settle whether the full combination is responsible for the gains; if participants still record as many informational-support points and show the same engagement with the reduced interface, the three-panel design is not the cause. A longitudinal deployment that tracks returning users would also test whether the benefits persist after the novelty of the new tool wears off.","tokens_in":28903,"feed_emoji":"💬","tokens_out":9840,"duration_ms":91372,"temperature":0.7,"pith_summary":"ComViewer aims to establish that passive viewers of online mental health communities—people who read rather than post—can be helped to seek social support by a purpose-built visual tool. From a ten-person formative study the authors distil five viewer challenges, from articulating needs in keywords to organizing fragmented advice and understanding unfamiliar terms, and they answer them with three linked panels: zoomable topic/post/comment circles with social-support filters, color-coded note-taking with LLM summaries, and an LLM-powered questioning panel. A within-subjects study with twenty participants found that with ComViewer viewers recorded significantly more meaningful informational-support points (5.25 vs 3.30, p<0.001), reported higher satisfaction and confidence, and rated the process more engaging and less cognitively demanding than with a baseline list-based interface. If the claim holds, the accumulated experience in mental-health forums becomes more usable for the large read-only population, and the design offers a template for turning noisy peer content into structured, personalized takeaways.","feed_headline":"Viewers mine 60% more support advice with zoomable forum tool","feed_subtitle":"In a 20-person test, the three-panel interface beat the standard list on information gained, engagement, and mental load.","key_machinery":"The carrying mechanism is ComViewer's three-panel design mapped onto a four-stage information-seeking flow of search, read, organize, and digest. The Zoomable Posts panel implements the overview-first, zoom-and-filter, details-on-demand principle: an LDA topic model clusters search results into topic-level circles, each topic circle nests post circles sized by comment count, and each post circle nests comment circles; hovering reveals titles and similar posts, while bar charts filter posts and comments by predicted levels of seeking or providing emotional and informational support. The Note-taking panel lets viewers highlight text in colors, automatically collects highlights into matching folders, generates an editable LLM summary and a mind map, and links each note back to its source location. The Questioning panel generates recommended what/why/how questions for any selected content and provides editable LLM answers in a branching whiteboard layout. Together the panels turn an overwhelming linear feed into a structured, self-paced exploration with in-place comprehension support.","core_discovery":"The paper's central claim, stated by the authors, is that ComViewer improves both the outcome and the experience of support seeking for OMHC viewers compared to a standard community interface. In a within-subjects study with 20 participants, ComViewer produced significantly more recorded meaningful informational-support points (mean 5.25 vs 3.30; p<0.001), significantly higher satisfaction and confidence in the received support, significantly higher engagement (concentration, clarity, doability, and a sense of discovery), and significantly lower perceived cognitive load for both searching and sensemaking. Participants also rated ComViewer more useful, easier to use, and more likely to be reused. The gain in recorded emotional-support points was not significant, and the authors attribute the improvements to the simplified visual search, the organization and summarization of highlighted content, and the ability to question unfamiliar content in place.","pith_inferences":["A natural test of the paper's mechanism is an ablation that removes one panel at a time; participants' interview comments suggest the note-taking and questioning panels may contribute as much as the visual search to the recorded advice gain.","Because emotional-support points did not improve, a designer following this paper might add automatic highlighting of emotionally supportive sentences or an LLM that responds as an emotional supporter, to see whether both support types can be lifted.","The same zoom-organize-question loop should transfer to other large peer-advice corpora, such as design forums or technical Q&A sites, if the topic model and support classifiers are retrained, which would make the contribution a general information-seeking pattern rather than a mental-health-specific tool.","A longer deployment with returning users and behavioral logs, rather than a single-session questionnaire, would reveal whether the measured confidence and reduced mental load translate into sustained coping actions."],"forward_implications":["OMHC interfaces should offer a hierarchical, filterable overview of posts and comments rather than only a linear list, since viewers' search and read steps benefit from the zoomable circle packing.","Interactive note-taking with color-coded organization and automatic summaries can lower the cognitive load of turning fragmented comments into a personal action plan.","LLM-generated recommended questions and answers can support sensemaking of user-generated community content, not just of LLM output, by resolving unfamiliar terms and prompting deeper reflection.","If ComViewer's effect transfers, viewers of online mental health communities would leave sessions with more concrete, actionable advice and more confidence in applying it to their own situations.","The three-panel design can be adapted to other online communities by swapping topic models, support classifiers, summary formats, and LLM prompts, per the paper's design considerations."],"supporting_citations":[{"why":"The prior visual-analytics system for online health community administrators; ComViewer shifts the target user from administrator to viewer.","marker":"[22]"},{"why":"Documents the linear, list-based navigation of online conversations that the Zoomable Posts panel is designed to replace.","marker":"[16]"},{"why":"Supplies the Latent Dirichlet Allocation topic model used to cluster search results into topic-level circles.","marker":"[3]"},{"why":"Provides the classifiers that label posts' sought emotional and informational support levels in ComViewer's dataset.","marker":"[44]"},{"why":"Provides the classifiers that label comments' provided emotional and informational support levels and characterizes comment support quality.","marker":"[42]"},{"why":"The prior tool for multilevel LLM sensemaking; ComViewer extends that idea from LLM-generated content to user-generated community content.","marker":"[58]"},{"why":"Furnishes the overview-first, zoom-and-filter, details-on-demand design principle that shapes the Zoomable Posts panel.","marker":"[57]"}],"fun_headline_variants":["ComViewer helps forum viewers find 60% more support info","Visual tool boosts viewer support points and cuts mental load","ComViewer lowers mental load and increases support advice","Zoomable interface aids support seeking in mental health forums","Tool helps viewers get more support in mental health forums"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The measured gains are credited to ComViewer's three panels on the assumption that the baseline interface differs from ComViewer only by those panels, so a whole-package difference or simple novelty could in principle explain the same results.","fun_headline_variants_meta":{"raw":{"variants":["ComViewer helps forum viewers find 60% more support info","Visual tool boosts viewer support points and cuts mental load","ComViewer lowers mental load and increases support advice","Zoomable interface aids support seeking in mental health forums","Tool helps viewers get more support in mental health forums"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000807,"raw_usage":{"total_tokens":3531,"prompt_tokens":923,"completion_tokens":2608,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":539,"completion_tokens_details":{"reasoning_tokens":2529}},"tokens_in":539,"tokens_out":2608,"duration_ms":18835,"temperature":1.0,"reasoning_tokens":2529,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:26:49.255922+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An ablation study that removes one of the three panels at a time—for example, replacing the zoomable circle packing with a plain list while keeping highlighting and questioning—would settle whether the full combination is responsible for the gains; if participants still record as many informational-support points and show the same engagement with the reduced interface, the three-panel design is not the cause. A longitudinal deployment that tracks returning users would also test whether the benefits persist after the novelty of the new tool wears off.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The prior visual-analytics system for online health community administrators; ComViewer shifts the target user from administrator to viewer."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the linear, list-based navigation of online conversations that the Zoomable Posts panel is designed to replace."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Furnishes the overview-first, zoom-and-filter, details-on-demand design principle that shapes the Zoomable Posts panel."}],"review_version":1}