{"id":"68c671b9-87c3-493d-a99b-5cd36d0faf7f","arxiv_id":"2505.20584","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors present a Streamlit dashboard for searching and visualizing mpox tweets and claim rising cynical sentiment, without releasing code or data.","lead":"This paper describes a researcher-focused dashboard for monitoring mpox-related tweets, built with Streamlit and combining multiple Twitter datasets. It reports increased tweet volume and dominant cynicism after the CDC's 2024 mpox designation, but provides few quantitative details.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The dashboard's key finding—cynicism becoming the dominant sentiment in mpox discourse—is produced by an unspecified, unvalidated labeling algorithm, and the abstract's volume comparison lacks any described comparable 2023 baseline; both claims are unverifiable as written.","rationale":"Reading the paper in good faith, it is a systems/dashboard paper whose scientific value rests on two empirical assertions: (1) the dashboard detected a marked increase in mpox tweets after August 2024 compared to 2023, and (2) the analysis reveals that cynicism and misinformation dominate recent mpox discourse. Both require a dataset that is representative across time and labels that are valid. The weaker of these two conditions is the labeling step in Section 5.2, which is described only as 'we applied an algorithm.' Without the algorithm's identity, training data, parameters, or validation, the cluster proportions in Figure 3 cannot be reproduced or audited. This is precisely the reader's weakest assumption, and I agree. I add a related concern: the volume comparison in the abstract is also under-specified—no counts, denominators, or collection-source breakdowns are given, and the methods section does not explicitly describe a 2023 dataset, creating a possible internal inconsistency with Section 6. Both issues point to the same remedy: the authors must release or precisely document the data pipeline and labeling method. The verdict should remain conditional rather than reject because the underlying dashboard is plausible and the missing evidence could be supplied in a revision. No ad hominem is intended; the concern is about the completeness of the scientific record, not the authors' conduct.","tokens_in":6000,"tokens_out":7496,"duration_ms":78827,"concrete_test":"Ask the authors for the full labeling pipeline and labeled data behind Figure 3; have two independent annotators label a stratified random sample of 500 tweets into the Figure 3 categories, and compute Cohen's kappa and per-class F1 against the system labels. If kappa is below 0.6 or the pipeline is unavailable, the 'cynicism dominant' claim and the misinformation-trend claim are not established. Also inspect the date/source distribution of the raw data to confirm a 2023 baseline collected with the same method exists.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.2 states that 'we applied an algorithm to label tweets according to their primary topics and sentiment' and then reports that 'cynicism became the dominant sentiment' and that misinformation clusters are visible over time. No name, architecture, training data, prompt, threshold, validation set, or accuracy measure is provided for this algorithm. The label set also mixes levels of analysis—'cynicism' is an attitude, 'COVID-19 comparisons' is a topic, and 'misinformation' is a veracity claim—so Figure 3 cannot be interpreted without an explicit labeling rule. If this labeling step is biased, every downstream claim about sentiment and misinformation trends is unsupported. A separate internal problem affects the abstract's volume claim: no raw counts or collection-source distributions are reported, and the methods section only describes Zeeschuimer for 2024 tweets without describing a comparable 2023 sample, even though the abstract compares 2024 to 2023. The limitations section (Section 6) acknowledges the short timeframe but does not address either the missing validation or the comparability of collection methods. Because the paper's central value proposition is that the dashboard surfaces reliable sentiment and misinformation trends, these omissions are load-bearing rather than cosmetic.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes a Streamlit-based dashboard for searching and visualizing mpox-related tweets, intended for local public health agencies. It combines three datasets (a UT Austin Computational Media Lab Spritzer-stream sample, a publicly available May 2022 Monkeypox dataset, and 3,000 tweets collected via Zeeschuimer in 2024), applies a Python filtering pipeline, and visualizes keyword trends, engagement metrics, and a color-coded cluster graph of sentiment/topic categories over time. The paper's central empirical claims are that the dashboard recorded a marked increase in tweet volume in 2024 compared with 2023, and that cynicism became the dominant sentiment in mpox-related discussions. The manuscript is primarily a system/tool description with an accompanying literature review, and it does not report quantitative validation of its measurement pipeline.","tokens_in":6358,"tokens_out":2751,"duration_ms":30044,"significance":"If the central claims were quantitatively supported, the dashboard would be a useful proof-of-concept for real-time infodemic monitoring by public health stakeholders, complementing prior tools such as PoxVerifi and the Johns Hopkins COVID-19 dashboard. The paper also usefully situates the work within the misinformation and health-communication literature. However, as written, the two headline findings—a marked 2024 versus 2023 volume increase and the dominance of cynicism—are not backed by reported counts, error bars, or any validation of the sentiment/topic labeling algorithm. The contribution is therefore better described as a dashboard demonstration than as a validated empirical study of mpox discourse.","major_comments":[{"comment":"The sentiment/topic labeling algorithm is not named, described, or validated. Section 5.2 states only that 'we applied an algorithm to label tweets according to their primary topics and sentiment' before reporting that 'cynicism became the dominant sentiment.' No training data, model architecture, prompt template, lexicon, decision threshold, or accuracy measure is given, and the category set mixes attitudes ('cynicism'), topics ('COVID-19 comparisons'), and veracity claims ('misinformation'). Since the paper's main analytical result depends entirely on these labels, the authors should specify the algorithm, report validation against a human-coded gold standard (including inter-annotator agreement or precision/recall/F1), and define each category operationally.","section":"§5.2"},{"comment":"The claim of a 'marked increase in tweet volume compared to 2023' is not supported by any quantitative evidence: no raw tweet counts, rates, collection durations, or sample sizes are reported for either year. Moreover, §4 describes a Zeeschuimer collection of 3,000 tweets from 2024 but gives no comparable 2023 collection procedure, so the baseline for the comparison is undefined. The authors should report the underlying numbers and describe how the 2023 and 2024 samples were collected, including any differences in collection mechanism that could confound the volume comparison.","section":"Abstract and §5.2"},{"comment":"The three constituent datasets use different collection mechanisms (Twitter Spritzer Stream random sample versus browser-capture-based Zeeschuimer versus a third-party May 2022 dataset), yet the Methods section does not report the date range, keyword filter, deduplication, or per-dataset tweet counts. As a result, the time series in Figure 3, which begins in April 2024, may reflect collection artifacts rather than genuine shifts in discourse. The authors should state precisely which dataset contributes which dates, how duplicate tweets were handled, and how the daily proportions in Figure 3 were computed, including the denominator.","section":"§4 and Figure 3"}],"minor_comments":[{"comment":"The running head contains a typo: 'MPOX-R ELATED' should be 'MPOX-RELATED.'","section":"Title page"},{"comment":"Reference [11] is cited for the claim that public confusion was intensified by the absence of FDA-approved at-home testing in the United States, but the cited paper addresses COVID-19 vaccination hesitancy in South Africa; the citation appears mismatched and should be corrected or replaced.","section":"References"},{"comment":"References [11] and [14] are the same arXiv preprint (Perikli et al., 2307.15072) listed twice; the duplicate should be removed and the citation numbering adjusted.","section":"References"},{"comment":"The platform is referred to inconsistently as 'X (formerly Twitter),' 'Twitter,' and 'tweets'; the authors should choose one terminology and apply it consistently.","section":"Throughout"},{"comment":"Figure 3 has no axis labels or y-axis units, so the reader cannot tell what 'proportion' means; the axes and the denominator of the proportion should be labeled clearly.","section":"§5.1 and Figure 3"},{"comment":"The paper reports that Zeeschuimer collected 3,000 tweets from 2024 but does not report the sizes of the other two datasets; reporting all dataset sizes and date ranges would improve reproducibility.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a tool/demo paper whose empirical claims currently outrun its evidence. The central fixes—reporting counts and baselines, and specifying/validating the labeling algorithm—are within the scope of a major revision rather than a rejection. I would also encourage the editor to consider whether the system-description nature of the paper is a good fit for a research journal; if the authors instead reframe the contribution as a dashboard demonstration without the unverified empirical claims, a shorter paper or a demo track might be more appropriate. No concerns about citation ethics beyond the reference mismatch noted in the minor comments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Hi — this one is a practical dashboard paper with a real gap between what it claims and what it actually shows. The artifact itself is plausible; the analytics are not substantiated.\n\nWhat's genuinely new: applying the dashboard-monitoring approach to mpox discourse, with an interactive Streamlit interface, keyword search, and engagement filters. That is a reasonable extension of prior work like PoxVerifi and the COVID dashboards, and the authors are appropriately citing their own tool only as prior context. The literature review is broad and the design rationale is sensible, even if not groundbreaking.\n\nThe soft spots are where the paper makes its load-bearing claims. The sentiment/topic labeling algorithm in Section 5.2 is unnamed, undescribed, and unvalidated. No training data, no prompt, no thresholds, no accuracy measure. The label set mixes attitudes ('cynicism'), topics ('COVID-19 comparisons'), and veracity claims ('misinformation'), so Figure 3 cannot be interpreted without an explicit labeling rule. The abstract's volume spike claim also lacks raw counts, and the methods section only describes Zeeschuimer for 2024 tweets without specifying a comparable 2023 sample, even though the comparison is central. The limitations section (Section 6) acknowledges the short timeframe but does not address these missing details.\n\nThese are not minor omissions — they are what the paper is selling. If the labeling step is biased, the 'cynicism became dominant' finding is unsupported. However, the dashboard itself could still be useful to local public health teams as a search and visualization tool, even without the analytic claims. The paper is not incoherent; it is just under-supported.\n\nMy recommendation: send it to serious peer review, but with a clear instruction that the empirical claims must either be removed, downgraded to observations, or fully supported with methodology, validation, and data or code release. As written, it reads more like a systems/design report than a rigorous empirical study. A workshop or an applied health-informatics venue could be a good fit.","headline":"A well-intentioned dashboard paper whose central sentiment claims rest on an unnamed, unvalidated labeling step — useful as a design report, not yet as an empirical study.","tokens_in":6751,"tokens_out":1591,"would_cite":false,"duration_ms":18732,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that an interactive dashboard can surface near-real-time shifts in mpox discourse, and that in its 2023–2024 data tweet volume rose sharply after the CDC's August 2024 designation while cynicism became the dominant…","keywords":["mpox","dashboard","misinformation","social media","health communication","sentiment analysis","topic clustering","public health surveillance"],"falsifier":"Take a random sample of the dashboard's 2024 mpox tweets, have independent human annotators label each tweet's sentiment and topic without seeing the dashboard's cluster labels, and compare the two distributions; if the human-labeled data does not show cynicism as the dominant category, or shows the algorithm systematically mislabels tweets, the central claim fails.","tokens_in":5801,"feed_emoji":"📊","tokens_out":5494,"duration_ms":52650,"temperature":0.7,"pith_summary":"The paper argues that an interactive social-media dashboard can support local public health agencies by making mpox-related discourse and misinformation visible in near real time. It reports that after the CDC designated mpox an emerging threat in August 2024, dashboard-recorded tweet volume rose markedly compared with 2023. It further reports that in the more recent data, cynicism toward public health institutions became the dominant sentiment in mpox discussions on X (formerly Twitter). If these observations hold, public health communicators would have an early signal of both rising attention and falling trust, allowing them to intervene before misinformation solidifies. The paper is a proof-of-concept application rather than a validation study.","feed_headline":"Mpox tweets surged and turned cynical after CDC alert, dashboard shows","feed_subtitle":"Tracking 2023-2024 posts, it ties the August CDC designation to higher volume and rising distrust.","key_machinery":"The carrying mechanism is the dashboard itself: a Streamlit web application that ingests filtered mpox tweets from three sources, presents searchable tables, lets users filter by keywords and engagement metrics such as like, reply, and retweet counts, and plots a time series of topic and sentiment clusters. The clustering relies on an algorithm that labels each tweet into thematic categories including cynicism, COVID-19 comparisons, government action, and misinformation; the proportion of tweets in each category per day is then shown as a color-coded scatterplot.","core_discovery":"The paper's central claim is that a researcher-focused dashboard built from 2023–2024 mpox-related tweets can reveal real-time changes in public discourse, and that when applied to its collected data it recorded a marked increase in tweet volume after the CDC designated mpox an emerging virus in August 2024. The same dashboard analysis found that cynicism, defined as distrust in public health institutions and traditional news sources, became the dominant sentiment in the more recent mpox-related discussions, a pattern the authors interpret as a warning that audiences may be increasingly receptive to unofficial or misleading health information.","pith_inferences":["If the volume surge is real, the same dashboard pattern could serve as a generic early-warning signal for other emerging pathogens, though the paper does not control for platform-wide volume changes or differences in data collection between years.","The dominance of cynicism may be specific to X's user base or driven by particular news cycles; applying the same pipeline to Reddit or Facebook data would test whether the mood shift reflects broader public sentiment.","A natural extension would be to use the dashboard's keyword filters in an intervention study: if agencies respond to cynicism-heavy clusters with trust-building messages, one could measure whether the proportion of cynicism-labeled tweets falls in subsequent weeks."],"forward_implications":["Local public health agencies could use the dashboard to detect when discourse volume jumps after official announcements, enabling faster communication responses.","If cynicism is truly dominant, health communication strategies may need to prioritize trust repair over simply providing more facts.","The keyword and engagement filters allow researchers to track specific misinformation narratives and see which ones gain traction.","The cluster graph could reveal when COVID-19 comparisons or government-action narratives spike, helping agencies tailor messages to current concerns."],"supporting_citations":[{"why":"Supplies the hydrated, pre-filtered Monkeypox dataset from May 2022 that anchors the historical side of the dashboard.","marker":"[27]"},{"why":"Provides the research-oriented collection tool used to gather the 3,000 recent 2024 tweets integrated into the dashboard.","marker":"[28]"},{"why":"Supplies the prior mpox misinformation verification tool that this dashboard positions itself against as a researcher-focused alternative.","marker":"[9]"},{"why":"Provides the real-time pandemic dashboard design model that the authors say shaped their interface.","marker":"[25]"},{"why":"Documents negative sentiment in mpox and COVID-19 tweets, the prior finding that the cynicism result extends.","marker":"[13]"},{"why":"Establishes that misinformative mpox posts circulated before official emergency declarations, motivating the misinformation-monitoring goal.","marker":"[12]"},{"why":"Demonstrates the precedent of a sentiment dashboard for a health crisis that this project adapts.","marker":"[22]"}],"fun_headline_variants":["Mpox tweet volume spiked post-CDC alert; cynicism now dominant","Dashboard reveals mpox tweets surged, distrust rose after CDC alert","Post-CDC mpox tweets surge, with cynicism overtaking discourse","Mpox dashboard: tweet surge, rising distrust after 2024 CDC alert","CDC alert linked to mpox tweet spike and surge in cynicism"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central finding that cynicism dominates recent mpox discussion rests entirely on the unnamed topic-and-sentiment labeling algorithm, whose accuracy is never tested, so if that algorithm mislabels tweets the main result collapses.","fun_headline_variants_meta":{"raw":{"variants":["Mpox tweet volume spiked post-CDC alert; cynicism now dominant","Dashboard reveals mpox tweets surged, distrust rose after CDC alert","Post-CDC mpox tweets surge, with cynicism overtaking discourse","Mpox dashboard: tweet surge, rising distrust after 2024 CDC alert","CDC alert linked to mpox tweet spike and surge in cynicism"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000197,"raw_usage":{"total_tokens":1294,"prompt_tokens":804,"completion_tokens":490,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":420,"completion_tokens_details":{"reasoning_tokens":395}},"tokens_in":420,"tokens_out":490,"duration_ms":4651,"temperature":1.0,"reasoning_tokens":395,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:50:00.549382+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the dashboard's 2024 mpox tweets, have independent human annotators label each tweet's sentiment and topic without seeing the dashboard's cluster labels, and compare the two distributions; if the human-labeled data does not show cynicism as the dominant category, or shows the algorithm systematically mislabels tweets, the central claim fails.","supporting_citations":[{"cited_title":"A Twitter dataset for Monkeypox, May 2022","cited_arxiv_id":null,"evidence_quote":"Supplies the hydrated, pre-filtered Monkeypox dataset from May 2022 that anchors the historical side of the dashboard."},{"cited_title":"Zeeschuimer","cited_arxiv_id":null,"evidence_quote":"Provides the research-oriented collection tool used to gather the 3,000 recent 2024 tweets integrated into the dashboard."},{"cited_title":"PoxVerifi: An Information Verification System to Combat Monkeypox Misinformation","cited_arxiv_id":"2209.09300","evidence_quote":"Supplies the prior mpox misinformation verification tool that this dashboard positions itself against as a researcher-focused alternative."},{"cited_title":"The Johns Hopkins University Center for Systems Science and Engineering COVID-19 Dashboard: data collection process, challenges faced, and lessons learned","cited_arxiv_id":null,"evidence_quote":"Provides the real-time pandemic dashboard design model that the authors say shaped their interface."},{"cited_title":"Sentiment analysis and text analysis of the public discourse on Twitter about COVID-19 and MPox","cited_arxiv_id":null,"evidence_quote":"Documents negative sentiment in mpox and COVID-19 tweets, the prior finding that the cynicism result extends."},{"cited_title":"Misinformation and Public Health Messaging in the Early Stages of the Mpox Outbreak: Mapping the Twitter Narrative with Deep Learning (Preprint)","cited_arxiv_id":null,"evidence_quote":"Establishes that misinformative mpox posts circulated before official emergency declarations, motivating the misinformation-monitoring goal."},{"cited_title":"Dashboard of sentiment in Austrian social media during COVID-19","cited_arxiv_id":"2006.11158","evidence_quote":"Demonstrates the precedent of a sentiment dashboard for a health crisis that this project adapts."}],"review_version":1}