{"id":"e5a186f4-1f83-4205-b1de-9fe35d0731e4","arxiv_id":"2504.18081","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"The paper interprets sentiment and emotion curves from tweets as evidence that generative AI adoption follows the Gartner Hype Cycle and Kübler-Ross Change Curve.","lead":"Analysis of 100 days of tweets about ChatGPT and similar tools claims public sentiment followed the Gartner Hype Cycle and emotions followed the Kübler-Ross Change Curve. A generalist might read it to see how social media data are used to test popular models of technology hype and adoption.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Kübler-Ross validation is circular: stage labels are assigned after seeing which month each emotion peaks, so the observed emotional sequence is constructed rather than independently tested.","rationale":"The reader's weakest-assumption analysis correctly identifies the post hoc mapping in Table 1 as the central flaw in the Kübler-Ross validation. The paper's own methodology section describes mapping emotions onto seven stages but does not state any rule for doing so before the results are reported; the mapping appears only in Table 1, after the month-of-peak data have been inspected. That makes the observed emotional sequence unfalsifiable: any arrangement of peaks could be relabeled to match the target curve by choosing which emotions belong to each stage. The overlap between adjacent stages (shock/denial both peaking in month 1, and sadness/decision both peaking in month 3) further undermines the claim of a clean sequential emotional journey. I agree that this concern is load-bearing because Hypothesis 2 is one of the two pillars of the paper's central claim. The sentiment analysis for Hypothesis 1 is also too weak to carry the claim independently: the entire observed range is positive, the 'trough' is defined after seeing the curve, and no statistical significance tests are reported. Thus the verdict should remain REJECT. I see no reason to adjust the reader's verdict, which is why verdict_should_be is UNCHANGED.","tokens_in":12828,"tokens_out":3328,"duration_ms":37789,"concrete_test":"Run a pre-registered replication on the same or an equivalent tweet stream: before inspecting monthly peaks, define an explicit dictionary from EmoRoBERTa's 28 labels to Kübler-Ross stages using only the stage definitions in Section 2.2 (e.g., surprise and joy should not be assigned to shock, and relief should not be assigned to denial). Then test whether the median peak month of each stage's assigned emotion set increases monotonically across stages and whether adjacent stages' peak-month distributions are separable, using bootstrap confidence intervals or a permutation test. If the pre-registered mapping produces overlapping or non-monotone stage peaks, Table 1's post hoc assignment is the source of the pattern and Hypothesis 2 is not validated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that Hypothesis 2 is 'effectively validated' rests on Table 1, where each of the 28 emotion categories is assigned to a Kübler-Ross stage only after its month of highest score is known. The paper never specifies an a priori mapping from EmoRoBERTa labels to the stage definitions in Section 2.2; instead, the mapping is chosen post hoc from the data. Consequently, the claimed progression from shock/denial (month 1) through frustration (month 2), sadness/decision (month 3), experimentation (month 4), and integration (months 4–5) is true by construction and cannot serve as independent evidence for the Kübler-Ross order. The overlap in the authors' own mapping makes this worse: shock and denial both peak in month 1, and sadness and decision both peak in month 3, so even the constructed sequence is not the clean monotone emotional journey described in the discussion. The sentiment-based test of Hypothesis 1 is similarly unsecured: all sentiment values remain positive, the 'trough of disillusionment' is inferred from a modest dip without confidence intervals or pre-registered thresholds, and no code or data are provided for independent verification. Because the Kübler-Ross validation is the paper's distinctive theoretical contribution and it is circular, the strongest claim of the paper is not supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper analyzes tweets about ChatGPT, Bing AI, and Microsoft Office Copilot collected over the first 100 days after the public release of ChatGPT. The authors compute daily sentiment scores with VADER and daily emotion scores with EmoRoBERTa, then interpret the sentiment curve as evidence for the Gartner Hype Cycle and the month-by-month peak order of 28 emotions as evidence for the Kübler-Ross Change Curve. The paper concludes that both hypotheses are 'effectively validated.'","tokens_in":13083,"tokens_out":4166,"duration_ms":41419,"significance":"If the empirical link were established with rigorous methods, this study would be a valuable longitudinal complement to snapshot sentiment analyses and an interesting extension of the Gartner/Kübler-Ross frameworks to generative AI. The shift from information-seeking to content-creating users is a worthwhile framing. However, the current evidence is entirely descriptive: no inferential statistics, no pre-registered thresholds, and—most importantly—the Kübler-Ross stage mapping is constructed after inspecting the data, so the central validation claim cannot be accepted on the evidence presented.","major_comments":[{"comment":"The assignment of the 28 emotions to Kübler-Ross stages is post hoc: each emotion is placed in the stage corresponding to its observed month of highest score, and this same ordering is then presented as evidence for the Kübler-Ross sequence. Because the stage labels are derived from the data, the observed progression from shock/denial to integration is true by construction and cannot independently validate Hypothesis 2; the paper needs an a priori mapping from EmoRoBERTa labels to stages or a separate confirmatory test. The overlap in the authors' own table (shock and denial both peak in month 1; sadness and decision both peak in month 3) further weakens the claimed monotone progression.","section":"Section 4.2, Table 1"},{"comment":"The Gartner Hype Cycle interpretation is not supported by quantitative evidence: all reported sentiment scores are positive, the 'trough' is a drop from about 0.37 to 0.24 with no confidence intervals, no error bars, and no statistical test against a null or an alternative trajectory. The paper should report uncertainty, specify pre-registered thresholds for stage boundaries, and test whether the observed curve fits a hype-cycle shape better than a flat or monotone trend.","section":"Section 4.1, Figure 3"},{"comment":"The methodology omits essential reproducibility details: no dataset size, collection dates, exact search queries, language filters, deduplication procedure, tweet-level or aggregation-level unit of analysis, or details of how EmoRoBERTa scores were aggregated. Without these, and without code or data, independent verification is impossible. This also affects the validity of the slope values reported in Table 1.","section":"Section 3"},{"comment":"The statement that both hypotheses are 'effectively validated' is not supported by the preceding analyses, which are descriptive and contain no inferential statistics, significance tests, or effect sizes. The paper should temper its conclusions to 'consistent with' at most, or provide proper hypothesis tests.","section":"Section 5"}],"minor_comments":[{"comment":"The emotion numbering is inconsistent and the rows are not ordered by number, which makes the table hard to read.","section":"Section 4.2, Table 1"},{"comment":"The x-axis runs from -15 to 125 and the axes have no labels; the figure should be made self-contained.","section":"Figure 3"},{"comment":"The sentiment tool is called 'VandeSentiment' but the reference is to VADER (Hutto & Gilbert, 2014); correct the name and capitalization.","section":"Section 3"},{"comment":"The abstract and method list experimentation and decision as Kübler-Ross stages, but the standard Kübler-Ross model has five stages; clarify that a seven-stage adaptation is being used.","section":"Section 2.2"},{"comment":"Figure 2 is referenced but not described or explained in the text; add a caption and discussion.","section":"Figure 2"}],"recommendation":"reject","confidential_remarks":"The manuscript fits the journal's scope, but the fundamental problem is the post hoc construction of the Kübler-Ross mapping; this is not a presentational issue and cannot be fixed by local edits. The paper also lacks data/code availability statements. If the authors can provide the dataset and a pre-registered, independent mapping, a substantially revised version might be worth reconsidering."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Thank you for the report. Here's my read.\n\nThe paper's central claim—that GenAI adoption follows the Gartner Hype Cycle and Kübler-Ross curve—is not supported by the evidence as presented. The main problem is circularity: Table 1 assigns each of 28 emotions to a Kübler-Ross stage based on the month in which the emotion peaked, and then the same peak ordering is cited as confirmation of the theoretical sequence. That's not validation; it's relabeling. The overlap (shock and denial both peak in month 1; sadness and decision both in month 3) makes even the constructed sequence messier than the narrative admits.\n\nThe sentiment analysis has a similar issue. Scores stay positive throughout—0.37 at the peak, 0.24 at the \"trough,\" 0.27 at day 100. Calling a 0.13 dip a 'trough of disillusionment' is a stretch, especially with no confidence intervals or pre-registered thresholds. No code or data are provided, so the numbers can't be checked.\n\nTo give credit where it's due: the paper is clearly written and the empirical setup is reasonable for a descriptive case study. Collecting fine-grained emotion data over the first 100 days and focusing on content creators rather than information seekers is a legitimate angle. The literature review correctly notes that prior work combined Kübler-Ross and the Hype Cycle (Williams & Braddock, 2019), so the authors aren't claiming that combination as new, though they do claim empirical validation that doesn't hold up.\n\nThe descriptive patterns might be of interest to practitioners tracking public reaction to GenAI, but as a theoretical validation the paper doesn't work. The circular mapping is a load-bearing flaw, not a minor one. If it goes to peer review, it would need an a priori emotion-to-stage mapping, statistical tests, and data/code availability. But I'd lean toward desk reject rather than spending referee time, because the core inference is unfixable without redesigning the study.\n\nOverall: not a serious thinker on its own terms, because the validation is constructed after the fact. I wouldn't cite it. Maybe bring it to reading group as a cautionary example.","headline":"A well-written case study undone by a post hoc emotion-to-stage mapping that makes the Kübler-Ross validation circular and a sentiment curve that never goes negative.","tokens_in":13604,"tokens_out":2943,"would_cite":false,"duration_ms":26761,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that generative AI adoption follows both the Gartner Hype Cycle and the Kübler-Ross Change Curve, based on tweet sentiment and emotions over the first 100 days after ChatGPT's release.","keywords":["generative AI","Gartner Hype Cycle","Kübler-Ross Change Curve","sentiment analysis","emotion analysis","Twitter data","technology adoption","ChatGPT"],"falsifier":"A direct test would pre-register the mapping between emotions and Kübler-Ross stages before data analysis, then apply the same pipeline to tweets about a later generative-AI release and check whether the predicted stage order appears. The claim would be undermined if, for example, surprise peaked after anger, or if the sentiment trough did not align with the peak of fear and confusion; more pointedly, since several emotions assigned to 'shock' and 'denial' both peak in month 1 and 'sadness' and 'decision' both peak in month 3, a re-analysis that reassigns those emotions to different stages would show how much of the Kübler-Ross fit is an artifact of the chosen mapping.","tokens_in":12584,"feed_emoji":"📈","tokens_out":7073,"duration_ms":57074,"temperature":0.7,"pith_summary":"The paper claims that public reactions to generative AI, as seen in tweets from the first 100 days after ChatGPT's release, follow two classic models at once: the Gartner Hype Cycle, a five-stage model of technology expectations, and the Kübler-Ross Change Curve, a seven-stage model of emotional response to change. Using daily sentiment scores and fine-grained emotion classifications, the author reports that enthusiasm peaked around day 30, fell into disappointment, then stabilized at moderate optimism, matching the hype cycle. On the emotional side, early surprise, fear, and denial gave way to frustration, sadness, and finally neutrality and acceptance, matching the change curve. If right, the paper offers a compact, real-world demonstration that both models apply to generative AI and provides a timeline for businesses and policymakers to manage expectations and support adoption.","feed_headline":"Generative AI's first 100 days follow hype and grief curves","feed_subtitle":"Twitter sentiment peaks, crashes, stabilizes; emotions move from shock to acceptance, tracking two classic models.","key_machinery":"The central machinery is the combination of two scoring pipelines applied to timestamped tweets. Sentiment is measured with VADER's compound score to produce a daily sentiment curve that is compared stage-by-stage to the Gartner Hype Cycle. Emotion is measured with EmoRoBERTa, which assigns each tweet to 28 fine-grained emotion categories; the paper assigns those emotions to seven Kübler-Ross stages by the month in which each emotion's score is highest (Table 1), then plots daily emotion scores as curves. The load-bearing step is that mapping: a post hoc assignment of, for example, surprise and joy to 'shock,' disgust and fear to 'denial,' and annoyance and confusion to 'frustration,' which turns independent emotion scores into a Kübler-Ross trajectory.","core_discovery":"On the paper's own terms, the central discovery is that the movement of sentiment scores closely follows the pattern predicted by the Gartner Hype Cycle, and that emotional reactions evolved from initial shock and denial to eventual acceptance and integration, so both stated hypotheses are effectively validated. The sentiment curve is bell-shaped: it rises to a peak of inflated expectations around day 30, dips to a trough of disillusionment near 0.24, then climbs to a plateau near 0.27, never turning strongly negative. The emotion curves, mapped from 28 categories to seven Kübler-Ross stages, show early peaks for surprise, joy, disgust, fear, and relief, followed by anger and confusion, then sadness and disappointment, and finally gratitude, curiosity, neutrality, and approval. The author interprets this as evidence that generative AI adoption is a dual-stage process, cognitive and emotional, and argues that the plateau of productivity sits higher than the starting point, reflecting genuine integration.","pith_inferences":["The paper's post hoc emotion-to-stage mapping is descriptive, not predictive; a stronger test would pre-register the mapping and apply it to a later AI release, such as a new model generation, to see if the stage order repeats.","The large positive slope for 'neutral' (2.22) suggests the main late-stage signal is not growing enthusiasm but growing habituation—users stop feeling strongly about the tool as it becomes background infrastructure, which is consistent with the plateau of productivity but also means the curve may be measuring attention fatigue, not satisfaction.","The dissociation between positive sentiment (never negative) and negative emotions (fear, disgust rising) hints that people can evaluate a tool as useful while still being affectively uneasy; future work could separate evaluative approval from emotional acceptance as distinct adoption barriers.","The 100-day window ends as the plateau begins; extending the data to a full year would reveal whether the plateau holds or whether a second hype cycle appears with new model releases, a test the paper's design leaves open."],"forward_implications":["Businesses and policymakers gain a predictable timeline: enthusiasm peaks around day 30, criticism troughs near day 60, and moderate optimism stabilizes by day 100.","Change-management teams can treat early resistance as expected and design training and communication for the frustration and sadness phases rather than treating them as product failures.","Because sentiment never turns strongly negative, the paper implies that generative AI is being integrated cautiously but successfully, so continued investment and rollout are justifiable."],"supporting_citations":[{"why":"Defines and visually presents the Gartner Hype Cycle that Hypothesis 1 is derived from.","marker":"(Fenn & Raskino, 2008)"},{"why":"Proposes superimposing the Kübler-Ross Change Curve on the Hype Cycle, the dual-model idea this study tests.","marker":"(Williams, 2019)"},{"why":"Supplies VADER, the sentiment scoring tool producing the compound scores used in Section 4.1.","marker":"(Hutto & Gilbert, 2014)"},{"why":"Introduces EmoRoBERTa, the emotion classifier providing the 28 emotion categories used in Section 4.2.","marker":"(Kamath et al., 2022)"},{"why":"Provides GoEmotions, the fine-grained emotion dataset that enables 28-category classification.","marker":"(Demszky et al., 2020)"},{"why":"Analyzes 151 technologies on the Gartner Hype Cycle and notes generative AI was not included, motivating the gap this study fills.","marker":"(Kaivo-Oja, 2020)"}],"fun_headline_variants":["Sentiment peaks and dips map AI hype and grief curves","Tweets track Gartner hype and Kübler-Ross grief for AI","Generative AI adoption mirrors hype cycle and grief stages","AI sentiment follows classic hype and acceptance curves","From shock to acceptance: AI tweets follow two classic models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is the post hoc mapping of 28 detected emotions to the seven Kübler-Ross stages based on the month in which each emotion peaked, a choice made after inspecting the data rather than derived from the change-curve literature; if that mapping is not valid, the emotion analysis provides no independent evidence for the Kübler-Ross order.","fun_headline_variants_meta":{"raw":{"variants":["Sentiment peaks and dips map AI hype and grief curves","Tweets track Gartner hype and Kübler-Ross grief for AI","Generative AI adoption mirrors hype cycle and grief stages","AI sentiment follows classic hype and acceptance curves","From shock to acceptance: AI tweets follow two classic models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000703,"raw_usage":{"total_tokens":3204,"prompt_tokens":1009,"completion_tokens":2195,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":625,"completion_tokens_details":{"reasoning_tokens":2114}},"tokens_in":625,"tokens_out":2195,"duration_ms":15418,"temperature":1.0,"reasoning_tokens":2114,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:24:53.588733+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would pre-register the mapping between emotions and Kübler-Ross stages before data analysis, then apply the same pipeline to tweets about a later generative-AI release and check whether the predicted stage order appears. The claim would be undermined if, for example, surprise peaked after anger, or if the sentiment trough did not align with the peak of fear and confusion; more pointedly, since several emotions assigned to 'shock' and 'denial' both peak in month 1 and 'sadness' and 'decision' both peak in month 3, a re-analysis that reassigns those emotions to different stages would show how much of the Kübler-Ross fit is an artifact of the chosen mapping.","supporting_citations":[],"review_version":1}