{"id":"5ed86793-049d-4bea-a553-5aaf55bcc6a6","arxiv_id":"2502.01682","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Humorous English words have higher phonemic bigram surprisal and better recall accuracy despite being positively valenced.","lead":"This meta-study combines existing word-rating datasets to show that humorous words are phonetically more surprising and are recalled more accurately than typical words. The finding is presented as an exception to the usual pattern where negative words stand out, suggesting humor may engage similar cognitive mechanisms.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'exception' claim rests on separate regressions; without a joint model including valence, humor's surprisal and recall effects could be confounded by valence.","rationale":"The reader's weakest assumption correctly identifies the load-bearing gap: the exception claim requires a direct statistical comparison of humor and valence effects within the same model, which is absent. My independent review of Tables 1, 3, 4, and 6 confirms that each relation is estimated separately, and Section 4 draws the exception conclusion without an interaction or mediation test. This is the single most important threat to correctness because if humor's effects on surprisal and recall are mediated or moderated by valence, the central finding reduces to a restatement of valence effects. The proposed test is feasible with the public dataset and would settle the issue. No additional concerns, such as multiple-comparison corrections or effect-size reporting, are as load-bearing as this missing joint model. The reader's conditional verdict is appropriate; the paper should not be accepted as establishing the exception claim without the additional analysis.","tokens_in":10350,"tokens_out":1810,"duration_ms":22548,"concrete_test":"Re-run the humor-surprisal model (Table 3) with NRC_Valence and G_Valence included as covariates, plus the Humor × Valence interaction term. Then re-run the humor-recall model (Table 6) likewise with valence covariates and the Humor × Valence interaction. If the Humor coefficient on surprisal remains significant and the interaction is significant (e.g., humor's surprisal slope differs by valence), the exception claim survives; if the Humor coefficient attenuates to non-significance or the interaction is null, the claim should be reframed as valence-driven rather than humor-specific.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that humor is exceptional: positively valenced yet high in phonemic surprisal and memory recall. This is supported only by comparing separate models: Table 1 shows negative valence predicts higher surprisal; Table 3 shows humor predicts higher surprisal; Table 4 shows negative valence predicts higher recall; Table 6 shows humor predicts higher recall. Because humor is positively correlated with positive valence (Section 3), the exception claim requires demonstrating that humor's positive associations with surprisal and recall are not explained by valence or by an interaction between humor and valence. But valence is never entered as a covariate in the humor models, and humor is never entered alongside valence in a single model. Consequently, the apparent 'exception' could arise if humorous words in the Engelthaler and Hills (2018) norms carry residual negative valence (e.g., sarcasm, dark humor), or if the humor effect is mediated by iconicity and its control does not fully remove that confound. The manuscript itself notes dark humor and sarcasm as mixtures of negativity and low probability, but does not test whether these drive the effect. This is not an internal inconsistency in the regressions, but a correctness risk: the headline conclusion is an interpretive leap from separate marginal associations to a claim about interaction and exception. The conclusion in Section 4 ('humor follows the same patterns as negative stimuli') is therefore not statistically established by the reported analyses.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This meta-study combines several existing English-language datasets to test whether humorous words, despite their positive emotional valence, are marked by higher phonemic bigram surprisal and better memory recall than non-humorous words. The authors compute average bigram surprisal from SUBTLEX-US and CMU pronunciations, then run multiple linear regressions with emotion variables (NRC lexicon, Glasgow/NRC valence, Engelthaler & Hills humor ratings) as either dependent or independent variables and Cortese et al. recognition-memory accuracy as the recall outcome. The reported tables show, in separate models, that negative valence is associated with higher surprisal and better recall, and that humor is also positively associated with surprisal and recall. The paper's headline claim is that humor is an exception: positively valenced yet high in surprisal and memorability, suggesting that humor shares cognitive mechanisms with negative stimuli.","tokens_in":10628,"tokens_out":3299,"duration_ms":37123,"significance":"If the central claim were established, the paper would make a useful empirical contribution to the emerging literature on iconicity, phonological markedness, affective norms, and word memorability. It draws on publicly available datasets, and the underlying pipeline (surprisal computation, regression structure) is transparent enough to be reproduced or extended. The paper also connects its findings to a concrete theoretical proposal (Suls's two-stage model, Dingemanse & Thompson on playful iconicity), which gives the result interpretive value beyond a purely descriptive correlation. However, the key 'exception' claim is currently an interpretive leap from separate marginal associations; the evidence as reported does not yet support it. The manuscript's strengths are its data reuse and transparency, but its central inference requires additional model specification before the conclusions can be accepted.","major_comments":[{"comment":"The central 'exception' claim is not tested in any joint model. Table 3 shows Humor positively predicting Average_Surprisal in a model without valence covariates, and Table 6 shows Humor positively predicting recall in a model without valence covariates. The first paragraph of Section 3 establishes that Humor is positively correlated with positive valence. Because positive valence is negatively associated with surprisal (Table 1) and with recall (Table 4), a valence-mediated or valence-confounded account of the humor effects is plausible. The manuscript needs either (a) a model containing Humor and NRC_Valence or G_Valence simultaneously, with the Humor coefficient reported, or (b) an explicit Humor x Valence interaction test. Without this, the headline conclusion that humor is 'an exception' is unsupported, and the Discussion's claim that 'humor follows the same patterns as negative stimuli' overstates what the reported regressions show.","section":"Section 3 (Tables 3 and 6)"},{"comment":"The two simple regressions of Humor on valence are described only by p-values ('p < 0.001 in both models'), with no coefficients, standard errors, or fit statistics. The claim that humor is 'stochastically positive' is load-bearing for the exception narrative, because the whole argument depends on the strength and direction of the humor-valence association. A weak or noisy association would leave room for substantial residual negative valence in the humorous word set, which would undermine the 'positive valence but high surprisal' framing. Please report the full regression output, including effect sizes and ideally the distribution of valence ratings among high-humor words.","section":"Section 3, first paragraph"},{"comment":"The interpretive statement that 'humor follows the same patterns as negative stimuli' is not justified by the analysis. The manuscript compares the sign and significance of Humor coefficients in Tables 3 and 6 with the sign and significance of Negative/Valence coefficients in Tables 1, 2, 4, and 5, but these models use different response scales (binary NRC emotions, 0-7 valence, 1-5 humor), different covariate sets, and no common metric for effect comparison. To support 'same patterns', the authors should either standardize coefficients, fit a common model containing both humor and valence, or provide a formal equivalence/contrast test. As written, the comparison is not statistically grounded.","section":"Section 4"}],"minor_comments":[{"comment":"Engelthaler and Hills is cited as 2017 in the Introduction but as 2018 in Methods and the reference list; please make the year consistent.","section":"References"},{"comment":"In the NRC_Valence column, the PoS_Interjection row appears to contain a stray '0.371' on a separate line rather than in the table cell, making the coefficient placement unclear.","section":"Table 4"},{"comment":"Table 5 omits rows for PoS_Determiner, PoS_Preposition, PoS_Pronoun, and PoS_Unclassified in all ten models, whereas Table 2 includes these categories. The authors should state whether these categories were dropped because of collinearity or reference-level coding, or whether this is an omission.","section":"Table 5"},{"comment":"The manuscript says there are 'two series' of regression models, but the results section actually reports four families of models (valence-surprisal, emotion-surprisal, valence-recall, emotion-recall) plus the humor models. Please adjust the wording and report the number of observations and R-squared for each model.","section":"Section 2 and all tables"},{"comment":"The data link is given as a short URL (shorturl.at/2SXvO), which is hard to verify and not archival. A stable repository DOI or a permanent data citation would be preferable.","section":"Data availability"},{"comment":"There are several small language errors, e.g., 'valanced' for 'valenced' in Section 4, and the use of 'Surprisal' as a variable name where the NRC lexicon's emotion is 'Surprise'. A careful proofreading pass would improve clarity.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper's claims are plausibly testable with the data the authors already have, and the central problem is a missing model specification rather than an unrecoverable flaw. I would encourage the editor to request a revision that includes joint humor-valence models and interaction tests, rather than rejecting the manuscript outright. One additional editorial concern is that the manuscript's framing in cs.CL is somewhat thin on the machine-learning side despite the Introduction's gesture toward sentiment analysis; the paper reads primarily as a psycholinguistics/empirical-linguistics contribution, which may affect fit with a computer science venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper does something new: it brings phonemic bigram surprisal to bear on humor norms and shows that humorous words tend to have more surprising phoneme sequences and better recall, even after controlling for iconicity, length, and word class. That is a real descriptive finding, and it is worth taking seriously. The data work is transparent—public datasets, straightforward regressions, controls included—and the writing is clear about what was measured.\n\nThe soft spot is exactly what the stress-test note says. The central 'exception' claim—humor is positively valenced but behaves like negative stimuli—is never actually tested. The paper shows in separate models that negative valence predicts surprisal and recall, and that humor predicts surprisal and recall. But humor and valence are never entered into the same equation. Without that, the exception could be an artifact: humorous words might carry residual negative associations (dark humor, sarcasm) or the surprisal effect could be mediated by iconicity, which the paper itself flags as a possible confound. The Discussion's statement that 'humor follows the same patterns as negative stimuli' is an interpretive leap from separate marginal associations to a claim about interaction. That needs a joint model—ideally with an interaction term or a mediation analysis.\n\nTwo smaller issues. The motivation leans heavily on two self-citations that are unpublished or in press (Kilpatrick, Under Review; Flaksman & Kilpatrick, In Press). That is not a fatal flaw, but it does mean the central premise cannot be checked. And the report omits effect sizes and multiple-comparison corrections across the many models, which would help the reader judge how robust these correlations really are.\n\nNone of this kills the paper. The descriptive pattern is interesting, and the analysis is honest about its controls. But the conclusion needs to be scaled back to what the current analyses support: humor is associated with surprisal and recall; whether that is an 'exception' to the valence pattern is untested.\n\nMy take: send it to a serious referee, but with a clear request for a joint model including valence and humor. The authors can probably do this with the same datasets. For readers in psycholinguistics or sentiment analysis, there is useful material here, but I would not cite the exception claim until it is tested properly.","headline":"A genuinely new descriptive pattern—humorous words are more surprising and more memorable—but the 'exception' claim needs a joint model with valence before it can be believed.","tokens_in":11118,"tokens_out":2426,"would_cite":false,"duration_ms":24623,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Humor is an exception to the negativity-bias pattern in word memory: although humorous words carry positive valence, they also carry higher phonemic surprisal and are recalled more accurately, according to this meta-analysis of American…","keywords":["humor","phonemic surprisal","emotional valence","memory recall","iconicity","negativity bias","phonological markedness","meta-study"],"falsifier":"Run one statistical model on the same word data that predicts how surprising a word sounds from its humor rating, its negative-versus-positive rating, and the combination of the two. If humor no longer predicts surprisal once negativity is accounted for, or if the combination term is not significant, the claimed exception fails; a parallel recall model would test the memory side.","tokens_in":10179,"feed_emoji":"😄","tokens_out":8417,"duration_ms":79184,"temperature":0.7,"pith_summary":"This paper asks whether humorous words break the usual emotional-memory pattern. Prior work shows negative words are phonologically surprising and easy to recall, while positive words are not; humor is rated positive, yet the paper finds humorous words are also high in phonemic surprisal and recall accuracy. The authors combine existing norming and recall datasets for American English and run regressions with surprisal as either outcome or predictor. A sympathetic reader would care because the result suggests humor recruits the same attention-and-memory machinery as negative stimuli, and gives phonology a concrete role in humor and memorability.","feed_headline":"Humor words are surprising and memorable despite positivity","feed_subtitle":"Humorous words track negative stimuli in sound unpredictability and recall, hinting at shared cognitive mechanisms.","key_machinery":"The load-bearing instrument is phonemic bigram surprisal, defined as $-\\log_2 P(\\text{phoneme}_i \\mid \\text{phoneme}_{i-1})$ averaged over all consecutive phoneme pairs of a word; it turns unpredictable sound sequences into a bit-count measure. The paper pairs this with Likert humor norms, two valence datasets, iconicity ratings, and recall accuracy from prior experiments, then runs two series of multiple regressions: one with emotions, valence, and humor as outcomes predicted by surprisal plus controls, and one with recall as the outcome predicted by emotions, valence, or humor plus surprisal. The machinery works by showing that humor's coefficient is positive and significant in both directions—higher surprisal and higher recall—while valence measures still show the usual negative-emotion pattern.","core_discovery":"On the paper's own terms, the central claim is that humor is an exception to the negativity-bias pattern in word memory. Humor ratings positively correlate with average phonemic bigram surprisal and with iconicity (Table 3), and humor positively predicts recall accuracy even with surprisal and other word features in the model (Table 6). Because humor is also shown to correlate with positive valence, the paper concludes that humorous words behave like negatively valenced words—surprising and memorable—despite being positive, suggesting that humor may exploit cognitive mechanisms similar to negativity while resolving into a positive social response. The discussion ties this to incongruity-resolution and to structural markedness as a shared basis for funniness and iconicity.","pith_inferences":["If the humor–surprisal correlation is causal, then deliberately choosing phonologically rare sound sequences could make humorous content stickier; the paper only establishes correlation, not direction.","A direct extension would rank humor subtypes by expected memorability: dark humor and sarcasm combine negative valence with high surprisal, so they should be remembered best of all.","Because surprisal is computed from phoneme transition probabilities without semantic input, it could be added as a feature in humor-recognition systems; whether it helps is an untested engineering question.","Replicating with non-English norming data would show whether the humor exception is a universal cognitive signature or an artifact of American English sound patterns."],"forward_implications":["Humor becomes a documented counterexample to the general rule that positive valence goes with lower surprisal.","Humorous words' memorability holds with average surprisal in the model, so humor and sound unpredictability each contribute to recall.","Phonemic surprisal can serve as an objective, quantitative proxy for phonological markedness in humor and iconicity research.","The study's cross-linguistic prediction can be tested directly: if non-English humor norms show the same surprisal and recall pattern, the effect is likely cognitive rather than language-specific.","Distinguishing humor types (colloquial iconic words, situational farce, dark humor) may reveal which subtypes drive the surprisal effect."],"supporting_citations":[{"why":"Supplies the Humor ratings on a 5-point Likert scale, which define which words count as humorous and serve as the central predictor of interest.","marker":"Engelthaler & Hills (2018)"},{"why":"Supplies the memory recall accuracy data from a two-session word recall experiment, the dependent variable in the recall models.","marker":"Cortese, Khanna, & Hacker (2010)"},{"why":"Provides the SUBTLEX-US word frequency norms used with phonemic transcriptions to estimate the bigram probabilities behind surprisal.","marker":"Brysbaert & New (2009)"},{"why":"Provides the CMU Pronouncing Dictionary phonemic transcriptions from which phonemic bigram surprisal is computed.","marker":"Weide (1999)"},{"why":"Supplies the iconicity ratings used as a covariate in all models and as a link between sound-meaning correspondence and surprisal.","marker":"Winter et al. (2023)"},{"why":"Supplies the NRC emotion lexicon variables, including the ten emotion categories and the NRC valence measure used as both outcomes and predictors.","marker":"Mohammad & Turney (2013)"},{"why":"Supplies the Glasgow Norms valence ratings, giving a second independent measure of emotional valence.","marker":"Scott et al. (2019)"},{"why":"Established the earlier result that high-surprisal words are harder to process but more memorable, the processing baseline this study extends.","marker":"Kilpatrick & Bundgaard-Nielsen (2024)"},{"why":"Provides the theoretical claim that structural markedness underlies both funniness and iconicity, motivating the prediction that humorous words carry more surprisal.","marker":"Dingemanse & Thompson (2020)"}],"fun_headline_variants":["Humour words are the surprising exception to negativity bias","Positive but memorable: humour breaks the negativity pattern","Why funny words are as memorable as negative ones","Humour: positive words that surprise and stick","The humour paradox: positive, surprising, and memorable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusion that humor is exceptional depends on putting separate findings together—humor predicts surprisal, negative emotion predicts surprisal, and humor predicts recall—without running one analysis that checks whether humor's effect is genuinely different from negativity's effect in the same model.","fun_headline_variants_meta":{"raw":{"variants":["Humour words are the surprising exception to negativity bias","Positive but memorable: humour breaks the negativity pattern","Why funny words are as memorable as negative ones","Humour: positive words that surprise and stick","The humour paradox: positive, surprising, and memorable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00039,"raw_usage":{"total_tokens":1979,"prompt_tokens":794,"completion_tokens":1185,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":410,"completion_tokens_details":{"reasoning_tokens":1126}},"tokens_in":410,"tokens_out":1185,"duration_ms":13084,"temperature":1.0,"reasoning_tokens":1126,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T18:07:18.121689+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run one statistical model on the same word data that predicts how surprising a word sounds from its humor rating, its negative-versus-positive rating, and the combination of the two. If humor no longer predicts surprisal once negativity is accounted for, or if the combination term is not significant, the claimed exception fails; a parallel recall model would test the memory side.","supporting_citations":[],"review_version":1}