Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

The exception of humour: Iconicity, Phonemic Surprisal, Memory Recall, and Emotional Associations

T0 review · 3 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Humor is an exception to the negativity-bias pattern in word memory: although humorous words carry positive valence, they also carry higher phonemic surprisal and are recalled more accurately, according to this meta-analysis of American…

desk verdict A genuinely new descriptive pattern—humorous words are more surprising and more memorable—but the 'exception' claim needs a joint model with valence before it can be believed. read the letter →

arxiv 2502.01682 v1 pith:MA74NI3J submitted 2025-02-02 cs.CL

classification cs.CL
keywords humorphonemicsurprisalemotionalvalencememoryrecalliconicitynegativitybiasphonologicalmarkednessmeta-study
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether humorous words break the usual emotional-memory pattern. Prior work shows negative words are phonologically surprising and easy to recall, while positive words are not; humor is rated positive, yet the paper finds humorous words are also high in phonemic surprisal and recall accuracy. The authors combine existing norming and recall datasets for American English and run regressions with surprisal as either outcome or predictor. A sympathetic reader would care because the result suggests humor recruits the same attention-and-memory machinery as negative stimuli, and gives phonology a concrete role in humor and memorability.

What carries the argument

The load-bearing instrument is phonemic bigram surprisal, defined as $-\log_2 P(\text{phoneme}_i \mid \text{phoneme}_{i-1})$ averaged over all consecutive phoneme pairs of a word; it turns unpredictable sound sequences into a bit-count measure. The paper pairs this with Likert humor norms, two valence datasets, iconicity ratings, and recall accuracy from prior experiments, then runs two series of multiple regressions: one with emotions, valence, and humor as outcomes predicted by surprisal plus controls, and one with recall as the outcome predicted by emotions, valence, or humor plus surprisal. The machinery works by showing that humor's coefficient is positive and significant in both directions—higher surprisal and higher recall—while valence measures still show the usual negative-emotion pattern.

What would settle it

Run one statistical model on the same word data that predicts how surprising a word sounds from its humor rating, its negative-versus-positive rating, and the combination of the two. If humor no longer predicts surprisal once negativity is accounted for, or if the combination term is not significant, the claimed exception fails; a parallel recall model would test the memory side.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central claim is that humor is an exception to the negativity-bias pattern in word memory. Humor ratings positively correlate with average phonemic bigram surprisal and with iconicity (Table 3), and humor positively predicts recall accuracy even with surprisal and other word features in the model (Table 6). Because humor is also shown to correlate with positive valence, the paper concludes that humorous words behave like negatively valenced words—surprising and memorable—despite being positive, suggesting that humor may exploit cognitive mechanisms similar to negativity while resolving into a positive social response. The discussion ties this to incongruity-resolution and to structural markedness as a shared basis for funniness and iconicity.

Load-bearing premise

The conclusion that humor is exceptional depends on putting separate findings together—humor predicts surprisal, negative emotion predicts surprisal, and humor predicts recall—without running one analysis that checks whether humor's effect is genuinely different from negativity's effect in the same model.

Editorial extensions

If this is right

  • Humor becomes a documented counterexample to the general rule that positive valence goes with lower surprisal.
  • Humorous words' memorability holds with average surprisal in the model, so humor and sound unpredictability each contribute to recall.
  • Phonemic surprisal can serve as an objective, quantitative proxy for phonological markedness in humor and iconicity research.
  • The study's cross-linguistic prediction can be tested directly: if non-English humor norms show the same surprisal and recall pattern, the effect is likely cognitive rather than language-specific.
  • Distinguishing humor types (colloquial iconic words, situational farce, dark humor) may reveal which subtypes drive the surprisal effect.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the humor–surprisal correlation is causal, then deliberately choosing phonologically rare sound sequences could make humorous content stickier; the paper only establishes correlation, not direction.
  • A direct extension would rank humor subtypes by expected memorability: dark humor and sarcasm combine negative valence with high surprisal, so they should be remembered best of all.
  • Because surprisal is computed from phoneme transition probabilities without semantic input, it could be added as a feature in humor-recognition systems; whether it helps is an untested engineering question.
  • Replicating with non-English norming data would show whether the humor exception is a universal cognitive signature or an artifact of American English sound patterns.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This meta-study combines several existing English-language datasets to test whether humorous words, despite their positive emotional valence, are marked by higher phonemic bigram surprisal and better memory recall than non-humorous words. The authors compute average bigram surprisal from SUBTLEX-US and CMU pronunciations, then run multiple linear regressions with emotion variables (NRC lexicon, Glasgow/NRC valence, Engelthaler & Hills humor ratings) as either dependent or independent variables and Cortese et al. recognition-memory accuracy as the recall outcome. The reported tables show, in separate models, that negative valence is associated with higher surprisal and better recall, and that humor is also positively associated with surprisal and recall. The paper's headline claim is that humor is an exception: positively valenced yet high in surprisal and memorability, suggesting that humor shares cognitive mechanisms with negative stimuli.

Significance. If the central claim were established, the paper would make a useful empirical contribution to the emerging literature on iconicity, phonological markedness, affective norms, and word memorability. It draws on publicly available datasets, and the underlying pipeline (surprisal computation, regression structure) is transparent enough to be reproduced or extended. The paper also connects its findings to a concrete theoretical proposal (Suls's two-stage model, Dingemanse & Thompson on playful iconicity), which gives the result interpretive value beyond a purely descriptive correlation. However, the key 'exception' claim is currently an interpretive leap from separate marginal associations; the evidence as reported does not yet support it. The manuscript's strengths are its data reuse and transparency, but its central inference requires additional model specification before the conclusions can be accepted.

major comments (3)
  1. [Section 3 (Tables 3 and 6)] The central 'exception' claim is not tested in any joint model. Table 3 shows Humor positively predicting Average_Surprisal in a model without valence covariates, and Table 6 shows Humor positively predicting recall in a model without valence covariates. The first paragraph of Section 3 establishes that Humor is positively correlated with positive valence. Because positive valence is negatively associated with surprisal (Table 1) and with recall (Table 4), a valence-mediated or valence-confounded account of the humor effects is plausible. The manuscript needs either (a) a model containing Humor and NRC_Valence or G_Valence simultaneously, with the Humor coefficient reported, or (b) an explicit Humor x Valence interaction test. Without this, the headline conclusion that humor is 'an exception' is unsupported, and the Discussion's claim that 'humor follows the same patterns as negative stimuli' overstates what the reported regressions show.
  2. [Section 3, first paragraph] The two simple regressions of Humor on valence are described only by p-values ('p < 0.001 in both models'), with no coefficients, standard errors, or fit statistics. The claim that humor is 'stochastically positive' is load-bearing for the exception narrative, because the whole argument depends on the strength and direction of the humor-valence association. A weak or noisy association would leave room for substantial residual negative valence in the humorous word set, which would undermine the 'positive valence but high surprisal' framing. Please report the full regression output, including effect sizes and ideally the distribution of valence ratings among high-humor words.
  3. [Section 4] The interpretive statement that 'humor follows the same patterns as negative stimuli' is not justified by the analysis. The manuscript compares the sign and significance of Humor coefficients in Tables 3 and 6 with the sign and significance of Negative/Valence coefficients in Tables 1, 2, 4, and 5, but these models use different response scales (binary NRC emotions, 0-7 valence, 1-5 humor), different covariate sets, and no common metric for effect comparison. To support 'same patterns', the authors should either standardize coefficients, fit a common model containing both humor and valence, or provide a formal equivalence/contrast test. As written, the comparison is not statistically grounded.
minor comments (6)
  1. [References] Engelthaler and Hills is cited as 2017 in the Introduction but as 2018 in Methods and the reference list; please make the year consistent.
  2. [Table 4] In the NRC_Valence column, the PoS_Interjection row appears to contain a stray '0.371' on a separate line rather than in the table cell, making the coefficient placement unclear.
  3. [Table 5] Table 5 omits rows for PoS_Determiner, PoS_Preposition, PoS_Pronoun, and PoS_Unclassified in all ten models, whereas Table 2 includes these categories. The authors should state whether these categories were dropped because of collinearity or reference-level coding, or whether this is an omission.
  4. [Section 2 and all tables] The manuscript says there are 'two series' of regression models, but the results section actually reports four families of models (valence-surprisal, emotion-surprisal, valence-recall, emotion-recall) plus the humor models. Please adjust the wording and report the number of observations and R-squared for each model.
  5. [Data availability] The data link is given as a short URL (shorturl.at/2SXvO), which is hard to verify and not archival. A stable repository DOI or a permanent data citation would be preferable.
  6. [Throughout] There are several small language errors, e.g., 'valanced' for 'valenced' in Section 4, and the use of 'Surprisal' as a variable name where the NRC lexicon's emotion is 'Surprise'. A careful proofreading pass would improve clarity.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation; the exception claim is an interpretive extrapolation from separate regressions, not a reduction to the paper's inputs.

full rationale

The study's central result—humor correlates with positive valence yet with higher phonemic surprisal and better recall—is computed directly from public datasets (Engelthaler and Hills 2018; NRC; Glasgow; Cortese et al. 2010; SUBTLEX/CMU) in new regressions reported in Tables 1-6. The 'exception' conclusion is obtained by comparing separate marginal models (negative valence predicts surprisal/recall; humor predicts surprisal/recall), without a joint interaction model; this is a statistical-inference limitation, not circularity, because no coefficient is reused as an output and no equation is definitionally equivalent to its input. The paper cites the authors' own work (Kilpatrick, Under Review; Flaksman and Kilpatrick, In Press; Kilpatrick and Bundgaard-Nielsen 2024) for background associations between negativity, iconicity, and surprisal, but those associations are independently estimated in the present Tables 1, 2, 4, and 5, so the self-citations are not load-bearing. The humor coefficients in Tables 3 and 6 are new products of the current dataset and do not reduce to the cited claims. Thus there is no self-definitional, fitted-input, or citation-forced circularity. Score 2 reflects the presence of several self-citations in the framing, not a circular derivation.

Assumptions & free parameters 0 free parameters · 6 assumptions · 0 invented entities

The paper introduces no new entities or fitted constants; the load-bearing assumptions are the validity of the corpus-derived surprisal measure, the representativeness of the merged norm datasets, and the correctness of the authors' prior unpublished findings on surprisal, negativity, and memory.

assumptions (6)
  • standard math Surprisal is defined as -log2 P(bigram), with probabilities estimated from SUBLEX-US corpus frequencies.
    Shannon information theory; the corpus is treated as the true distribution over phoneme bigrams in American English.
  • domain assumption The Engelthaler and Hills humor ratings capture a one-dimensional humor construct.
    Acknowledged in Section 1; the study uses this single humor score without distinguishing humor types.
  • domain assumption Negative valence is associated with higher surprisal, as claimed in Kilpatrick (Under Review).
    This premise is loaded from an unpublished, self-cited manuscript and is not independently verifiable in the present paper.
  • domain assumption Higher phonemic surprisal enhances memory recall, from Kilpatrick and Bundgaard-Nielsen (2024).
    The surprisal-memory link is imported from the first author's prior work and is central to interpreting the recall regressions.
  • standard math Linear regression with part-of-speech controls is an adequate model for the relationships.
    The paper uses multiple linear regression throughout; this assumes linearity and that included controls are sufficient.
  • domain assumption The merged norm datasets are compatible and missing data is missing at random.
    The paper states no samples were excluded except for missing data, implying the merged set is representative.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The exception of humour: Iconicity, Phonemic Surprisal, Memory Recall, and Emotional Associations." pith.science (2026). https://pith.science/paper/MA74NI3J

@misc{pith2026250201682,
  author       = {Pith},
  title        = {Pith review of: The exception of humour: Iconicity, Phonemic Surprisal, Memory Recall, and Emotional Associations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MA74NI3J}},
  note         = {Machine review of arXiv:2502.01682}
}
read the original abstract

This meta-study explores the relationships between humor, phonemic bigram surprisal, emotional valence, and memory recall. Prior research indicates that words with higher phonemic surprisal are more readily remembered, suggesting that unpredictable phoneme sequences promote long-term memory recall. Emotional valence is another well-documented factor influencing memory, with negative experiences and stimuli typically being remembered more easily than positive ones. Building on existing findings, this study highlights that words with negative associations often exhibit greater surprisal and are easier to recall. Humor, however, presents an exception: while associated with positive emotions, humorous words also display heightened surprisal and enhanced memorability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Spectral Rewiring for Exploration, Purification, and Model Merging

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Subspace-Aligned Rewiring projects RL weight updates onto the base model’s SVD basis, retaining a compact rewiring matrix that preserves reasoning and improves exploration and multi-domain merging.

Reference graph

Works this paper leans on

7 extracted references · 7 canonical work pages · cited by 1 Pith paper

  1. [7]

    John Benjamins Publishing Company

    Ideophones. John Benjamins Publishing Company. https://doi.org/10.1075/tsl.45. Weide, R

  2. [2001]

    Personality and Social Psychology Review, 5(4), 296 -320

    Negativity bias, negativity dominance, and contagion. Personality and Social Psychology Review, 5(4), 296 -320. https://doi.org/10.1207/S15327957PSPR0504_2. Sánchez-Gutiérrez, C. H., et al

  3. [2003]

    Nature Reviews Neuroscience, 4(3):193 -202

    Neural mechanisms for detecting and remembering novel events. Nature Reviews Neuroscience, 4(3):193 -202. https://doi.org/10.1038/nrn1052. Rozin, P., & Royzman, E. B

  4. [2004]

    Current Directions in Psychological Science, 13(3), 102 -105

    Human emotion and memory: Interactions of the amygdala and hippocampal complex. Current Directions in Psychological Science, 13(3), 102 -105. https://doi.org/10.1111/j.0963-7214.2004.00293.x. Ranganath, C., & Rainer, G

  5. [2006]

    Nature Reviews Neuroscience, 7(1), 54 -64

    Cognitive neuroscience of emotional memory. Nature Reviews Neuroscience, 7(1), 54 -64. https://doi.org/10.1038/nrn1825. Mohammad, S. M., & Turney, P. D

  6. [2007]

    Current Directions in Psychological Science, 16(4):213-218

    Negative emotion enhances memory accuracy: Behavioral and neuroimaging evidence. Current Directions in Psychological Science, 16(4):213-218. Kilpatrick, A. (2023). Sound Symbolism in Automatic Emotion Recognition and Sentiment Analysis. In P. L. Villagrá, & X. Li (Eds.), Proceedings of the International Workshop on Cognitive AI 2023 co - located with the ...

  7. [2018]

    In Proceedings of the 2018 conference of the North American chapter of the association for computational linguistics: Human language technologies, 2, 113-117

    Humor recognition using deep learning. In Proceedings of the 2018 conference of the North American chapter of the association for computational linguistics: Human language technologies, 2, 113-117. Cortese, M. J., Khanna, M. M., & Hacker, S

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.