REVIEW 4 major objections 4 minor 110 references
Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects
T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This systematic review of 40 studies argues that speech sarcasm recognition requires multimodal fusion, and its meta-analysis finds no statistically significant performance difference between attention and encoder-decoder fusion methods.
desk verdict Useful review of speech-based sarcasm recognition, but its meta-analysis of fusion methods is statistically shaky and needs correction. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The analytical engine is a systematic-review screening pipeline (five databases, two search phases, 40 studies) feeding three structured analyses: dataset comparison, feature taxonomy, and classification taxonomy adopting a four-way fusion scheme (encoder-decoder, attention, collaborative gating, quantum-based). The load-bearing quantitative machinery is a random-effects meta-analysis that pools nine MUStARD-based studies using Hedges' g, the Q-statistic and I² for heterogeneity, and tau² for between-study variance; this supports the conclusion that mean F1 differences between attention and encoder-decoder fusion (about 4.5 points speaker-independent, 1.1 points speaker-dependent) are not st
What would settle it
Run a pre-registered benchmark on a new spontaneous, multilingual sarcasm corpus: train matched attention and encoder-decoder fusion models with identical features and compute the F1 difference with confidence intervals. If the gap becomes significant, or if pooling five additional MUStARD 5-fold studies flips the random-effects p-value below 0.05, the review's equivalence conclusion fails to generalize.
Extended reading notes
Core claim
Speech sarcasm recognition has moved from audio-only systems (MFCCs, pitch, intensity; GMMs, HMMs, SVMs) to multimodal systems fusing audio with BERT text and ResNet video features; no modality or feature set suffices. The quantitative core is a random-effects meta-analysis of nine comparable MUStARD 5-fold studies: encoder-decoder versus attention fusion gives weighted F1 means of 67.2 vs 71.7 (speaker-independent, p=0.463) and 74.0 vs 75.1 (speaker-dependent, p=0.630), with I²=83–89%. The authors conclude the two fusion families perform equivalently, and progress needs spontaneous multilingual datasets, standardized prosodic features, and linguistically informed fusion.
Load-bearing premise
The review's conclusions stand on the assumption that the 40 selected studies—especially the nine MUStARD-based 5-fold studies pooled in the meta-analysis—are representative of the speech-sarcasm literature and comparable enough to combine; if the search or comparability criteria bias the pool, the no-difference fusion result could change.
Editorial extensions
If this is right
- Fusion architecture is not the main lever: teams choosing between attention and encoder-decoder designs should decide on compute, interpretability, or data fit, not expected F1.
- Datasets are the field's limiting resource: expect progress from spontaneous, multilingual, richly labeled corpora rather than from new fusion layers on MUStARD.
- Feature engineering still matters: audio features need standardized, controlled comparisons rather than ad hoc MFCC and pitch sets.
- Sarcasm should be treated as multimodal in benchmarks: single-modality scores understate what is linguistically needed.
- Failure modes point to labels: errors concentrate on modal mismatch, neutral cues, and annotation limits, so granular labels beyond binary sarcasm could move performance.
Reading between the lines
- If attention and encoder-decoder are truly tied, then published performance differences likely come from feature extraction, context handling, or data splits, so re-ranking models on one standardized benchmark could reshuffle the leaderboard.
- The equivalence result is fragile because only nine studies met inclusion criteria and heterogeneity was high; adding a few new MUStARD-based studies to the meta-analysis could plausibly flip significance, making this a testable rather than settled claim.
- The review's own dataset critique suggests that MUStARD's acted, laugh-tracked English prosody may mask larger gaps between fusion methods; a spontaneous cross-lingual corpus could reveal differences that MUStARD does not.
- Cross-linguistic prosody findings—pitch rises with sarcasm in Cantonese, French, and Italian but falls in German—imply that English-trained multimodal systems should be stress-tested for culturally specific cue weighting.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript presents a PRISMA-guided systematic review of speech-based sarcasm recognition, synthesizing 40 studies retrieved from five databases up to December 2024. It organizes the literature according to three research questions: available datasets and their limitations, the evolution of feature extraction, and the evolution of classification/fusion methods. It also includes a random-effects meta-analysis comparing F1-scores of attention-based versus encoder-decoder fusion methods on MUStARD-based studies, an error analysis, and a discussion of cross-cultural, explainability, and practical issues. The central claims are that speech-based sarcasm recognition has moved from unimodal audio analysis to multimodal fusion, that no single modality or feature set is sufficient, and that the meta-analysis found no statistically significant performance difference between attention and encoder-decoder fusion in either speaker-independent or speaker-dependent settings.
Significance. If the descriptive synthesis is taken on its own, the paper makes a useful contribution: it is the first systematic review specifically focused on speech-based sarcasm recognition, with a clearly documented search protocol, a PRISMA flow diagram, structured data-extraction tables, a taxonomy of fusion strategies, and a grounded error analysis. These elements provide a practical map for researchers entering the field and are a genuine strength. The quantitative meta-analytic claim, however, is not statistically sound as presented: Hedges' g is applied to single per-study F1-scores without within-study variances or sample sizes, the reported Q, I², and p-values are not reproducible from the tables, and the small number of studies with very high heterogeneity makes 'no difference' an overstatement. The descriptive review can stand, but the meta-analysis and the conclusions built on it require correction.
major comments (4)
- [§III.C.3.b, Tables III and IV] The meta-analysis is not reproducible from the reported data. Hedges' g is a standardized mean difference between two groups within a primary study; it requires per-study group means, within-group standard deviations, and sample sizes. Here each study contributes a single F1 percentage, and no variance, sample size, or per-study effect size is reported. The pooled means, confidence intervals, Q = 30.00 and 72.00, I² = 83.33% and 88.89%, τ², and p-values therefore cannot be verified. Because this is the basis for the headline 'no statistically significant difference' result, the quantitative synthesis must either be conducted with a defensible effect measure (e.g., a single-proportion meta-analysis with an explicit variance for F1, or a proper contrast using per-study group data) with full input data supplied, or removed and replaced by a clearly labeled descriptive comparison.
- [§III.C.3.b, text before and after Tables III and IV] With k = 6 (speaker-independent) and k = 9 (speaker-dependent) and I² between 83% and 89%, the random-effects estimate of between-study variance is poorly identified and the comparison has very low statistical power. The p-values 0.4630 and 0.6303 reflect absence of evidence, not evidence of no difference. The conclusion 'no statistically significant difference' and the abstract's framing of this as a null result overstate what the data can show. The authors should rephrase the conclusion as 'insufficient evidence to establish a difference' and present the descriptive trend (attention mean numerically higher in both settings) as a hypothesis for future work.
- [§III.C.2 and §V] The claim that multimodal approaches 'significantly enhance' sarcasm recognition performance compared to unimodal approaches is asserted without a quantitative synthesis. The appendix tables (A1 and A2) intermix different datasets, metrics, feature sets, and evaluation protocols, and no paired comparisons or common benchmark are provided. Since the review itself repeatedly stresses comparability issues, this central narrative claim should be either substantiated with same-dataset comparisons (e.g., MUStARD audio-only versus multimodal systems with controlled model families) or downgraded to a qualitative trend.
- [§II, p. iii] The manuscript states that PRISMA items on effect measures (#13) and risk of bias (#12, #15, #19, #22) are excluded as beyond the scope, yet §III.C.3 performs a meta-analysis. PRISMA 2020 requires specifying the effect measure and assessing risk of bias for any quantitative synthesis; the excluded items are exactly the components needed for a valid meta-analysis. The reporting standard claimed in Section II is therefore not met for the meta-analytic portion. The authors should either include these PRISMA elements for the quantitative synthesis or remove the meta-analysis from the review and treat the comparison as a descriptive summary.
minor comments (4)
- [§III.B.3, Visual features] 'Hasen and Patil [40]' should be 'Hasan and Patil'; later in the same section 'incoprprating' should be 'incorporating'.
- [Table A1] Minor typos: 'Wav2cev2.0' should be 'Wav2Vec2.0', and 'datsset' should be 'dataset'.
- [§III.A.1.f] The citation to Fan et al. [34] appears mismatched: the text attributes to this work the finding that 'anticipation of sarcastic intent is crucial for the efficient comprehension of sarcasm,' but the reference is the 'exposure advantage' paper on multilingual exposure. Please verify and correct the citation.
- [§III.C.3.a] The attrition description ('19 articles ... excluded five ... and five ... resulting in nine articles') is confusing on first reading. A small inclusion-attrition table or a step-by-step count would improve clarity.
Circularity Check
No significant circularity: the review and meta-analysis rest on external published results; the authors' own papers are non-load-bearing examples.
full rationale
The derivation chain of this paper is a systematic literature review plus a meta-analysis of published F1-scores, not a first-principles derivation. The central claims—that sarcasm recognition has moved from unimodal to multimodal fusion and that no single modality or feature set is sufficient—are supported by external studies (e.g., [21], [29], [35], [40]) and by the survey's own tabulated evidence. The meta-analysis in Section III.C.3 compares attention mechanisms with encoder-decoder fusion using F1-scores reported in external MUStARD-based studies (Tables III and IV); no fitted parameter is renamed as a prediction, and the null result is a statistical aggregation of those external numbers, not a consequence of the authors' own model outputs. The authors' self-citations ([42], [70]) appear in the narrative as examples of feature use, and [70] is excluded from the meta-analysis because it uses a train-test split rather than 5-fold cross-validation; [42] is a unimodal audio study and is not part of the fusion comparison. Thus the self-citations are not load-bearing. The paper itself flags limitations: it excludes PRISMA risk-of-bias items and effect measures ('Items related to risk of bias (#12, 15, 19, 22), effect measures (#13), additional analyses (#16, 23) are excluded'), and the meta-analysis is under-reported with no per-study effect sizes or variances, which is a methodological validity concern but not circularity. No equation or definition in the paper reduces a claimed result to its own inputs. Therefore the appropriate finding is no significant circularity, score 0.
Assumptions & free parameters
assumptions (5)
- domain assumption PRISMA 2020 guidelines are an appropriate methodological framework for this systematic review.
- domain assumption The five selected databases (Scopus, IEEE Xplore, ISCA, ScienceDirect, ACM DL) and English-only filter capture the relevant literature.
- domain assumption Zhao et al.'s fusion taxonomy (encoder-decoder, attention, gating, quantum) adequately classifies multimodal sarcasm architectures.
- standard math Random-effects meta-analysis with Hedges' g, Q and I-squared is valid for pooling F1-scores across heterogeneous studies.
- domain assumption F1-score is a balanced metric appropriate for class-imbalanced sarcasm datasets.
Cite this review
Pith. "Pith review of Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects." pith.science (2026). https://pith.science/paper/JG2KF7Z6
@misc{pith2026250904605,
author = {Pith},
title = {Pith review of: Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects},
year = {2026},
howpublished = {\url{https://pith.science/paper/JG2KF7Z6}},
note = {Machine review of arXiv:2509.04605}
}
read the original abstract
Sarcasm, a common feature of human communication, poses challenges in interpersonal interactions and human-machine interactions. Linguistic research has highlighted the importance of prosodic cues, such as variations in pitch, speaking rate, and intonation, in conveying sarcastic intent. Although previous work has focused on text-based sarcasm detection, the role of speech data in recognizing sarcasm has been underexplored. Recent advancements in speech technology emphasize the growing importance of leveraging speech data for automatic sarcasm recognition, which can enhance social interactions for individuals with neurodegenerative conditions and improve machine understanding of complex human language use, leading to more nuanced interactions. This systematic review is the first to focus on speech-based sarcasm recognition, charting the evolution from unimodal to multimodal approaches. It covers datasets, feature extraction, and classification methods, and aims to bridge gaps across diverse research domains. The findings include limitations in datasets for sarcasm recognition in speech, the evolution of feature extraction techniques from traditional acoustic features to deep learning-based representations, and the progression of classification methods from unimodal approaches to multimodal fusion techniques. In so doing, we identify the need for greater emphasis on cross-cultural and multilingual sarcasm recognition, as well as the importance of addressing sarcasm as a multimodal phenomenon, rather than a text-based challenge.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Irony and use-mention distinction,
D. Sperber and D. Wilson, “Irony and use-mention distinction,” in Radical Pragmatics , P. Cole, Ed. New York, NY , USA: Academic Press, 1981, pp. 295–318
1981
-
[2]
The functions of sarcastic irony in speech,
J. Jorgensen, “The functions of sarcastic irony in speech,” J. Pragmat., vol. 26, no. 5, pp. 613–634, Nov. 1996
1996
-
[3]
Brown and S
P. Brown and S. C. Levinson, Politeness: Some Universals in Language Usage. Cambridge, UK: Cambridge University Press, 1987
1987
-
[4]
Jocularity, sarcasm, and relation- ships: An empirical study,
M. A. Seckman and C. J. Couch, “Jocularity, sarcasm, and relation- ships: An empirical study,” J. Contemp. Ethnogr. , vol. 18, no. 3, pp. 327–344, July 1989
1989
-
[5]
On irony and negation,
R. Giora, “On irony and negation,” Discourse Process., vol. 19, no. 2, pp. 239–264, Nov. 1995
1995
-
[6]
Exploring the process of inference generation in sar- casm: A review of normal and clinical studies,
S. McDonald, “Exploring the process of inference generation in sar- casm: A review of normal and clinical studies,” Brain Lang., vol. 68, no. 3, pp. 486–506, July 1999
1999
-
[7]
Logic and conversation,
H. P. Grice, “Logic and conversation,” in Syntax and Semantics, Vol. 3: Speech Acts, P. Cole and J. L. Morgan, Eds. New York, NY , USA: Academic Press, 1975, pp. 41–58
1975
-
[8]
Test of the mention theory of irony,
J. Jorgensen, G. A. Miller, and D. Sperber, “Test of the mention theory of irony,” J. Exp. Psychol. Gen. , vol. 113, no. 1, pp. 112–120, 1984
1984
Show all 110 references
-
[9]
Understanding social dysfunction in the behavioural variant of frontotemporal dementia: The role of emotion and sarcasm processing,
C. M. Kipps, P. J. Nestor, J. Acosta-Cabronero, R. Arnold, and J. R. Hodges, “Understanding social dysfunction in the behavioural variant of frontotemporal dementia: The role of emotion and sarcasm processing,” Brain, vol. 132, no. 3, pp. 592–603, Jan. 2009
2009
-
[10]
Teaching children with autism to detect and respond to sarcasm,
A. Persicke, J. Tarbox, J. Ranick, and M. S. Clair, “Teaching children with autism to detect and respond to sarcasm,” Res. Autism Spectr. Disord., vol. 7, no. 1, pp. 193–198, Jan. 2013
2013
-
[11]
Lower, slower, louder: V ocal cues of sarcasm,
P. Rockwell, “Lower, slower, louder: V ocal cues of sarcasm,” J. Psycholinguist. Res., vol. 29, no. 5, pp. 483–495, Sep. 2000
2000
-
[12]
Acoustic markers of sarcasm in cantonese and english,
H. S. Cheang and M. D. Pell, “Acoustic markers of sarcasm in cantonese and english,” J. Acoust. Soc. Am. , vol. 126, no. 3, pp. 1394– 1405, Sep. 2009
2009
-
[13]
V oice modulations in german ironic speech,
L. Scharrer and U. Christmann, “V oice modulations in german ironic speech,” Lang. Speech, vol. 54, no. 4, pp. 435–465, Dec. 2011. xviii
2011
-
[14]
Prosodic cues of sarcastic speech in french: Slower, higher, wider,
H. Lœvenbruck, M. A. B. Jannet, M. D’Imperio, M. Spini, and M. Champagne-Lavau, “Prosodic cues of sarcastic speech in french: Slower, higher, wider,” in Proc. Interspeech, Lyon, France, Aug. 2013, pp. 3537–3541
2013
-
[15]
‘yeah right’: Sarcasm recognition for spoken dialogue systems,
J. Tepperman, D. Traum, and S. Narayanan, “‘yeah right’: Sarcasm recognition for spoken dialogue systems,” in Proc. Interspeech, Pitts- burgh, PA, USA, Sep. 2006, pp. 1838–1841
2006
-
[16]
‘sure, i did the right thing’: A system for sarcasm detection in speech,
R. Rakov and A. Rosenberg, “‘sure, i did the right thing’: A system for sarcasm detection in speech,” in Proc. Interspeech, Lyon, France, Aug. 2013, pp. 842–846
2013
-
[17]
Automatic sarcasm detection: A survey,
A. Joshi, P. Bhattacharyya, and M. J. Carman, “Automatic sarcasm detection: A survey,” ACM Comput. Surv. , vol. 50, no. 5, pp. 1–22, Sep. 2017
2017
-
[18]
Sarcasm identification in textual data: Systematic review, research challenges and open directions,
C. I. Eke, A. A. Norman, L. Shuib, and H. F. Nweke, “Sarcasm identification in textual data: Systematic review, research challenges and open directions,” Artif. Intell. Rev., vol. 53, no. 6, pp. 4215–4258, Nov. 2020
2020
-
[19]
Literature survey of sarcasm detection,
P. Chaudhari and C. Chandankhede, “Literature survey of sarcasm detection,” in Proc. 2017 Int. Conf. Wireless Commun., Signal Process. Netw. (WiSPNET), Chennai, India, Mar. 2017, pp. 2041–2046
2017
-
[20]
Pre- ferred reporting items for systematic reviews and meta-analyses: The prisma statement,
D. Moher, A. Liberati, J. Tetzlaff, D. G. Altman, and T. P. Group, “Pre- ferred reporting items for systematic reviews and meta-analyses: The prisma statement,” PLOS Med., vol. 6, no. 7, 2009, doi: 10.1371/jour- nal.pmed.1000097
2009 doi
-
[21]
Towards multimodal sarcasm detection (an obviously perfect paper),
S. Castro, D. Hazarika, V . P ´erez-Rosas, R. Zimmermann, R. Mihalcea, and S. Poria, “Towards multimodal sarcasm detection (an obviously perfect paper),” in Proc. 57th Annu. Meeting Assoc. Comput. Linguis- tics (ACL), Florence, Italy, Aug. 2019, pp. 4619–4629
2019
-
[22]
IITKGP-SEHSC: Hindi speech corpus for emotion analysis,
S. G. Koolagudi, R. Reddy, J. Yadav, and K. S. Rao, “IITKGP-SEHSC: Hindi speech corpus for emotion analysis,” in Proc. 2011 Int. Conf. Devices Commun. (ICDeCom) , Mesra, Ranchi, India, Feb. 2011, pp. 1–5
2011
-
[23]
Multi-modal sarcasm detection and humor classification in code-mixed conversa- tions,
M. Bedi, S. Kumar, M. S. Akhtar, and T. Chakraborty, “Multi-modal sarcasm detection and humor classification in code-mixed conversa- tions,” IEEE Trans. Affective Comput. , vol. 14, no. 2, pp. 1363–1375, 2023
2023
-
[24]
¡Qu´e maravilla! multimodal sarcasm detection in spanish: A dataset and a baseline,
K. Alnajjar and M. H ¨am¨al¨ainen, “¡Qu´e maravilla! multimodal sarcasm detection in spanish: A dataset and a baseline,” in Proc. 3rd Workshop Multimodal Artif. Intell. , Mexico City, Mexico, Jun. 2021, pp. 63–68
2021
-
[25]
CMMA: Benchmarking multi-affection detection in Chinese multi-modal conversations,
Y . Zhang et al. , “CMMA: Benchmarking multi-affection detection in Chinese multi-modal conversations,” in Proc. 37th Int. Conf. Neural Information Processing Systems (NeurIPS) , New Orleans, LA, USA, 2023, pp. 18 794–18 805, [Online]. Available: https://openreview.net/f orum?...
2023
-
[26]
Deep learning for acous- tic irony classification in spontaneous speech,
H. Gent, C. Adams, C. Shih, and Y . Tang, “Deep learning for acous- tic irony classification in spontaneous speech,” in Proc. Interspeech, Incheon, South Korea, Sep. 2022, pp. 3993–3997
2022
-
[27]
Sentiment and emotion help sarcasm? a multi-task learning framework for multi- modal sarcasm, sentiment, and emotion analysis,
D. S. Chauhan, D. S. R, A. Ekbal, and P. Bhattacharyya, “Sentiment and emotion help sarcasm? a multi-task learning framework for multi- modal sarcasm, sentiment, and emotion analysis,” in Proc. 58th Annu. Meet. Assoc. Comput. Linguistics (ACL) , Seattle, W A, USA, Jul. 2020, p...
2020
-
[28]
An emoji-aware multitask framework for multimodal sarcasm detection,
D. S. Chauhan, G. Singh, A. Arora, A. Ekbal, and P. Bhat- tacharyya, “An emoji-aware multitask framework for multimodal sarcasm detection,” Knowl.-Based Syst. , vol. 257, 2022, doi: 10.1016/j.knosys.2022.109924
2022
-
[29]
A multimodal corpus for emotion recognition in sarcasm,
A. Ray, S. Mishra, A. Nunna, and P. Bhattacharyya, “A multimodal corpus for emotion recognition in sarcasm,” in Proc. 13th Lang. Resour. Eval. Conf. (LREC) , Marseille, France, June 2022, pp. 6992–7003
2022
-
[30]
MELD: A multimodal multi-party dataset for emotion recognition in conversations,
S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea, “MELD: A multimodal multi-party dataset for emotion recognition in conversations,” in Proc. 57th Annu. Meet. Assoc. Com- put. Linguistics (ACL) , Florence, Italy, Jul. 2019, pp. 527–536
2019
-
[31]
Which contextual and socio- cultural information predict irony perception?
E. Rivi `ere and M. Champagne-Lavau, “Which contextual and socio- cultural information predict irony perception?” Discourse Process. , vol. 57, no. 3, pp. 259–277, 2020
2020
-
[32]
Verbal irony as implicit display of ironic environment: Distinguishing ironic utterances from nonirony,
A. Utsumi, “Verbal irony as implicit display of ironic environment: Distinguishing ironic utterances from nonirony,” J. Pragmat., vol. 32, no. 12, pp. 1777–1806, Nov. 2000
2000
-
[33]
How about another piece of pie: The allusional pretense theory of discourse irony,
S. Kumon-Nakamura, S. Glucksberg, and M. Brown, “How about another piece of pie: The allusional pretense theory of discourse irony,” J. Exp. Psychol. Gen. , vol. 124, no. 1, pp. 3–21, 1995
1995
-
[34]
The exposure advantage: Early exposure to a multilingual environment promotes effective communication,
S. P. Fan, Z. Liberman, B. Keysar, and K. D. Kinzler, “The exposure advantage: Early exposure to a multilingual environment promotes effective communication,” Psychol. Sci., vol. 26, no. 7, pp. 1090–1097, Jul. 2015
2015
-
[35]
Modeling incongruity between modalities for multimodal sarcasm detection,
Y . Wu et al., “Modeling incongruity between modalities for multimodal sarcasm detection,” IEEE MultiMedia, vol. 28, no. 2, pp. 86–95, Apr. 2021
2021
-
[36]
The coding strategy for the mandarin speech conveying sarcasm in acoustic and articulatory domain,
P. Geng, S. Shi, and H. Guo, “The coding strategy for the mandarin speech conveying sarcasm in acoustic and articulatory domain,” in Proc. 5th Int. Conf. Digit. Signal Process. (DSP) , Chengdu, China, Feb. 2021, pp. 195–200
2021
-
[37]
Detecting vocal irony,
F. Burkhardt, B. Weiss, F. Eyben, J. Deng, and B. Schuller, “Detecting vocal irony,” in Proc. 27th Int. Conf. Lang. Technol. Challenges Digit. Age (GSCL), vol. 10713, 2018, pp. 11–22
2018
-
[38]
Emotional vocal expressions recognition using the cost 2102 italian database of emotional speech,
H. Atassi, M. T. Riviello, Z. Sm ´ekal, A. Hussain, and A. Esposito, “Emotional vocal expressions recognition using the cost 2102 italian database of emotional speech,” in Development of Multimodal Inter- faces: Active Listening and Synchrony . Berlin, Germany: Springer, 2010,...
2010
-
[39]
Emotion recognition in speech using machine learning techniques,
A. Arun, I. Rallabhandi, S. Hebbar, A. Nair, and R. Jayashree, “Emotion recognition in speech using machine learning techniques,” in Proc. 12th Int. Conf. Comput. Commun. Netw. Technol. (ICCCNT) , Kharagpur, India, Jul. 2021, pp. 1–7
2021
-
[40]
Humor knowledge enriched transformer for understanding multimodal humor,
M. K. Hasan et al. , “Humor knowledge enriched transformer for understanding multimodal humor,” in Proc. AAAI Conf. Artif. Intell. , vol. 35, no. 14, May 2021, pp. 12 972–12 980
2021
-
[41]
Measuring nominal scale agreement among many raters,
J. L. Fleiss, “Measuring nominal scale agreement among many raters,” Psychol. Bull., vol. 76, no. 5, pp. 378–382, Nov. 1971
1971
-
[42]
Deep cnn-based inductive trans- fer learning for sarcasm detection in speech,
X. Gao, S. Nayak, and M. Coler, “Deep cnn-based inductive trans- fer learning for sarcasm detection in speech,” in Proc. Interspeech , Incheon, South Korea, Sep. 2022, pp. 2323–2327
2022
-
[43]
Emotion recognition of speech,
N. S. Sacheth and R. Jayashree, “Emotion recognition of speech,” in Proc. 14th Int. Conf. Soft Comput. Pattern Recognit. (SoCPaR) . Springer, 2023, vol. 490, pp. 359–371
2023
-
[44]
Characterization of emotions using the dynamics of prosodic features,
K. S. Rao, R. Reddy, S. Maity, and S. G. Koolagudi, “Characterization of emotions using the dynamics of prosodic features,” in Proc. Speech Prosody, Chicago, IL, USA, May 2010, [Online]. Available: https://ww w.isca-archive.org/speechprosody 2010/rao10 speechprosody.html
2010
-
[45]
Understanding sarcasm in speech using mel-frequency cepstral coefficient,
A. Mathur, V . Saxena, and S. K. Singh, “Understanding sarcasm in speech using mel-frequency cepstral coefficient,” in Proc. 7th Int. Conf. Cloud Comput., Data Sci. Eng. (Confluence) , Noida, India, Jan. 2017, pp. 728–732
2017
-
[46]
The interspeech 2013 computational paralinguistics challenge: Social signals, conflict, emotion, autism,
B. Schuller et al., “The interspeech 2013 computational paralinguistics challenge: Social signals, conflict, emotion, autism,” in Proc. Inter- speech, Lyon, France, Aug. 2013, pp. 148–152
2013
-
[47]
Multi-modal sarcasm detection based on contrastive attention mechanism,
X. Zhang, Y . Chen, and G. Li, “Multi-modal sarcasm detection based on contrastive attention mechanism,” inProc. Nat. Lang. Process. Chin. Comput., vol. 13028, 2021, pp. 822–833
2021
-
[48]
Multimodal sarcasm detection: A deep learning approach,
S. K. Bharti, R. Gupta, P. K. Shukla, W. A. Hatamleh, H. Tarazi, and S. J. Nuagah, “Multimodal sarcasm detection: A deep learning approach,” Wirel. Commun. Mob. Comput., vol. 2022, pp. 1–10, 2022, doi: 10.1155/2022/1653696
2022 doi
-
[49]
A multimodal fusion method for sarcasm detection based on late fusion,
N. Ding, S.-W. Tian, and L. Yu, “A multimodal fusion method for sarcasm detection based on late fusion,” Multimed. Tools Appl., vol. 81, no. 6, pp. 8597–8616, Feb. 2022
2022
-
[50]
EFAFN: An efficient feature adaptive fusion network with facial feature for multimodal sarcasm detection,
Y . Sun, H. Zhang, S. Yang, and J. Wang, “EFAFN: An efficient feature adaptive fusion network with facial feature for multimodal sarcasm detection,” Appl. Sci. , vol. 12, no. 21, p. 11235, Nov. 2022, doi: 10.3390/app122111235
2022 doi
-
[51]
Multimodal learning using optimal transport for sarcasm and humor detection,
S. Pramanick, A. Roy, and V . M. Patel, “Multimodal learning using optimal transport for sarcasm and humor detection,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV) , Waikoloa, HI, USA, Jan. 2022, pp. 546–556
2022
-
[52]
Multimodal sarcasm detection (MSD) in videos using deep learning models,
A. Pandey and D. K. Vishwakarma, “Multimodal sarcasm detection (MSD) in videos using deep learning models,” in Proc. 2023 Int. Conf. Adv. Power, Signal, Inf. Technol. (APSIT), 2023, pp. 811–814
2023
-
[53]
Multi- modal sarcasm detection method using rnn and cnn,
B. Azahouani, H. Elfaik, E. H. Nfaoui, and S. E. Garouani, “Multi- modal sarcasm detection method using rnn and cnn,” in Proc. 6th Int. Conf. Intell. Comput. Data Sci. (ICDS) , Marrakech, Morocco, 2024, pp. 1–6
2024
-
[54]
An attention-based, context-aware multimodal fusion method for sarcasm detection using inter-modality inconsistency,
Y . Li et al. , “An attention-based, context-aware multimodal fusion method for sarcasm detection using inter-modality inconsistency,” Knowl.-Based Syst., vol. 287, p. 111457, Mar. 2024
2024
-
[55]
A smart video analytical frame- work for sarcasm detection using novel adaptive fusion network and sarcasnet-99 model,
J. S. Murthy and G. M. Siddesh, “A smart video analytical frame- work for sarcasm detection using novel adaptive fusion network and sarcasnet-99 model,” Vis. Comput. , vol. 40, no. 11, pp. 8085–8097, Nov. 2024
2024
-
[56]
V oice-based sarcasm detection in kannada language,
R. Manohar and S. Swamy, “V oice-based sarcasm detection in kannada language,” Int. J. Intell. Syst. Appl. Eng., vol. 12, no. 14s, pp. 356–367, Feb. 2024, [Online]. Available: https://ijisae.org/index.php/IJISAE/arti cle/view/4672. xix
2024
-
[57]
Sarcasm detection using cognitive features of visual data by learning model,
B. N. Hiremath and M. M. Patil, “Sarcasm detection using cognitive features of visual data by learning model,” Expert Syst. Appl., vol. 184, Dec. 2021
2021
-
[58]
Cnn architectures for large-scale audio classifi- cation,
S. Hershey et al. , “Cnn architectures for large-scale audio classifi- cation,” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), New Orleans, LA, USA, Mar. 2017, pp. 131–135
2017
-
[59]
wav2vec 2.0: A framework for self-supervised learning of speech representations,
A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Adv. Neural Inf. Process. Syst. , vol. 33, pp. 12 449–12 460, 2020
2020
-
[60]
A multitask learning model for multimodal sarcasm, sentiment and emotion recognition in conversations,
Y . Zhang et al., “A multitask learning model for multimodal sarcasm, sentiment and emotion recognition in conversations,” Inf. Fusion , vol. 93, pp. 282–301, May 2023
2023
-
[61]
Attention is all you need,
A. Vaswani et al., “Attention is all you need,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , vol. 30, Long Beach, CA, USA, Dec. 2017, pp. 5998–6008
2017
-
[62]
Learning phrase representations using rnn encoder–decoder for statistical machine translation,
K. Cho et al. , “Learning phrase representations using rnn encoder–decoder for statistical machine translation,” arXiv preprint arXiv:1406.1078, 2014, [Online]. Available: https://arxiv.org/abs/14 06.1078
2014 arXiv
-
[63]
PRAAT, a system for doing phonetics by computer,
P. Boersma, “PRAAT, a system for doing phonetics by computer,” Glot Int., vol. 5, no. 9/10, pp. 341–345, 2001
2001
-
[64]
OpenSMILE: The munich versatile and fast open-source audio feature extractor,
F. Eyben, M. W ¨ollmer, and B. Schuller, “OpenSMILE: The munich versatile and fast open-source audio feature extractor,” in Proc. 18th ACM Int. Conf. Multimedia, Florence, Italy, Oct. 2010, pp. 1459–1462
2010
-
[65]
Librosa: Audio and music signal analysis in python,
B. McFee et al., “Librosa: Audio and music signal analysis in python,” in Proc. Python in Sci. Conf. (SciPy) , Austin, TX, USA, 2015, pp. 18– 24, doi: 10.25080/Majora-7b98e3ed-003
2015 doi
-
[66]
CO- V AREP: A collaborative voice analysis repository for speech tech- nologies,
G. Degottex, J. Kane, T. Drugman, T. Raitio, and S. Scherer, “CO- V AREP: A collaborative voice analysis repository for speech tech- nologies,” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), Florence, Italy, May 2014, pp. 960–964
2014
-
[67]
BERT: Pre- training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre- training of deep bidirectional transformers for language understanding,” in Proc. Conf. North Amer. Chapter Assoc. Comput. Linguistics: Human Lang. Technol. (NAACL-HLT) , Minneapolis, MN, USA, Jun. 2019, pp. 4171–4186
2019
-
[68]
A quantum probability driven frame- work for joint multi-modal sarcasm, sentiment and emotion analysis,
Y . Liu, Y . Zhang, and D. Song, “A quantum probability driven frame- work for joint multi-modal sarcasm, sentiment and emotion analysis,” IEEE Trans. Affect. Comput. , pp. 1–15, 2023
2023
-
[69]
Learning multitask commonness and uniqueness for multimodal sarcasm detection and sentiment analysis in conversation,
Y . Zhang et al., “Learning multitask commonness and uniqueness for multimodal sarcasm detection and sentiment analysis in conversation,” IEEE Trans. Artif. Intell. , vol. 5, no. 3, pp. 1349–1361, 2024
2024
-
[70]
Improving sarcasm detection from speech and text through attention-based fusion exploiting the interplay of emotions and sentiments,
X. Gao, S. Nayak, and M. Coler, “Improving sarcasm detection from speech and text through attention-based fusion exploiting the interplay of emotions and sentiments,” in Proc. Meet. Acoust. , vol. 54, no. 1, Jan. 2024, p. 060002, doi:10.1121/2.0001918
2024 doi
-
[71]
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,
M. Lewis et al., “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proc. 58th Annu. Meet. Assoc. Comput. Linguistics (ACL) , Online, Jul. 2020, pp. 7871–7880
2020
-
[72]
Your tone speaks louder than your face! modality order infused multi-modal sarcasm detection,
M. Tomar, A. Tiwari, T. Saha, and S. Saha, “Your tone speaks louder than your face! modality order infused multi-modal sarcasm detection,” in Proc. 31st ACM Int. Conf. Multimedia (MM) , Ottawa, ON, Canada, Oct. 2023, pp. 3926–3933
2023
-
[73]
Quantum fuzzy neural network for multimodal sentiment and sarcasm detection,
P. Tiwari, L. Zhang, Z. Qu, and G. Muhammad, “Quantum fuzzy neural network for multimodal sentiment and sarcasm detection,” Inf. Fusion, vol. 97, p. 102085, Jan. 2024
2024
-
[74]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Las Vegas, NV , USA, Jun. 2016, pp. 770–778
2016
-
[75]
EfficientNet: Rethinking model scaling for convolutional neural networks,
M. Tan and Q. V . Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proc. 36th Int. Conf. Mach. Learn. (ICML), vol. 97, 2019, pp. 6105–6114
2019
-
[76]
Joint face detection and alignment using multi-task cascaded convolutional networks,
K. Zhang, Z. Zhang, Z. Li, and Y . Qiao, “Joint face detection and alignment using multi-task cascaded convolutional networks,” IEEE Signal Process. Lett. , vol. 23, no. 10, pp. 1499–1503, 2016
2016
-
[77]
Openface 2.0: Facial behavior analysis toolkit,
T. Baltrusaitis, A. Zadeh, Y . C. Lim, and L.-P. Morency, “Openface 2.0: Facial behavior analysis toolkit,” in Proc. 13th IEEE Int. Conf. Autom. Face Gesture Recognit. (FG) , Xi’an, China, May 2018, pp. 59–66
2018
-
[78]
Quo vadis, action recognition? a new model and the kinetics dataset,
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2017, pp. 4724–4733
2017
-
[79]
Multimodal markers of irony and sarcasm,
S. Attardo, J. Eisterhold, J. Hay, and I. Poggi, “Multimodal markers of irony and sarcasm,” Humor - Int. J. Humor Res. , vol. 16, no. 2, pp. 243–260, Jan. 2003
2003
-
[80]
Verbal irony use in face-to-face and computer-mediated conversations,
J. T. Hancock, “Verbal irony use in face-to-face and computer-mediated conversations,” J. Lang. Soc. Psychol. , vol. 23, no. 4, pp. 447–463, 2004
2004
-
[81]
Effects of emotional intelligence on the impression of irony created by the mismatch between verbal and nonverbal cues,
J. H., B. Kreifelts, S. Nizielski, A. Sch ¨utz, and D. Wildgruber, “Effects of emotional intelligence on the impression of irony created by the mismatch between verbal and nonverbal cues,” PLOS ONE , vol. 11, no. 10, p. e0163211, 2016, doi: 10.1371/journal.pone.0163211
2016 doi
-
[82]
Deep multimodal data fusion,
F. Zhao, C. Zhang, and B. Geng, “Deep multimodal data fusion,” ACM Comput. Surv., vol. 56, no. 9, pp. 1–36, Oct. 2024
2024
-
[83]
L. V . Hedges and I. Olkin, Statistical Methods for Meta-Analysis . Academic Press, 1985
1985
-
[84]
Quantifying heterogeneity in a meta-analysis,
J. P. Higgins and S. G. Thompson, “Quantifying heterogeneity in a meta-analysis,” Stat. Med., vol. 21, no. 11, pp. 1539–1558, Jun. 2002
2002
-
[85]
Meta-analysis in clinical trials,
R. DerSimonian and N. Laird, “Meta-analysis in clinical trials,” Con- trol. Clin. Trials, vol. 7, no. 3, pp. 177–188, Sep. 1986
1986
-
[86]
The cost 2102 italian audio and video emotional database,
A. Esposito, M. T. Riviello, and G. Maio, “The cost 2102 italian audio and video emotional database,” in Front. Artif. Intell. Appl., Jan. 2009, vol. 204, pp. 51–61
2009
-
[87]
Sarcasm, pretense, and the semantics/pragmatics distinc- tion,
E. Camp, “Sarcasm, pretense, and the semantics/pragmatics distinc- tion,” Noˆus, vol. 46, no. 4, pp. 587–634, Dec. 2012
2012
-
[88]
WavLM: Large-scale self-supervised pre-training for full stack speech processing,
S. Chen et al. , “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE J. Sel. Top. Signal Process., vol. 16, no. 6, pp. 1505–1518, Oct. 2022
2022
-
[89]
SpeechLM: Enhanced speech pre-training with unpaired textual data,
Z. Zhang et al. , “SpeechLM: Enhanced speech pre-training with unpaired textual data,” IEEE/ACM Trans. Audio Speech Lang. Process., vol. 32, pp. 2177–2187, 2024
2024
-
[90]
Adapting WavLM for speech emotion recognition,
D. Diatlova, A. Udalov, V . Shutov, and E. Spirin, “Adapting WavLM for speech emotion recognition,” in Proc. Odyssey 2024: The Speaker and Language Recognition Workshop , 2024, pp. 303–308
2024
- [91]
- [92]
-
[93]
SarcasmBench: Towards evaluating large language models on sarcasm understanding,
Y . Zhang, C. Zou, Z. Lian, P. Tiwari, and J. Qin, “SarcasmBench: Towards evaluating large language models on sarcasm understanding,” arXiv preprint arXiv:2408.11319, 2024, [Online]. Available: https://do i.org/10.48550/arXiv.2408.11319
- [94]
-
[95]
Paraverbal expression of verbal irony: V ocal cues matter and facial cues even more,
M. Aguert, “Paraverbal expression of verbal irony: V ocal cues matter and facial cues even more,” J. Nonverbal Behav. , vol. 46, no. 1, pp. 45–70, Mar. 2022
2022
-
[96]
Tracking nonliteral language processing using audiovisual scenarios,
K. Rothermich et al. , “Tracking nonliteral language processing using audiovisual scenarios,” Can. J. Exp. Psychol. , vol. 75, no. 2, pp. 211– 220, 2021
2021
-
[97]
The role of auditory and visual cues in the interpretation of mandarin ironic speech,
S. Li, A. Chen, Y . Chen, and P. Tang, “The role of auditory and visual cues in the interpretation of mandarin ironic speech,” J. Pragmat., vol. 201, pp. 3–14, 2022
2022
-
[98]
Modality matters: Testing bilingual irony comprehension in the textual, auditory, and audio-visual modality,
K. Bromberek-Dyzman, K. Jankowiak, and P. Chełminiak, “Modality matters: Testing bilingual irony comprehension in the textual, auditory, and audio-visual modality,” J. Pragmat., vol. 180, pp. 219–231, 2021
2021
-
[99]
Context, facial expression and prosody in irony processing,
G. Deliens, K. Antoniou, E. Clin, E. Ostashchenko, and M. Kissine, “Context, facial expression and prosody in irony processing,” J. Mem. Lang., vol. 99, pp. 35–48, 2018
2018
-
[100]
The sound of sarcasm,
H. S. Cheang and M. D. Pell, “The sound of sarcasm,” Speech Commun., vol. 50, no. 5, pp. 366–381, May 2008
2008
-
[101]
From ‘blame by praise’ to ‘praise by blame’: Analysis of vocal patterns in ironic communication,
L. Anolli, R. Ciceri, and M. G. Infantino, “From ‘blame by praise’ to ‘praise by blame’: Analysis of vocal patterns in ironic communication,” Int. J. Psychol. , vol. 37, no. 5, pp. 266–276, Oct. 2002
2002
-
[102]
Categorical and schematic organization in memory,
J. M. Mandler, “Categorical and schematic organization in memory,” in Memory, Organization, and Structure , R. C. Puff, Ed. New York, NY , USA: Academic Press, 1979, pp. 259–299
1979
-
[103]
How Korean EFL learners understand sarcasm in L2 English,
J. Kim, “How Korean EFL learners understand sarcasm in L2 English,” J. Pragmat., vol. 60, pp. 193–206, 2014
2014
-
[104]
Self- adaptive representation learning model for multi-modal sentiment and sarcasm joint analysis,
Y . Zhang, Y . Yu, M. Wang, M. Huang, and M. S. Hossain, “Self- adaptive representation learning model for multi-modal sentiment and sarcasm joint analysis,” ACM Trans. Multimedia Comput. Commun. Appl., vol. 20, no. 5, May 2024
2024
-
[105]
Modified convolutional neural network with multiple features for multimodal sarcasm detection,
R. Krishnamaneni, M. Kurni, S. Sen, and A. Murthy, “Modified convolutional neural network with multiple features for multimodal sarcasm detection,” in Proc. 2nd Int. Conf. Recent Adv. Inf. Technol. Sustain. Dev. (ICRAIS), Manipal, India, 2024, pp. 160–165. xx
2024
-
[106]
The geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing,
F. Eyben et al. , “The geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing,” IEEE Trans. Affect. Comput., vol. 7, no. 2, pp. 190–202, 2016
2016
-
[107]
Understanding non-verbal irony markers: Machine learning insights versus human judgment,
M. Spitale, F. Catania, and F. Panzeri, “Understanding non-verbal irony markers: Machine learning insights versus human judgment,” in Proc. 26th ACM Int. Conf. Multimodal Interact. (ICMI), Nov. 2024, pp. 164– 172
2024
-
[108]
SWITCHBOARD: Telephone speech corpus for research and development,
J. J. Godfrey, E. C. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone speech corpus for research and development,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) , vol. 1, San Francisco, CA, USA, Mar. 1992, pp. 517–520
1992
-
[109]
The Fisher corpus: A resource for the next generations of speech-to-text,
C. Cieri, D. Miller, and K. Walker, “The Fisher corpus: A resource for the next generations of speech-to-text,” in Proc. 4th Int. Conf. Lang. Resour. Eval. (LREC) , Lisbon, Portugal, May 2004, pp. 69–71, [Online]. Available: https://aclanthology.org/L04-1500/. Xiyuan Gao recei...
2004
-
[2010]
In 2012, he returned to academia
After a postdoctoral appointment at his alma mater, he joined an AI start-up focusing on acoustic sensors, where he served as Head of the Cognitive Systems Unit. In 2012, he returned to academia. Dr. Coler is currently an Associate Professor of Speech Technology at the Faculty...
2012
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.