Pith. sign in

REVIEW 4 major objections 4 minor 110 references

Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects

T0 review · 4 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This systematic review of 40 studies argues that speech sarcasm recognition requires multimodal fusion, and its meta-analysis finds no statistically significant performance difference between attention and encoder-decoder fusion methods.

desk verdict Useful review of speech-based sarcasm recognition, but its meta-analysis of fusion methods is statistically shaky and needs correction. read the letter →

arxiv 2509.04605 v1 pith:JG2KF7Z6 submitted 2025-09-04 cs.CL

classification cs.CL
keywords speech-basedsarcasmrecognitionmultimodalfusionsystematicreviewmeta-analysisprosodicfeaturesMUStARDdatasetaffectivecomputingmultilingual
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is the first systematic review devoted to speech-based sarcasm recognition, spanning 40 empirical studies from unimodal audio-only systems to multimodal systems that fuse audio, text, and video. Its central claim is that sarcasm in speech cannot be captured by any single modality or feature set; robust recognition requires combining acoustic, lexical, and visual cues. The review charts dataset evolution (mostly English, mostly TV-sourced, with MUStARD as the common benchmark), feature evolution (from MFCCs and prosodic statistics toward deep embeddings such as VGGish and Wav2Vec 2.0), and classifier evolution (from rule-based and statistical models to attention-based and quantum-inspired fusion). Its most pointed result is a random-effects meta-analysis of nine MUStARD-based studies: encoder-decoder and attention-mechanism fusion do not differ statistically in F1, in either speaker-dependent or speaker-independent settings, even though attention shows a higher point estimate. If this holds, the practical message is that fusion architecture choice is not the current bottleneck; data scarcity, annotation quality, and linguistic diversity are.

What carries the argument

The analytical engine is a systematic-review screening pipeline (five databases, two search phases, 40 studies) feeding three structured analyses: dataset comparison, feature taxonomy, and classification taxonomy adopting a four-way fusion scheme (encoder-decoder, attention, collaborative gating, quantum-based). The load-bearing quantitative machinery is a random-effects meta-analysis that pools nine MUStARD-based studies using Hedges' g, the Q-statistic and I² for heterogeneity, and tau² for between-study variance; this supports the conclusion that mean F1 differences between attention and encoder-decoder fusion (about 4.5 points speaker-independent, 1.1 points speaker-dependent) are not st

What would settle it

Run a pre-registered benchmark on a new spontaneous, multilingual sarcasm corpus: train matched attention and encoder-decoder fusion models with identical features and compute the F1 difference with confidence intervals. If the gap becomes significant, or if pooling five additional MUStARD 5-fold studies flips the random-effects p-value below 0.05, the review's equivalence conclusion fails to generalize.

Watch

Extended reading notes

Core claim

Speech sarcasm recognition has moved from audio-only systems (MFCCs, pitch, intensity; GMMs, HMMs, SVMs) to multimodal systems fusing audio with BERT text and ResNet video features; no modality or feature set suffices. The quantitative core is a random-effects meta-analysis of nine comparable MUStARD 5-fold studies: encoder-decoder versus attention fusion gives weighted F1 means of 67.2 vs 71.7 (speaker-independent, p=0.463) and 74.0 vs 75.1 (speaker-dependent, p=0.630), with I²=83–89%. The authors conclude the two fusion families perform equivalently, and progress needs spontaneous multilingual datasets, standardized prosodic features, and linguistically informed fusion.

Load-bearing premise

The review's conclusions stand on the assumption that the 40 selected studies—especially the nine MUStARD-based 5-fold studies pooled in the meta-analysis—are representative of the speech-sarcasm literature and comparable enough to combine; if the search or comparability criteria bias the pool, the no-difference fusion result could change.

Editorial extensions

If this is right

  • Fusion architecture is not the main lever: teams choosing between attention and encoder-decoder designs should decide on compute, interpretability, or data fit, not expected F1.
  • Datasets are the field's limiting resource: expect progress from spontaneous, multilingual, richly labeled corpora rather than from new fusion layers on MUStARD.
  • Feature engineering still matters: audio features need standardized, controlled comparisons rather than ad hoc MFCC and pitch sets.
  • Sarcasm should be treated as multimodal in benchmarks: single-modality scores understate what is linguistically needed.
  • Failure modes point to labels: errors concentrate on modal mismatch, neutral cues, and annotation limits, so granular labels beyond binary sarcasm could move performance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If attention and encoder-decoder are truly tied, then published performance differences likely come from feature extraction, context handling, or data splits, so re-ranking models on one standardized benchmark could reshuffle the leaderboard.
  • The equivalence result is fragile because only nine studies met inclusion criteria and heterogeneity was high; adding a few new MUStARD-based studies to the meta-analysis could plausibly flip significance, making this a testable rather than settled claim.
  • The review's own dataset critique suggests that MUStARD's acted, laugh-tracked English prosody may mask larger gaps between fusion methods; a spontaneous cross-lingual corpus could reveal differences that MUStARD does not.
  • Cross-linguistic prosody findings—pitch rises with sarcasm in Cantonese, French, and Italian but falls in German—imply that English-trained multimodal systems should be stress-tested for culturally specific cue weighting.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This manuscript presents a PRISMA-guided systematic review of speech-based sarcasm recognition, synthesizing 40 studies retrieved from five databases up to December 2024. It organizes the literature according to three research questions: available datasets and their limitations, the evolution of feature extraction, and the evolution of classification/fusion methods. It also includes a random-effects meta-analysis comparing F1-scores of attention-based versus encoder-decoder fusion methods on MUStARD-based studies, an error analysis, and a discussion of cross-cultural, explainability, and practical issues. The central claims are that speech-based sarcasm recognition has moved from unimodal audio analysis to multimodal fusion, that no single modality or feature set is sufficient, and that the meta-analysis found no statistically significant performance difference between attention and encoder-decoder fusion in either speaker-independent or speaker-dependent settings.

Significance. If the descriptive synthesis is taken on its own, the paper makes a useful contribution: it is the first systematic review specifically focused on speech-based sarcasm recognition, with a clearly documented search protocol, a PRISMA flow diagram, structured data-extraction tables, a taxonomy of fusion strategies, and a grounded error analysis. These elements provide a practical map for researchers entering the field and are a genuine strength. The quantitative meta-analytic claim, however, is not statistically sound as presented: Hedges' g is applied to single per-study F1-scores without within-study variances or sample sizes, the reported Q, I², and p-values are not reproducible from the tables, and the small number of studies with very high heterogeneity makes 'no difference' an overstatement. The descriptive review can stand, but the meta-analysis and the conclusions built on it require correction.

major comments (4)
  1. [§III.C.3.b, Tables III and IV] The meta-analysis is not reproducible from the reported data. Hedges' g is a standardized mean difference between two groups within a primary study; it requires per-study group means, within-group standard deviations, and sample sizes. Here each study contributes a single F1 percentage, and no variance, sample size, or per-study effect size is reported. The pooled means, confidence intervals, Q = 30.00 and 72.00, I² = 83.33% and 88.89%, τ², and p-values therefore cannot be verified. Because this is the basis for the headline 'no statistically significant difference' result, the quantitative synthesis must either be conducted with a defensible effect measure (e.g., a single-proportion meta-analysis with an explicit variance for F1, or a proper contrast using per-study group data) with full input data supplied, or removed and replaced by a clearly labeled descriptive comparison.
  2. [§III.C.3.b, text before and after Tables III and IV] With k = 6 (speaker-independent) and k = 9 (speaker-dependent) and I² between 83% and 89%, the random-effects estimate of between-study variance is poorly identified and the comparison has very low statistical power. The p-values 0.4630 and 0.6303 reflect absence of evidence, not evidence of no difference. The conclusion 'no statistically significant difference' and the abstract's framing of this as a null result overstate what the data can show. The authors should rephrase the conclusion as 'insufficient evidence to establish a difference' and present the descriptive trend (attention mean numerically higher in both settings) as a hypothesis for future work.
  3. [§III.C.2 and §V] The claim that multimodal approaches 'significantly enhance' sarcasm recognition performance compared to unimodal approaches is asserted without a quantitative synthesis. The appendix tables (A1 and A2) intermix different datasets, metrics, feature sets, and evaluation protocols, and no paired comparisons or common benchmark are provided. Since the review itself repeatedly stresses comparability issues, this central narrative claim should be either substantiated with same-dataset comparisons (e.g., MUStARD audio-only versus multimodal systems with controlled model families) or downgraded to a qualitative trend.
  4. [§II, p. iii] The manuscript states that PRISMA items on effect measures (#13) and risk of bias (#12, #15, #19, #22) are excluded as beyond the scope, yet §III.C.3 performs a meta-analysis. PRISMA 2020 requires specifying the effect measure and assessing risk of bias for any quantitative synthesis; the excluded items are exactly the components needed for a valid meta-analysis. The reporting standard claimed in Section II is therefore not met for the meta-analytic portion. The authors should either include these PRISMA elements for the quantitative synthesis or remove the meta-analysis from the review and treat the comparison as a descriptive summary.
minor comments (4)
  1. [§III.B.3, Visual features] 'Hasen and Patil [40]' should be 'Hasan and Patil'; later in the same section 'incoprprating' should be 'incorporating'.
  2. [Table A1] Minor typos: 'Wav2cev2.0' should be 'Wav2Vec2.0', and 'datsset' should be 'dataset'.
  3. [§III.A.1.f] The citation to Fan et al. [34] appears mismatched: the text attributes to this work the finding that 'anticipation of sarcastic intent is crucial for the efficient comprehension of sarcasm,' but the reference is the 'exposure advantage' paper on multilingual exposure. Please verify and correct the citation.
  4. [§III.C.3.a] The attrition description ('19 articles ... excluded five ... and five ... resulting in nine articles') is confusing on first reading. A small inclusion-attrition table or a step-by-step count would improve clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the review and meta-analysis rest on external published results; the authors' own papers are non-load-bearing examples.

full rationale

The derivation chain of this paper is a systematic literature review plus a meta-analysis of published F1-scores, not a first-principles derivation. The central claims—that sarcasm recognition has moved from unimodal to multimodal fusion and that no single modality or feature set is sufficient—are supported by external studies (e.g., [21], [29], [35], [40]) and by the survey's own tabulated evidence. The meta-analysis in Section III.C.3 compares attention mechanisms with encoder-decoder fusion using F1-scores reported in external MUStARD-based studies (Tables III and IV); no fitted parameter is renamed as a prediction, and the null result is a statistical aggregation of those external numbers, not a consequence of the authors' own model outputs. The authors' self-citations ([42], [70]) appear in the narrative as examples of feature use, and [70] is excluded from the meta-analysis because it uses a train-test split rather than 5-fold cross-validation; [42] is a unimodal audio study and is not part of the fusion comparison. Thus the self-citations are not load-bearing. The paper itself flags limitations: it excludes PRISMA risk-of-bias items and effect measures ('Items related to risk of bias (#12, 15, 19, 22), effect measures (#13), additional analyses (#16, 23) are excluded'), and the meta-analysis is under-reported with no per-study effect sizes or variances, which is a methodological validity concern but not circularity. No equation or definition in the paper reduces a claimed result to its own inputs. Therefore the appropriate finding is no significant circularity, score 0.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The review introduces no new entities or fitted parameters. Its conclusions rest on the methodological assumptions of the systematic search and meta-analysis, which the authors document. The main unverified axiom is that the search captured all relevant work, and the main quantitative assumption is that the nine selected studies are sufficiently homogeneous for a meaningful pooled comparison.

assumptions (5)
  • domain assumption PRISMA 2020 guidelines are an appropriate methodological framework for this systematic review.
    Invoked in Section II; the review follows PRISMA but explicitly excludes risk-of-bias and effect-measure items (#12, 13, 15, 16, 19, 22, 23).
  • domain assumption The five selected databases (Scopus, IEEE Xplore, ISCA, ScienceDirect, ACM DL) and English-only filter capture the relevant literature.
    Section II search strategy; any omitted venues or non-English studies could change the synthesis.
  • domain assumption Zhao et al.'s fusion taxonomy (encoder-decoder, attention, gating, quantum) adequately classifies multimodal sarcasm architectures.
    Adopted in Section III.C.2 to organize fusion methods.
  • standard math Random-effects meta-analysis with Hedges' g, Q and I-squared is valid for pooling F1-scores across heterogeneous studies.
    Section III.C.3.b; standard meta-analytic approach, though F1-scores are bounded and study heterogeneity is high (I-squared > 80%).
  • domain assumption F1-score is a balanced metric appropriate for class-imbalanced sarcasm datasets.
    Section III.C.3.a justification for selecting studies reporting F1.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects." pith.science (2026). https://pith.science/paper/JG2KF7Z6

@misc{pith2026250904605,
  author       = {Pith},
  title        = {Pith review of: Spoken in Jest, Detected in Earnest: A Systematic Review of Sarcasm Recognition -- Multimodal Fusion, Challenges, and Future Prospects},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JG2KF7Z6}},
  note         = {Machine review of arXiv:2509.04605}
}
read the original abstract

Sarcasm, a common feature of human communication, poses challenges in interpersonal interactions and human-machine interactions. Linguistic research has highlighted the importance of prosodic cues, such as variations in pitch, speaking rate, and intonation, in conveying sarcastic intent. Although previous work has focused on text-based sarcasm detection, the role of speech data in recognizing sarcasm has been underexplored. Recent advancements in speech technology emphasize the growing importance of leveraging speech data for automatic sarcasm recognition, which can enhance social interactions for individuals with neurodegenerative conditions and improve machine understanding of complex human language use, leading to more nuanced interactions. This systematic review is the first to focus on speech-based sarcasm recognition, charting the evolution from unimodal to multimodal approaches. It covers datasets, feature extraction, and classification methods, and aims to bridge gaps across diverse research domains. The findings include limitations in datasets for sarcasm recognition in speech, the evolution of feature extraction techniques from traditional acoustic features to deep learning-based representations, and the progression of classification methods from unimodal approaches to multimodal fusion techniques. In so doing, we identify the need for greater emphasis on cross-cultural and multilingual sarcasm recognition, as well as the importance of addressing sarcasm as a multimodal phenomenon, rather than a text-based challenge.

Figures

Figures reproduced from arXiv: 2509.04605 by the authors.

Figure 2
Figure 2. Trend of the publications. “Multimodal” approaches include audio [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 1
Figure 1. PRISMA-guided systematic review process. The initial search phase [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 4
Figure 4. Distribution of datasets by source, language, and modality. [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗
Figures from the paper (3 more)
Figure 5
Figure 5. Figure 5: Hierarchical taxonomy of the most prominent features. [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: The overview of classification strategies in unimodal sarcasm [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: The distribution of current multimodal fusion strategies. [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

110 extracted references · 72 canonical work pages

  1. [1]

    Irony and use-mention distinction,

    D. Sperber and D. Wilson, “Irony and use-mention distinction,” in Radical Pragmatics , P. Cole, Ed. New York, NY , USA: Academic Press, 1981, pp. 295–318

  2. [2]

    The functions of sarcastic irony in speech,

    J. Jorgensen, “The functions of sarcastic irony in speech,” J. Pragmat., vol. 26, no. 5, pp. 613–634, Nov. 1996

  3. [3]

    Brown and S

    P. Brown and S. C. Levinson, Politeness: Some Universals in Language Usage. Cambridge, UK: Cambridge University Press, 1987

  4. [4]

    Jocularity, sarcasm, and relation- ships: An empirical study,

    M. A. Seckman and C. J. Couch, “Jocularity, sarcasm, and relation- ships: An empirical study,” J. Contemp. Ethnogr. , vol. 18, no. 3, pp. 327–344, July 1989

  5. [5]

    On irony and negation,

    R. Giora, “On irony and negation,” Discourse Process., vol. 19, no. 2, pp. 239–264, Nov. 1995

  6. [6]

    Exploring the process of inference generation in sar- casm: A review of normal and clinical studies,

    S. McDonald, “Exploring the process of inference generation in sar- casm: A review of normal and clinical studies,” Brain Lang., vol. 68, no. 3, pp. 486–506, July 1999

  7. [7]

    Logic and conversation,

    H. P. Grice, “Logic and conversation,” in Syntax and Semantics, Vol. 3: Speech Acts, P. Cole and J. L. Morgan, Eds. New York, NY , USA: Academic Press, 1975, pp. 41–58

  8. [8]

    Test of the mention theory of irony,

    J. Jorgensen, G. A. Miller, and D. Sperber, “Test of the mention theory of irony,” J. Exp. Psychol. Gen. , vol. 113, no. 1, pp. 112–120, 1984

Show all 110 references
  1. [9]

    Understanding social dysfunction in the behavioural variant of frontotemporal dementia: The role of emotion and sarcasm processing,

    C. M. Kipps, P. J. Nestor, J. Acosta-Cabronero, R. Arnold, and J. R. Hodges, “Understanding social dysfunction in the behavioural variant of frontotemporal dementia: The role of emotion and sarcasm processing,” Brain, vol. 132, no. 3, pp. 592–603, Jan. 2009

  2. [10]

    Teaching children with autism to detect and respond to sarcasm,

    A. Persicke, J. Tarbox, J. Ranick, and M. S. Clair, “Teaching children with autism to detect and respond to sarcasm,” Res. Autism Spectr. Disord., vol. 7, no. 1, pp. 193–198, Jan. 2013

  3. [11]

    Lower, slower, louder: V ocal cues of sarcasm,

    P. Rockwell, “Lower, slower, louder: V ocal cues of sarcasm,” J. Psycholinguist. Res., vol. 29, no. 5, pp. 483–495, Sep. 2000

  4. [12]

    Acoustic markers of sarcasm in cantonese and english,

    H. S. Cheang and M. D. Pell, “Acoustic markers of sarcasm in cantonese and english,” J. Acoust. Soc. Am. , vol. 126, no. 3, pp. 1394– 1405, Sep. 2009

  5. [13]

    V oice modulations in german ironic speech,

    L. Scharrer and U. Christmann, “V oice modulations in german ironic speech,” Lang. Speech, vol. 54, no. 4, pp. 435–465, Dec. 2011. xviii

  6. [14]

    Prosodic cues of sarcastic speech in french: Slower, higher, wider,

    H. Lœvenbruck, M. A. B. Jannet, M. D’Imperio, M. Spini, and M. Champagne-Lavau, “Prosodic cues of sarcastic speech in french: Slower, higher, wider,” in Proc. Interspeech, Lyon, France, Aug. 2013, pp. 3537–3541

  7. [15]

    ‘yeah right’: Sarcasm recognition for spoken dialogue systems,

    J. Tepperman, D. Traum, and S. Narayanan, “‘yeah right’: Sarcasm recognition for spoken dialogue systems,” in Proc. Interspeech, Pitts- burgh, PA, USA, Sep. 2006, pp. 1838–1841

  8. [16]

    ‘sure, i did the right thing’: A system for sarcasm detection in speech,

    R. Rakov and A. Rosenberg, “‘sure, i did the right thing’: A system for sarcasm detection in speech,” in Proc. Interspeech, Lyon, France, Aug. 2013, pp. 842–846

  9. [17]

    Automatic sarcasm detection: A survey,

    A. Joshi, P. Bhattacharyya, and M. J. Carman, “Automatic sarcasm detection: A survey,” ACM Comput. Surv. , vol. 50, no. 5, pp. 1–22, Sep. 2017

  10. [18]

    Sarcasm identification in textual data: Systematic review, research challenges and open directions,

    C. I. Eke, A. A. Norman, L. Shuib, and H. F. Nweke, “Sarcasm identification in textual data: Systematic review, research challenges and open directions,” Artif. Intell. Rev., vol. 53, no. 6, pp. 4215–4258, Nov. 2020

  11. [19]

    Literature survey of sarcasm detection,

    P. Chaudhari and C. Chandankhede, “Literature survey of sarcasm detection,” in Proc. 2017 Int. Conf. Wireless Commun., Signal Process. Netw. (WiSPNET), Chennai, India, Mar. 2017, pp. 2041–2046

  12. [20]

    Pre- ferred reporting items for systematic reviews and meta-analyses: The prisma statement,

    D. Moher, A. Liberati, J. Tetzlaff, D. G. Altman, and T. P. Group, “Pre- ferred reporting items for systematic reviews and meta-analyses: The prisma statement,” PLOS Med., vol. 6, no. 7, 2009, doi: 10.1371/jour- nal.pmed.1000097

  13. [21]

    Towards multimodal sarcasm detection (an obviously perfect paper),

    S. Castro, D. Hazarika, V . P ´erez-Rosas, R. Zimmermann, R. Mihalcea, and S. Poria, “Towards multimodal sarcasm detection (an obviously perfect paper),” in Proc. 57th Annu. Meeting Assoc. Comput. Linguis- tics (ACL), Florence, Italy, Aug. 2019, pp. 4619–4629

  14. [22]

    IITKGP-SEHSC: Hindi speech corpus for emotion analysis,

    S. G. Koolagudi, R. Reddy, J. Yadav, and K. S. Rao, “IITKGP-SEHSC: Hindi speech corpus for emotion analysis,” in Proc. 2011 Int. Conf. Devices Commun. (ICDeCom) , Mesra, Ranchi, India, Feb. 2011, pp. 1–5

  15. [23]

    Multi-modal sarcasm detection and humor classification in code-mixed conversa- tions,

    M. Bedi, S. Kumar, M. S. Akhtar, and T. Chakraborty, “Multi-modal sarcasm detection and humor classification in code-mixed conversa- tions,” IEEE Trans. Affective Comput. , vol. 14, no. 2, pp. 1363–1375, 2023

  16. [24]

    ¡Qu´e maravilla! multimodal sarcasm detection in spanish: A dataset and a baseline,

    K. Alnajjar and M. H ¨am¨al¨ainen, “¡Qu´e maravilla! multimodal sarcasm detection in spanish: A dataset and a baseline,” in Proc. 3rd Workshop Multimodal Artif. Intell. , Mexico City, Mexico, Jun. 2021, pp. 63–68

  17. [25]

    CMMA: Benchmarking multi-affection detection in Chinese multi-modal conversations,

    Y . Zhang et al. , “CMMA: Benchmarking multi-affection detection in Chinese multi-modal conversations,” in Proc. 37th Int. Conf. Neural Information Processing Systems (NeurIPS) , New Orleans, LA, USA, 2023, pp. 18 794–18 805, [Online]. Available: https://openreview.net/f orum?...

  18. [26]

    Deep learning for acous- tic irony classification in spontaneous speech,

    H. Gent, C. Adams, C. Shih, and Y . Tang, “Deep learning for acous- tic irony classification in spontaneous speech,” in Proc. Interspeech, Incheon, South Korea, Sep. 2022, pp. 3993–3997

  19. [27]

    Sentiment and emotion help sarcasm? a multi-task learning framework for multi- modal sarcasm, sentiment, and emotion analysis,

    D. S. Chauhan, D. S. R, A. Ekbal, and P. Bhattacharyya, “Sentiment and emotion help sarcasm? a multi-task learning framework for multi- modal sarcasm, sentiment, and emotion analysis,” in Proc. 58th Annu. Meet. Assoc. Comput. Linguistics (ACL) , Seattle, W A, USA, Jul. 2020, p...

  20. [28]

    An emoji-aware multitask framework for multimodal sarcasm detection,

    D. S. Chauhan, G. Singh, A. Arora, A. Ekbal, and P. Bhat- tacharyya, “An emoji-aware multitask framework for multimodal sarcasm detection,” Knowl.-Based Syst. , vol. 257, 2022, doi: 10.1016/j.knosys.2022.109924

  21. [29]

    A multimodal corpus for emotion recognition in sarcasm,

    A. Ray, S. Mishra, A. Nunna, and P. Bhattacharyya, “A multimodal corpus for emotion recognition in sarcasm,” in Proc. 13th Lang. Resour. Eval. Conf. (LREC) , Marseille, France, June 2022, pp. 6992–7003

  22. [30]

    MELD: A multimodal multi-party dataset for emotion recognition in conversations,

    S. Poria, D. Hazarika, N. Majumder, G. Naik, E. Cambria, and R. Mihalcea, “MELD: A multimodal multi-party dataset for emotion recognition in conversations,” in Proc. 57th Annu. Meet. Assoc. Com- put. Linguistics (ACL) , Florence, Italy, Jul. 2019, pp. 527–536

  23. [31]

    Which contextual and socio- cultural information predict irony perception?

    E. Rivi `ere and M. Champagne-Lavau, “Which contextual and socio- cultural information predict irony perception?” Discourse Process. , vol. 57, no. 3, pp. 259–277, 2020

  24. [32]

    Verbal irony as implicit display of ironic environment: Distinguishing ironic utterances from nonirony,

    A. Utsumi, “Verbal irony as implicit display of ironic environment: Distinguishing ironic utterances from nonirony,” J. Pragmat., vol. 32, no. 12, pp. 1777–1806, Nov. 2000

  25. [33]

    How about another piece of pie: The allusional pretense theory of discourse irony,

    S. Kumon-Nakamura, S. Glucksberg, and M. Brown, “How about another piece of pie: The allusional pretense theory of discourse irony,” J. Exp. Psychol. Gen. , vol. 124, no. 1, pp. 3–21, 1995

  26. [34]

    The exposure advantage: Early exposure to a multilingual environment promotes effective communication,

    S. P. Fan, Z. Liberman, B. Keysar, and K. D. Kinzler, “The exposure advantage: Early exposure to a multilingual environment promotes effective communication,” Psychol. Sci., vol. 26, no. 7, pp. 1090–1097, Jul. 2015

  27. [35]

    Modeling incongruity between modalities for multimodal sarcasm detection,

    Y . Wu et al., “Modeling incongruity between modalities for multimodal sarcasm detection,” IEEE MultiMedia, vol. 28, no. 2, pp. 86–95, Apr. 2021

  28. [36]

    The coding strategy for the mandarin speech conveying sarcasm in acoustic and articulatory domain,

    P. Geng, S. Shi, and H. Guo, “The coding strategy for the mandarin speech conveying sarcasm in acoustic and articulatory domain,” in Proc. 5th Int. Conf. Digit. Signal Process. (DSP) , Chengdu, China, Feb. 2021, pp. 195–200

  29. [37]

    Detecting vocal irony,

    F. Burkhardt, B. Weiss, F. Eyben, J. Deng, and B. Schuller, “Detecting vocal irony,” in Proc. 27th Int. Conf. Lang. Technol. Challenges Digit. Age (GSCL), vol. 10713, 2018, pp. 11–22

  30. [38]

    Emotional vocal expressions recognition using the cost 2102 italian database of emotional speech,

    H. Atassi, M. T. Riviello, Z. Sm ´ekal, A. Hussain, and A. Esposito, “Emotional vocal expressions recognition using the cost 2102 italian database of emotional speech,” in Development of Multimodal Inter- faces: Active Listening and Synchrony . Berlin, Germany: Springer, 2010,...

  31. [39]

    Emotion recognition in speech using machine learning techniques,

    A. Arun, I. Rallabhandi, S. Hebbar, A. Nair, and R. Jayashree, “Emotion recognition in speech using machine learning techniques,” in Proc. 12th Int. Conf. Comput. Commun. Netw. Technol. (ICCCNT) , Kharagpur, India, Jul. 2021, pp. 1–7

  32. [40]

    Humor knowledge enriched transformer for understanding multimodal humor,

    M. K. Hasan et al. , “Humor knowledge enriched transformer for understanding multimodal humor,” in Proc. AAAI Conf. Artif. Intell. , vol. 35, no. 14, May 2021, pp. 12 972–12 980

  33. [41]

    Measuring nominal scale agreement among many raters,

    J. L. Fleiss, “Measuring nominal scale agreement among many raters,” Psychol. Bull., vol. 76, no. 5, pp. 378–382, Nov. 1971

  34. [42]

    Deep cnn-based inductive trans- fer learning for sarcasm detection in speech,

    X. Gao, S. Nayak, and M. Coler, “Deep cnn-based inductive trans- fer learning for sarcasm detection in speech,” in Proc. Interspeech , Incheon, South Korea, Sep. 2022, pp. 2323–2327

  35. [43]

    Emotion recognition of speech,

    N. S. Sacheth and R. Jayashree, “Emotion recognition of speech,” in Proc. 14th Int. Conf. Soft Comput. Pattern Recognit. (SoCPaR) . Springer, 2023, vol. 490, pp. 359–371

  36. [44]

    Characterization of emotions using the dynamics of prosodic features,

    K. S. Rao, R. Reddy, S. Maity, and S. G. Koolagudi, “Characterization of emotions using the dynamics of prosodic features,” in Proc. Speech Prosody, Chicago, IL, USA, May 2010, [Online]. Available: https://ww w.isca-archive.org/speechprosody 2010/rao10 speechprosody.html

  37. [45]

    Understanding sarcasm in speech using mel-frequency cepstral coefficient,

    A. Mathur, V . Saxena, and S. K. Singh, “Understanding sarcasm in speech using mel-frequency cepstral coefficient,” in Proc. 7th Int. Conf. Cloud Comput., Data Sci. Eng. (Confluence) , Noida, India, Jan. 2017, pp. 728–732

  38. [46]

    The interspeech 2013 computational paralinguistics challenge: Social signals, conflict, emotion, autism,

    B. Schuller et al., “The interspeech 2013 computational paralinguistics challenge: Social signals, conflict, emotion, autism,” in Proc. Inter- speech, Lyon, France, Aug. 2013, pp. 148–152

  39. [47]

    Multi-modal sarcasm detection based on contrastive attention mechanism,

    X. Zhang, Y . Chen, and G. Li, “Multi-modal sarcasm detection based on contrastive attention mechanism,” inProc. Nat. Lang. Process. Chin. Comput., vol. 13028, 2021, pp. 822–833

  40. [48]

    Multimodal sarcasm detection: A deep learning approach,

    S. K. Bharti, R. Gupta, P. K. Shukla, W. A. Hatamleh, H. Tarazi, and S. J. Nuagah, “Multimodal sarcasm detection: A deep learning approach,” Wirel. Commun. Mob. Comput., vol. 2022, pp. 1–10, 2022, doi: 10.1155/2022/1653696

  41. [49]

    A multimodal fusion method for sarcasm detection based on late fusion,

    N. Ding, S.-W. Tian, and L. Yu, “A multimodal fusion method for sarcasm detection based on late fusion,” Multimed. Tools Appl., vol. 81, no. 6, pp. 8597–8616, Feb. 2022

  42. [50]

    EFAFN: An efficient feature adaptive fusion network with facial feature for multimodal sarcasm detection,

    Y . Sun, H. Zhang, S. Yang, and J. Wang, “EFAFN: An efficient feature adaptive fusion network with facial feature for multimodal sarcasm detection,” Appl. Sci. , vol. 12, no. 21, p. 11235, Nov. 2022, doi: 10.3390/app122111235

  43. [51]

    Multimodal learning using optimal transport for sarcasm and humor detection,

    S. Pramanick, A. Roy, and V . M. Patel, “Multimodal learning using optimal transport for sarcasm and humor detection,” inProc. IEEE/CVF Winter Conf. Appl. Comput. Vis. (WACV) , Waikoloa, HI, USA, Jan. 2022, pp. 546–556

  44. [52]

    Multimodal sarcasm detection (MSD) in videos using deep learning models,

    A. Pandey and D. K. Vishwakarma, “Multimodal sarcasm detection (MSD) in videos using deep learning models,” in Proc. 2023 Int. Conf. Adv. Power, Signal, Inf. Technol. (APSIT), 2023, pp. 811–814

  45. [53]

    Multi- modal sarcasm detection method using rnn and cnn,

    B. Azahouani, H. Elfaik, E. H. Nfaoui, and S. E. Garouani, “Multi- modal sarcasm detection method using rnn and cnn,” in Proc. 6th Int. Conf. Intell. Comput. Data Sci. (ICDS) , Marrakech, Morocco, 2024, pp. 1–6

  46. [54]

    An attention-based, context-aware multimodal fusion method for sarcasm detection using inter-modality inconsistency,

    Y . Li et al. , “An attention-based, context-aware multimodal fusion method for sarcasm detection using inter-modality inconsistency,” Knowl.-Based Syst., vol. 287, p. 111457, Mar. 2024

  47. [55]

    A smart video analytical frame- work for sarcasm detection using novel adaptive fusion network and sarcasnet-99 model,

    J. S. Murthy and G. M. Siddesh, “A smart video analytical frame- work for sarcasm detection using novel adaptive fusion network and sarcasnet-99 model,” Vis. Comput. , vol. 40, no. 11, pp. 8085–8097, Nov. 2024

  48. [56]

    V oice-based sarcasm detection in kannada language,

    R. Manohar and S. Swamy, “V oice-based sarcasm detection in kannada language,” Int. J. Intell. Syst. Appl. Eng., vol. 12, no. 14s, pp. 356–367, Feb. 2024, [Online]. Available: https://ijisae.org/index.php/IJISAE/arti cle/view/4672. xix

  49. [57]

    Sarcasm detection using cognitive features of visual data by learning model,

    B. N. Hiremath and M. M. Patil, “Sarcasm detection using cognitive features of visual data by learning model,” Expert Syst. Appl., vol. 184, Dec. 2021

  50. [58]

    Cnn architectures for large-scale audio classifi- cation,

    S. Hershey et al. , “Cnn architectures for large-scale audio classifi- cation,” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), New Orleans, LA, USA, Mar. 2017, pp. 131–135

  51. [59]

    wav2vec 2.0: A framework for self-supervised learning of speech representations,

    A. Baevski, H. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representations,” Adv. Neural Inf. Process. Syst. , vol. 33, pp. 12 449–12 460, 2020

  52. [60]

    A multitask learning model for multimodal sarcasm, sentiment and emotion recognition in conversations,

    Y . Zhang et al., “A multitask learning model for multimodal sarcasm, sentiment and emotion recognition in conversations,” Inf. Fusion , vol. 93, pp. 282–301, May 2023

  53. [61]

    Attention is all you need,

    A. Vaswani et al., “Attention is all you need,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS) , vol. 30, Long Beach, CA, USA, Dec. 2017, pp. 5998–6008

  54. [62]

    Learning phrase representations using rnn encoder–decoder for statistical machine translation,

    K. Cho et al. , “Learning phrase representations using rnn encoder–decoder for statistical machine translation,” arXiv preprint arXiv:1406.1078, 2014, [Online]. Available: https://arxiv.org/abs/14 06.1078

  55. [63]

    PRAAT, a system for doing phonetics by computer,

    P. Boersma, “PRAAT, a system for doing phonetics by computer,” Glot Int., vol. 5, no. 9/10, pp. 341–345, 2001

  56. [64]

    OpenSMILE: The munich versatile and fast open-source audio feature extractor,

    F. Eyben, M. W ¨ollmer, and B. Schuller, “OpenSMILE: The munich versatile and fast open-source audio feature extractor,” in Proc. 18th ACM Int. Conf. Multimedia, Florence, Italy, Oct. 2010, pp. 1459–1462

  57. [65]

    Librosa: Audio and music signal analysis in python,

    B. McFee et al., “Librosa: Audio and music signal analysis in python,” in Proc. Python in Sci. Conf. (SciPy) , Austin, TX, USA, 2015, pp. 18– 24, doi: 10.25080/Majora-7b98e3ed-003

  58. [66]

    CO- V AREP: A collaborative voice analysis repository for speech tech- nologies,

    G. Degottex, J. Kane, T. Drugman, T. Raitio, and S. Scherer, “CO- V AREP: A collaborative voice analysis repository for speech tech- nologies,” in Proc. IEEE Int. Conf. Acoust., Speech Signal Process. (ICASSP), Florence, Italy, May 2014, pp. 960–964

  59. [67]

    BERT: Pre- training of deep bidirectional transformers for language understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre- training of deep bidirectional transformers for language understanding,” in Proc. Conf. North Amer. Chapter Assoc. Comput. Linguistics: Human Lang. Technol. (NAACL-HLT) , Minneapolis, MN, USA, Jun. 2019, pp. 4171–4186

  60. [68]

    A quantum probability driven frame- work for joint multi-modal sarcasm, sentiment and emotion analysis,

    Y . Liu, Y . Zhang, and D. Song, “A quantum probability driven frame- work for joint multi-modal sarcasm, sentiment and emotion analysis,” IEEE Trans. Affect. Comput. , pp. 1–15, 2023

  61. [69]

    Learning multitask commonness and uniqueness for multimodal sarcasm detection and sentiment analysis in conversation,

    Y . Zhang et al., “Learning multitask commonness and uniqueness for multimodal sarcasm detection and sentiment analysis in conversation,” IEEE Trans. Artif. Intell. , vol. 5, no. 3, pp. 1349–1361, 2024

  62. [70]

    Improving sarcasm detection from speech and text through attention-based fusion exploiting the interplay of emotions and sentiments,

    X. Gao, S. Nayak, and M. Coler, “Improving sarcasm detection from speech and text through attention-based fusion exploiting the interplay of emotions and sentiments,” in Proc. Meet. Acoust. , vol. 54, no. 1, Jan. 2024, p. 060002, doi:10.1121/2.0001918

  63. [71]

    BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,

    M. Lewis et al., “BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension,” in Proc. 58th Annu. Meet. Assoc. Comput. Linguistics (ACL) , Online, Jul. 2020, pp. 7871–7880

  64. [72]

    Your tone speaks louder than your face! modality order infused multi-modal sarcasm detection,

    M. Tomar, A. Tiwari, T. Saha, and S. Saha, “Your tone speaks louder than your face! modality order infused multi-modal sarcasm detection,” in Proc. 31st ACM Int. Conf. Multimedia (MM) , Ottawa, ON, Canada, Oct. 2023, pp. 3926–3933

  65. [73]

    Quantum fuzzy neural network for multimodal sentiment and sarcasm detection,

    P. Tiwari, L. Zhang, Z. Qu, and G. Muhammad, “Quantum fuzzy neural network for multimodal sentiment and sarcasm detection,” Inf. Fusion, vol. 97, p. 102085, Jan. 2024

  66. [74]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), Las Vegas, NV , USA, Jun. 2016, pp. 770–778

  67. [75]

    EfficientNet: Rethinking model scaling for convolutional neural networks,

    M. Tan and Q. V . Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proc. 36th Int. Conf. Mach. Learn. (ICML), vol. 97, 2019, pp. 6105–6114

  68. [76]

    Joint face detection and alignment using multi-task cascaded convolutional networks,

    K. Zhang, Z. Zhang, Z. Li, and Y . Qiao, “Joint face detection and alignment using multi-task cascaded convolutional networks,” IEEE Signal Process. Lett. , vol. 23, no. 10, pp. 1499–1503, 2016

  69. [77]

    Openface 2.0: Facial behavior analysis toolkit,

    T. Baltrusaitis, A. Zadeh, Y . C. Lim, and L.-P. Morency, “Openface 2.0: Facial behavior analysis toolkit,” in Proc. 13th IEEE Int. Conf. Autom. Face Gesture Recognit. (FG) , Xi’an, China, May 2018, pp. 59–66

  70. [78]

    Quo vadis, action recognition? a new model and the kinetics dataset,

    J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” in Proc. IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR), 2017, pp. 4724–4733

  71. [79]

    Multimodal markers of irony and sarcasm,

    S. Attardo, J. Eisterhold, J. Hay, and I. Poggi, “Multimodal markers of irony and sarcasm,” Humor - Int. J. Humor Res. , vol. 16, no. 2, pp. 243–260, Jan. 2003

  72. [80]

    Verbal irony use in face-to-face and computer-mediated conversations,

    J. T. Hancock, “Verbal irony use in face-to-face and computer-mediated conversations,” J. Lang. Soc. Psychol. , vol. 23, no. 4, pp. 447–463, 2004

  73. [81]

    Effects of emotional intelligence on the impression of irony created by the mismatch between verbal and nonverbal cues,

    J. H., B. Kreifelts, S. Nizielski, A. Sch ¨utz, and D. Wildgruber, “Effects of emotional intelligence on the impression of irony created by the mismatch between verbal and nonverbal cues,” PLOS ONE , vol. 11, no. 10, p. e0163211, 2016, doi: 10.1371/journal.pone.0163211

  74. [82]

    Deep multimodal data fusion,

    F. Zhao, C. Zhang, and B. Geng, “Deep multimodal data fusion,” ACM Comput. Surv., vol. 56, no. 9, pp. 1–36, Oct. 2024

  75. [83]

    L. V . Hedges and I. Olkin, Statistical Methods for Meta-Analysis . Academic Press, 1985

  76. [84]

    Quantifying heterogeneity in a meta-analysis,

    J. P. Higgins and S. G. Thompson, “Quantifying heterogeneity in a meta-analysis,” Stat. Med., vol. 21, no. 11, pp. 1539–1558, Jun. 2002

  77. [85]

    Meta-analysis in clinical trials,

    R. DerSimonian and N. Laird, “Meta-analysis in clinical trials,” Con- trol. Clin. Trials, vol. 7, no. 3, pp. 177–188, Sep. 1986

  78. [86]

    The cost 2102 italian audio and video emotional database,

    A. Esposito, M. T. Riviello, and G. Maio, “The cost 2102 italian audio and video emotional database,” in Front. Artif. Intell. Appl., Jan. 2009, vol. 204, pp. 51–61

  79. [87]

    Sarcasm, pretense, and the semantics/pragmatics distinc- tion,

    E. Camp, “Sarcasm, pretense, and the semantics/pragmatics distinc- tion,” Noˆus, vol. 46, no. 4, pp. 587–634, Dec. 2012

  80. [88]

    WavLM: Large-scale self-supervised pre-training for full stack speech processing,

    S. Chen et al. , “WavLM: Large-scale self-supervised pre-training for full stack speech processing,” IEEE J. Sel. Top. Signal Process., vol. 16, no. 6, pp. 1505–1518, Oct. 2022

  81. [89]

    SpeechLM: Enhanced speech pre-training with unpaired textual data,

    Z. Zhang et al. , “SpeechLM: Enhanced speech pre-training with unpaired textual data,” IEEE/ACM Trans. Audio Speech Lang. Process., vol. 32, pp. 2177–2187, 2024

  82. [90]

    Adapting WavLM for speech emotion recognition,

    D. Diatlova, A. Udalov, V . Shutov, and E. Spirin, “Adapting WavLM for speech emotion recognition,” in Proc. Odyssey 2024: The Speaker and Language Recognition Workshop , 2024, pp. 303–308

  83. [91]

    GPT-4 technical report,

    J. Achiam et al. , “GPT-4 technical report,” arXiv preprint arXiv:2303.08774, 2023, [Online]. Available: https://doi.org/10.485 50/arXiv.2303.08774

  84. [92]

    LLaMA: Open and efficient foundation language models,

    H. Touvron et al. , “LLaMA: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023, [Online]. Available: https://doi.org/10.48550/arXiv.2302.13971

  85. [93]

    SarcasmBench: Towards evaluating large language models on sarcasm understanding,

    Y . Zhang, C. Zou, Z. Lian, P. Tiwari, and J. Qin, “SarcasmBench: Towards evaluating large language models on sarcasm understanding,” arXiv preprint arXiv:2408.11319, 2024, [Online]. Available: https://do i.org/10.48550/arXiv.2408.11319

  86. [94]

    Learning transferable visual models from natural language supervision,

    A. Radford et al. , “Learning transferable visual models from natural language supervision,” arXiv preprint arXiv:2103.00020, 2021, [On- line]. Available: https://doi.org/10.48550/arXiv.2103.00020

  87. [95]

    Paraverbal expression of verbal irony: V ocal cues matter and facial cues even more,

    M. Aguert, “Paraverbal expression of verbal irony: V ocal cues matter and facial cues even more,” J. Nonverbal Behav. , vol. 46, no. 1, pp. 45–70, Mar. 2022

  88. [96]

    Tracking nonliteral language processing using audiovisual scenarios,

    K. Rothermich et al. , “Tracking nonliteral language processing using audiovisual scenarios,” Can. J. Exp. Psychol. , vol. 75, no. 2, pp. 211– 220, 2021

  89. [97]

    The role of auditory and visual cues in the interpretation of mandarin ironic speech,

    S. Li, A. Chen, Y . Chen, and P. Tang, “The role of auditory and visual cues in the interpretation of mandarin ironic speech,” J. Pragmat., vol. 201, pp. 3–14, 2022

  90. [98]

    Modality matters: Testing bilingual irony comprehension in the textual, auditory, and audio-visual modality,

    K. Bromberek-Dyzman, K. Jankowiak, and P. Chełminiak, “Modality matters: Testing bilingual irony comprehension in the textual, auditory, and audio-visual modality,” J. Pragmat., vol. 180, pp. 219–231, 2021

  91. [99]

    Context, facial expression and prosody in irony processing,

    G. Deliens, K. Antoniou, E. Clin, E. Ostashchenko, and M. Kissine, “Context, facial expression and prosody in irony processing,” J. Mem. Lang., vol. 99, pp. 35–48, 2018

  92. [100]

    The sound of sarcasm,

    H. S. Cheang and M. D. Pell, “The sound of sarcasm,” Speech Commun., vol. 50, no. 5, pp. 366–381, May 2008

  93. [101]

    From ‘blame by praise’ to ‘praise by blame’: Analysis of vocal patterns in ironic communication,

    L. Anolli, R. Ciceri, and M. G. Infantino, “From ‘blame by praise’ to ‘praise by blame’: Analysis of vocal patterns in ironic communication,” Int. J. Psychol. , vol. 37, no. 5, pp. 266–276, Oct. 2002

  94. [102]

    Categorical and schematic organization in memory,

    J. M. Mandler, “Categorical and schematic organization in memory,” in Memory, Organization, and Structure , R. C. Puff, Ed. New York, NY , USA: Academic Press, 1979, pp. 259–299

  95. [103]

    How Korean EFL learners understand sarcasm in L2 English,

    J. Kim, “How Korean EFL learners understand sarcasm in L2 English,” J. Pragmat., vol. 60, pp. 193–206, 2014

  96. [104]

    Self- adaptive representation learning model for multi-modal sentiment and sarcasm joint analysis,

    Y . Zhang, Y . Yu, M. Wang, M. Huang, and M. S. Hossain, “Self- adaptive representation learning model for multi-modal sentiment and sarcasm joint analysis,” ACM Trans. Multimedia Comput. Commun. Appl., vol. 20, no. 5, May 2024

  97. [105]

    Modified convolutional neural network with multiple features for multimodal sarcasm detection,

    R. Krishnamaneni, M. Kurni, S. Sen, and A. Murthy, “Modified convolutional neural network with multiple features for multimodal sarcasm detection,” in Proc. 2nd Int. Conf. Recent Adv. Inf. Technol. Sustain. Dev. (ICRAIS), Manipal, India, 2024, pp. 160–165. xx

  98. [106]

    The geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing,

    F. Eyben et al. , “The geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing,” IEEE Trans. Affect. Comput., vol. 7, no. 2, pp. 190–202, 2016

  99. [107]

    Understanding non-verbal irony markers: Machine learning insights versus human judgment,

    M. Spitale, F. Catania, and F. Panzeri, “Understanding non-verbal irony markers: Machine learning insights versus human judgment,” in Proc. 26th ACM Int. Conf. Multimodal Interact. (ICMI), Nov. 2024, pp. 164– 172

  100. [108]

    SWITCHBOARD: Telephone speech corpus for research and development,

    J. J. Godfrey, E. C. Holliman, and J. McDaniel, “SWITCHBOARD: Telephone speech corpus for research and development,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP) , vol. 1, San Francisco, CA, USA, Mar. 1992, pp. 517–520

  101. [109]

    The Fisher corpus: A resource for the next generations of speech-to-text,

    C. Cieri, D. Miller, and K. Walker, “The Fisher corpus: A resource for the next generations of speech-to-text,” in Proc. 4th Int. Conf. Lang. Resour. Eval. (LREC) , Lisbon, Portugal, May 2004, pp. 69–71, [Online]. Available: https://aclanthology.org/L04-1500/. Xiyuan Gao recei...

  102. [2010]

    In 2012, he returned to academia

    After a postdoctoral appointment at his alma mater, he joined an AI start-up focusing on acoustic sensors, where he served as Head of the Cognitive Systems Unit. In 2012, he returned to academia. Dr. Coler is currently an Associate Professor of Speech Technology at the Faculty...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.