Pith. sign in

REVIEW 3 major objections 6 minor 300 references

Towards High-Level Semantic Intelligence

T0 review · 3 major / 6 minor · reviewed 2026-07-31 · deepseek-v4-flash

Pith's one-line read The paper argues that current AI has largely mastered literal, perception-grounded semantics and is now moving into high-level semantic intelligence: understanding and generating humor, sarcasm, metaphor, empathy, persuasion, and narrative

desk verdict A useful, well-organized survey that is agenda-setting rather than formal; the BLSI/HLSI boundary is asserted and the complexity equations don't yet do work, but the reference map and task taxonomy are genuinely valuable. read the letter →

arxiv 2607.24082 v1 pith:WBJXNTF3 submitted 2026-07-27 cs.AI

classification cs.AI
keywords high-levelsemanticssemanticcomplexityintelligencehumorunderstandingfigurativelanguagemultimodalreasoningsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that AI's progress can be read as a shift from Basic-Level Semantic Intelligence (BLSI) to High-Level Semantic Intelligence (HLSI): from literal recognition and production to meanings that require inference, context, and social grounding. It argues that six seemingly separate tasks—humor, sarcasm, metaphor, empathy, persuasion, and narrative—should be treated as one coherent research direction, unified by a formal language of semantic clues, chains, causal structure, and semantic effects. If true, the field gains a shared vocabulary for comparing methods and benchmarks across text, speech, vision, and multimodal settings, and a common target for future models. The survey organizes existing data, modeling, and evaluation work for HLS understanding and generation under this umbrella.

What carries the argument

The framework centers on four formal elements: Semantic Clues (perceptual evidence), Semantic Chains (inference from clues to derived meaning, called 'Semantic Level-Up'), Causal Chains (causal relations among factors and effects), and Semantic Effects (the holistic cognitive or affective result of composing chains). These feed two proposed metrics: Semantic Complexity Cs, proportional to the complexity of the semantic and causal chains, and Semantic Density Ds = Cs/|R|, the complexity packed into a physical representation. The work that this machinery does is to give the BLSI/HLSI boundary a common descriptive language, and to justify treating humor, sarcasm, metaphor, empathy, persuasion,

What would settle it

Take paired literal and sarcastic utterances matched in token count: if the paper's metric is real, the sarcastic one should have higher semantic density Ds. Since no procedure for computing Cs is supplied, the prediction cannot be evaluated. A concrete test would be to provide an algorithm for Complexity(SC,CC), compute Ds across a benchmark, and show a threshold that separates BLS from HLS; absent that, the formal apparatus is decorative.

Watch

Extended reading notes

Core claim

The central claim is that AI's semantic evolution follows a clear trajectory: early systems handled direct, literal, perception-grounded meanings, while current systems are increasingly expected to handle abstract, context-dependent, socially situated meanings. The paper names these two levels Basic-Level Semantic Intelligence (BLSI) and High-Level Semantic Intelligence (HLSI), and places six task families—humor, sarcasm, metaphor, empathy, persuasion, and narrative—under HLSI. It proposes that all six share underlying structures: semantic clues, semantic chains, causal chains, and semantic effects, with semantic complexity and semantic density as quantitative characterizations. The authors

Load-bearing premise

The whole BLSI/HLSI stage distinction rests on the claim that semantic complexity is a measurable quantity, but the paper never defines how to compute it; if complexity cannot be measured, the boundary between basic and high-level semantics is a label, not a discovery.

Editorial extensions

If this is right

  • If HLS is one coherent direction, progress on one task—say metaphor—can transfer to others through shared semantic mechanisms and evaluation tools.
  • If the BLSI/HLSI transition is real, benchmarks should measure semantic complexity and density, not just label accuracy, and should separate perceptual clue recognition from higher-level semantic reasoning.
  • If current large models sit near the BLS/HLS boundary, failures in humor, sarcasm, empathy, persuasion, and narrative point to a common reasoning bottleneck rather than task-specific gaps.
  • The framework predicts that training on HLS data can strengthen general reasoning; the paper cites evidence that metaphorical images and humorous content improve complex visual and logical reasoning.
  • Because the same mechanisms can both reveal implicit hate and enable metaphor- or persuasion-based jailbreaks, HLS is a double-edged factor in AI safety.
  • If the BLSI/HLSI distinction is correct, then measuring model capability should include not just correctness but whether the model can explain the mechanism—incongruity, source-target mapping, affective cause—that produces the high-level meaning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the proposed metrics were ever operationalized, semantic density could become a quantitative ranking of utterances and experiences, connecting HLS to information theory and cognitive-load measures; the paper leaves this connection implicit.
  • A testable extension would be a developmental-style benchmark: paired literal and nonliteral stimuli matched for surface length, presented across modalities, to see whether models acquire BLS skills first and HLS skills later in roughly the order humans do—literal, then metaphor/sarcasm/humor, then persuasion and narrative.
  • Because the paper explicitly excludes formal-rule domains like mathematical reasoning and code generation, HLSI should be read as a social-communicative axis of intelligence, not a complete account of high-level cognition.
  • The taxonomy's real contribution may be to force researchers to define what 'literal' and 'implicit' mean in concrete task terms, since the BLS/HLS boundary is likely to shift as models improve.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper argues that AI is undergoing a transition from Basic-Level Semantic Intelligence (BLSI) to High-Level Semantic Intelligence (HLSI), defined by a shift from literal, perceptually grounded semantics to context-dependent, socially situated, and cognitively complex meanings. It proposes a conceptual framework built on Semantic Clues, Semantic Chains, Causal Chains, and Semantic Effects, and introduces two quantitative notions, Semantic Complexity (Cs) and Semantic Density (Ds), meant to characterize the BLS/HLS boundary. The survey then organizes recent work on humor, sarcasm, metaphor, empathy, persuasion, and narrative across text, speech, vision, and multimodal settings, covering data construction, modeling/optimization, and evaluation for both understanding and generation. It also discusses applications in AI safety, mental health, social agents, negotiation, creativity, and cultural understanding, and closes with open challenges.

Significance. If the HLS framing is accepted, the paper provides a useful common vocabulary and a broad map of otherwise scattered tasks, spanning modalities and both understanding and generation. Its strengths are the extensive reference coverage, the systematic data/modeling/evaluation taxonomy in Sections IV and V, and its explicit treatment of subjectivity, quality control, and evaluation limitations. The paper also candidly acknowledges in Section VII.B that measuring abstract semantic complexity is particularly challenging and that systematic research is limited. The main risk is that the formal apparatus—especially the complexity metric—is not operational, so the paper's headline theoretical contribution currently rests on stipulation rather than measurement. The survey's organizational value, however, is largely independent of the formal metric, and the claimed trajectory could be supported with more transparent empirical evidence.

major comments (3)
  1. [§II-A2, Eqs. (4)–(5)] Semantic Complexity is defined only as Cs ∝ Complexity(SC, CC), but Complexity(SC, CC) is never defined: there is no function, unit, measurement protocol, or comparison procedure. SC is a semantic inference relation and CC a causal relation, but no account is given of how their 'complexity' is computed. Consequently Ds = Cs/|R| is also non-operational. The BLSI/HLSI boundary in §II-A3 and the stage transition in Fig. 2 are therefore asserted rather than derived from the formal apparatus. The paper itself concedes in §VII.B that measurement is 'particularly challenging' and 'systematic research remains limited.' This does not invalidate the survey, but the formal equations should either be operationalized or repositioned as informal illustrative notation, with the claimed contribution adjusted accordingly.
  2. [Figure 1, footnote 2] The 'clear trajectory' and 'pronounced growth' claims rest on the publication-density statistics, but the methodology is not disclosed. The caption says counts were compiled via the DBLP Search API across 15 top-tier venues, but omits the exact queries, keyword lists, search fields (title/abstract/full text), inclusion and exclusion criteria, normalization procedure, and year groupings. Without this information, the figure cannot be reproduced or checked for query artifacts, duplicate counts, or venue bias. A detailed methodology appendix or released query/count file is needed for this evidence to be load-bearing.
  3. [§II-A3, §III-B] The selection of ChatGPT as the 'rough boundary' between BLSI and HLSI is not connected to the proposed formal quantities. No computation shows that tasks before ChatGPT are below some threshold of Cs or Ds and tasks after are above it; the boundary is chronological. Similarly, the six core HLS phenomena are designated as such by stipulation, and their coherence is supported by shared informal features (§II-B) rather than by the formal metric. These choices may be reasonable as a research proposal, but the paper should present them as an organizing hypothesis, not as a consequence of the semantic-complexity framework.
minor comments (6)
  1. [Footnote 1] 'Working in progress' should be 'Work in progress.'
  2. [§II-A2, Eq. (5)] The notation |R| is explained only by parenthetical examples (token count, pixel area); a formal definition in the definition block would improve precision.
  3. [Figure 2] The row of milestone labels (e.g., 'AlexNetSeq2Seq', 'Deep Speech') is visually compressed and hard to read; adding spacing or separate lines would improve legibility.
  4. [§II-A1d] The abbreviation 'SE' is introduced for Semantic Effect but is rarely used later; consider using it consistently or dropping it to avoid unnecessary notation.
  5. [Figures 5 and 6] The taxonomy diagrams are dense. Several leaf terms (e.g., 'Theory-Guided Semantic Structuring' vs. 'Theory-Guided Semantic Annotation') are close; check that each term is exactly aligned with the section it summarizes.
  6. [References] In-text citations such as [239], [240], [245], and [246] appear in the body; please verify that all cited entries are present in the final reference list. The provided text is truncated before the end, so this is a consistency check rather than a claim of a missing entry.

Circularity Check

1 steps flagged · score 2.0 of 10

Mild definitional circularity: HLS is defined through the same cognitive-complexity constructs that Eq. (4) uses to measure complexity, making the BLS/HLS boundary stipulative.

  1. self definitional [Section II.A.3(b) (HLS definition); Section II.A.2(a), Eq. (4); Section II.B(c)]
    "High-Level Semantics (HLS) refers to abstract meaning constructed through the integration of basic-level observations and complex cognitive reasoning. ... High semantic complexity implies that multiple layers of inference, causal reasoning, and contextual integration are required. ... HLS typically involves a higher degree of semantic complexity, as its understanding and generation often requires multi-hop reasoning."

    HLS is defined as meaning produced by complex cognitive reasoning over Semantic Chains and Causal Chains, while Eq. (4) defines semantic complexity as Complexity(SC,CC). The claim that HLS has high semantic complexity is therefore true by construction, since the quantity supposed to measure the HLS/BLS distinction is defined in terms of the very chains that constitute HLS. No independent measurement of Cs is given to separate HLS from BLS. The paper concedes in Sec. VII.B that measuring abstract semantic complexity is 'particularly challenging' and systematic research is 'limited,' so the formal boundary is stipulative rather than empirically derived. This does not invalidate the survey's taxonomy, but the central stage-claim rests on definition, not on an operationalized metric.

full rationale

The paper is a survey rather than a prediction-driven study. I found no fitted-parameter circularity, no self-citation chain, and no equation used to generate a numerical prediction from fitted inputs. The only reduction-by-construction appears in the formal framework: HLS is defined as abstract meaning requiring complex cognitive reasoning over Semantic Chains and Causal Chains, and Eq. (4) then states Cs ∝ Complexity(SC,CC). Consequently, asserting that HLS has high semantic complexity is a restatement of the definition rather than an independent empirical finding. The BLS/HLS boundary is therefore stipulative, and the claimed 'clear trajectory' in Fig. 1 is not derived from an operationalized complexity measure; the paper itself acknowledges in Sec. VII.B that measuring abstract semantic complexity is difficult and underdeveloped. This is a mild definitional circularity, not a forced derivation. The survey's extensive literature organization and its characterization of HLS tasks remain independently useful.

Assumptions & free parameters 0 free parameters · 5 assumptions · 5 invented entities

The framework is entirely conceptual. There are no fitted numeric parameters, but the entire apparatus rests on stipulated constructs (Semantic Clue/Chain/Causal Chain/Effect) and an undefined Complexity function. The human-AI parallelism and the six-task selection are domain assumptions made for the survey's scope. None of the invented entities has an independent falsifiable handle, so the framework cannot be empirically confirmed or refuted as stated.

assumptions (5)
  • ad hoc to paper HLS phenomena can be decomposed into Semantic Clues, Semantic Chains, Causal Chains, and Semantic Effects.
    Introduced as the foundational decomposition in Sec. II-A1; no independent evidence that this decomposition is exhaustive or unique.
  • ad hoc to paper Semantic Complexity and Semantic Density are measurable quantities via Cs ∝ Complexity(SC,CC) and Ds = Cs/|R|.
    Sec. II-A2 presents these as formal metrics, but Complexity(SC,CC) is never defined, so the quantities are not computable or falsifiable.
  • domain assumption AI semantic development parallels human cognitive development from literal to abstract, socially embedded semantics.
    Used in Sec. I and Figure 2 to motivate the BLSI-HLSI trajectory, supported only by selected psycholinguistics references, not by a systematic comparative argument.
  • ad hoc to paper The Transformer and ChatGPT mark the boundary between BLSI and HLSI.
    Stated in Sec. I as a rough boundary; this is a modeling choice, not an empirically established threshold.
  • ad hoc to paper Humor, sarcasm, metaphor, empathy, persuasion, and narrative are the core HLS phenomena.
    The survey explicitly excludes math/code and other formal-rule domains (footnote 3) and selects six categories; this scope decision is not derived from a prior theory.
invented entities (5)
  • Semantic Clue (c)
    purpose: Atomic unit of semantic evidence that a system must perceive and reason over.
    No operational definition or measurement procedure is provided; it is a conceptual primitive.
  • Semantic Chain (SC)
    purpose: Represents the inference relation ⊢ from clues to a derived HLS conclusion.
    The ⊢ symbol is used informally in Eq. (1) without a formal semantics or inference rules.
  • Causal Chain (CC)
    purpose: Represents causal factors ⇝ effects within semantic interpretation.
    The ⇝ relation in Eq. (2) is not formally defined or grounded in a causal inference framework.
  • Semantic Effect (E)
    purpose: Composite cognitive/affective/communicative outcome E = Φ({SC},{CC}).
    Φ is unspecified in Eq. (3); no method is given to compute or evaluate E.
  • BLSI / HLSI stage labels
    purpose: Provide a high-level classification of AI semantic capability over time.
    The boundary examples (Transformer, ChatGPT) are illustrative; no quantitative measure is attached to the stages.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards High-Level Semantic Intelligence." pith.science (2026). https://pith.science/paper/WBJXNTF3

@misc{pith2026260724082,
  author       = {Pith},
  title        = {Pith review of: Towards High-Level Semantic Intelligence},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WBJXNTF3}},
  note         = {Machine review of arXiv:2607.24082}
}
read the original abstract

Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, the development of AI reveals a clear trajectory from simple to complex semantic processing. While early AI systems mainly addressed tasks involving direct and literal semantic perception or expression, contemporary systems are increasingly expected to perform more sophisticated cognitive reasoning, enabling the understanding and generation of High-Level Semantics (HLS). A similar trajectory can also be observed in human cognitive development. We define this transition as the shift from Basic-Level Semantic Intelligence (BLSI) to High-Level Semantic Intelligence (HLSI). However, this issue has not yet been systematically and comprehensively examined in prior work. Motivated by this gap, this survey reviews the development of AI semantic intelligence from the perspective of semantic complexity. We systematically survey existing research on HLS tasks, including humor, sarcasm, metaphor, empathy, persuasion, narrative, and other general HLS phenomena, across text, speech, vision, and multimodal scenarios. Specifically, we summarize data construction methods, modeling and optimization strategies, and evaluation methodologies for both understanding and generation. HLS is essential for advancing AI toward genuinely human-like intelligence. By synthesizing existing methods and insights from the perspective of semantic intelligence, this survey aims to support the continued development of AI toward HLSI.

Figures

Figures reproduced from arXiv: 2607.24082 by the authors.

Figure 1
Figure 1. Evolutionary trajectory and publication density of High-Level Se [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Conceptual taxonomy and evolutionary roadmap from Basic-Level Semantic Intelligence (BLSI) to High-Level Semantic Intelligence (HLSI). Top: [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Illustration of Semantic Level-Up through cognitive reasoning across different modalities. By combining lower-level perceptual or contextual semantic [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: The cognitive architecture of High-Level Semantic Understanding [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Taxonomy of High-Level Semantic (HLS) understanding methods, covering data construction, modeling approaches, and evaluation protocols. [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: Taxonomy of High-Level Semantic (HLS) generation methods, covering data construction, generation approaches, and evaluation protocols. [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

300 extracted references · 3 canonical work pages

  1. [1]

    Gradient-based learning applied to document recognition,

    Y . Lecun, L. Bottou, Y . Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,”Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998

  2. [2]

    Maximum mutual information estimation of hidden markov model parameters for speech recognition,

    L. Bahl, P. Brown, P. de Souza, and R. Mercer, “Maximum mutual information estimation of hidden markov model parameters for speech recognition,” inICASSP ’86. IEEE International Conference on Acous- tics, Speech, and Signal Processing, vol. 11, 1986, pp. 49–52

  3. [3]

    A neural proba- bilistic language model,

    Y . Bengio, R. Ducharme, P. Vincent, and C. Jauvin, “A neural proba- bilistic language model,”Journal of machine learning research, vol. 3, no. Feb, pp. 1137–1155, 2003

  4. [4]

    A mathematical theory of communication,

    C. E. Shannon, “A mathematical theory of communication,”The Bell System Technical Journal, vol. 27, no. 3, pp. 379–423, 1948

  5. [5]

    Recent contributions to the mathematical theory of com- munication,

    W. Weaver, “Recent contributions to the mathematical theory of com- munication,”ETC: a review of general semantics, vol. 74, no. 1/2, pp. 136–157, 2017

  6. [6]

    Tomasello, m., constructing a language: a usage-based theory of language acquisition. cambridge, ma: Harvard university press, 2003. pp. 388. hardback, £29.95. isbn 0-674-01030-2

    J. M. PINE, “Tomasello, m., constructing a language: a usage-based theory of language acquisition. cambridge, ma: Harvard university press, 2003. pp. 388. hardback, £29.95. isbn 0-674-01030-2.”Journal of Child Language, vol. 32, no. 3, p. 697–702, 2005

  7. [7]

    Bloom,How children learn the meanings of words

    P. Bloom,How children learn the meanings of words. MIT press, 2002

  8. [8]

    Winner,The point of words: Children’s understanding of metaphor and irony

    E. Winner,The point of words: Children’s understanding of metaphor and irony. Harvard University Press, 1988

Show all 300 references
  1. [9]

    Playing with expectations: A contextual view of humor development,

    G. Airenti, “Playing with expectations: A contextual view of humor development,”Frontiers in Psychology, vol. 7, p. 1392, 2016

  2. [10]

    Empathy and moral development,

    M. L. Hoffman, “Empathy and moral development,”The annual report of educational psychology in Japan, vol. 35, pp. 157–162, 1996

  3. [11]

    J. S. Bruner,Acts of meaning: Four lectures on mind and culture. Harvard university press, 1990, vol. 3

  4. [12]

    A theory of argumentative understand- ing: Relationships among position preference, judgments of goodness, memory and reasoning,

    N. L. Stein and C. A. Miller, “A theory of argumentative understand- ing: Relationships among position preference, judgments of goodness, memory and reasoning,”Argumentation, vol. 7, no. 2, pp. 183–204, 1993

  5. [13]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems, vol. 30, 2017

  6. [14]

    Gpt-4 technical report,

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkatet al., “Gpt-4 technical report,”arXiv preprint arXiv:2303.08774, 2023. 29

  7. [15]

    B. G. Buchanan and E. H. Shortliffe,Rule Based Expert Systems: The Mycin Experiments of the Stanford Heuristic Programming Project (The Addison-Wesley series in artificial intelligence). USA: Addison- Wesley Longman Publishing Co., Inc., 1984

  8. [16]

    Imagenet classification with deep convolutional neural networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” inAdvances in Neural Information Processing Systems, F. Pereira, C. Burges, L. Bottou, and K. Weinberger, Eds., vol. 25. Curran Associates, Inc., 2012. [Online]. A...

  9. [17]

    Deep speech: Scaling up end-to-end speech recognition,

    A. Hannun, C. Case, J. Casper, B. Catanzaro, G. Diamos, E. Elsen, R. Prenger, S. Satheesh, S. Sengupta, A. Coateset al., “Deep speech: Scaling up end-to-end speech recognition,”arXiv preprint arXiv:1412.5567, 2014

  10. [18]

    Rich feature hierarchies for accurate object detection and semantic segmentation,

    R. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich feature hierarchies for accurate object detection and semantic segmentation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2014

  11. [19]

    Bert: Pre- training of deep bidirectional transformers for language understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. N. Toutanova, “Bert: Pre- training of deep bidirectional transformers for language understanding,”

  12. [20]

    Improving language understanding by generative pre-training,

    A. Radford, K. Narasimhan, T. Salimans, I. Sutskeveret al., “Improving language understanding by generative pre-training,” 2018

  13. [21]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” inProceedings of the 38th International Conference on Machine Le...

  14. [22]

    Videobert: A joint model for video and language representation learning,

    C. Sun, A. Myers, C. V ondrick, K. Murphy, and C. Schmid, “Videobert: A joint model for video and language representation learning,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), October 2019

  15. [23]

    Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,

    H. Akbari, L. Yuan, R. Qian, W.-H. Chuang, S.-F. Chang, Y . Cui, and B. Gong, “Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text,” in Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y . Dauphin, P. Liang, a...

  16. [24]

    Flamingo: a visual language model for few-shot learning,

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y . Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynoldset al., “Flamingo: a visual language model for few-shot learning,”Advances in neural information processing systems, vol. 35, pp. 23 716–23 736, 2022

  17. [25]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” inAdvances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., vol. 36. Curran Associates, Inc., 2023, pp. 34 892–34 916. [Online]. Available: ht...

  18. [26]

    Paper review:’sparks of artificial general intelligence: Early experiments with gpt-4’,

    S. Bubeck, V . Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y . T. Lee, Y . Li, S. Lundberget al., “Paper review:’sparks of artificial general intelligence: Early experiments with gpt-4’,” 2023

  19. [27]

    Open-sora: Democratizing efficient video production for all,

    Z. Zheng, X. Peng, T. Yang, C. Shen, S. Li, H. Liu, Y . Zhou, T. Li, and Y . You, “Open-sora: Democratizing efficient video production for all,”CoRR, vol. abs/2412.20404, 2024. [Online]. Available: https://doi.org/10.48550/arXiv.2412.20404

  20. [28]

    Seedance 2.0: Advancing video generation for world complexity,

    T. Seedance, D. Chen, L. Chen, X. Chen, Y . Chen, Z. Chen, Z. Chen, F. Cheng, T. Cheng, Y . Chenget al., “Seedance 2.0: Advancing video generation for world complexity,”arXiv preprint arXiv:2604.14148, 2026

  21. [29]

    Large lan- guage models for subjective language understanding: A survey,

    C. Song, Y . Zhang, H. Gao, B. Yao, and P. Zhang, “Large lan- guage models for subjective language understanding: A survey,”arXiv preprint arXiv:2508.07959, 2025

  22. [30]

    Looking beyond the obvious: A survey on abstract concept recognition for video understanding: Looking beyond the obvious: A survey on abstract concept

    G. Mago, P. Mettes, and S. Rudinac, “Looking beyond the obvious: A survey on abstract concept recognition for video understanding: Looking beyond the obvious: A survey on abstract concept...”Int. J. Comput. Vision, vol. 134, no. 5, Apr. 2026. [Online]. Available: https://doi.o...

  23. [31]

    Class-basedn-gram models of natural language,

    P. F. Brown, V . J. Della Pietra, P. V . deSouza, J. C. Lai, and R. L. Mercer, “Class-basedn-gram models of natural language,” Computational Linguistics, vol. 18, no. 4, pp. 467–480, 1992. [Online]. Available: https://aclanthology.org/J92-4003/

  24. [32]

    Efficient estimation of word representations in vector space,

    T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,”arXiv preprint arXiv:1301.3781, 2013

  25. [33]

    Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition,

    E. F. Tjong Kim Sang and F. De Meulder, “Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition,” inProceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, 2003, pp. 142–147. [Online]. Available: https://aclantho...

  26. [34]

    Message Understanding Conference- 6: A brief history,

    R. Grishman and B. Sundheim, “Message Understanding Conference- 6: A brief history,” inCOLING 1996 Volume 1: The 16th International Conference on Computational Linguistics, 1996. [Online]. Available: https://aclanthology.org/C96-1079/

  27. [35]

    Statistical phrase-based transla- tion,

    P. Koehn, F. J. Och, and D. Marcu, “Statistical phrase-based transla- tion,” inProceedings of the 2003 human language technology confer- ence of the north american chapter of the association for computational linguistics, 2003, pp. 127–133

  28. [36]

    Neural machine translation by jointly learning to align and translate,

    D. Bahdanau, K. Cho, and Y . Bengio, “Neural machine translation by jointly learning to align and translate,” in3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, Y . Bengio and Y . LeCun, Eds.,...

  29. [37]

    SQuAD: 100,000+ questions for machine comprehension of text,

    P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang, “SQuAD: 100,000+ questions for machine comprehension of text,” inProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, J. Su, K. Duh, and X. Carreras, Eds. Austin, Texas: Association for Comput...

  30. [38]

    A neural attention model for abstractive sentence summarization,

    A. M. Rush, S. Chopra, and J. Weston, “A neural attention model for abstractive sentence summarization,” inProceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, L. M `arquez, C. Callison-Burch, and J. Su, Eds. Lisbon, Portugal: Association for...

  31. [39]

    Automatically constructing a corpus of sentential paraphrases,

    W. B. Dolan and C. Brockett, “Automatically constructing a corpus of sentential paraphrases,” inProceedings of the Third International Workshop on Paraphrasing (IWP2005), 2005. [Online]. Available: https://aclanthology.org/I05-5002/

  32. [40]

    Imagenet: A large-scale hierarchical image database,

    J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255

  33. [41]

    Faster r-cnn: Towards real-time object detection with region proposal networks,

    S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” inAdvances in Neural Information Processing Systems, C. Cortes, N. Lawrence, D. Lee, M. Sugiyama, and R. Garnett, Eds., vol. 28. Curran Associates, Inc., 2...

  34. [42]

    Fully convolutional networks for semantic segmentation,

    J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” inProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015

  35. [43]

    Mask r-cnn,

    K. He, G. Gkioxari, P. Doll ´ar, and R. Girshick, “Mask r-cnn,” in Proceedings of the IEEE international conference on computer vision, 2017, pp. 2961–2969

  36. [44]

    Realtime multi-person 2d pose estimation using part affinity fields,

    Z. Cao, T. Simon, S.-E. Wei, and Y . Sheikh, “Realtime multi-person 2d pose estimation using part affinity fields,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 7291–7299

  37. [45]

    State of the art in example-based texture synthesis,

    L.-Y . Wei, S. Lefebvre, V . Kwatra, and G. Turk, “State of the art in example-based texture synthesis,”Eurographics 2009, State of the Art Report, EG-STAR, pp. 93–117, 2009

  38. [46]

    Texture synthesis using convo- lutional neural networks,

    L. Gatys, A. S. Ecker, and M. Bethge, “Texture synthesis using convo- lutional neural networks,”Advances in neural information processing systems, vol. 28, 2015

  39. [47]

    Globally and locally consistent image completion,

    S. Iizuka, E. Simo-Serra, and H. Ishikawa, “Globally and locally consistent image completion,”ACM Transactions on Graphics (ToG), vol. 36, no. 4, pp. 1–14, 2017

  40. [48]

    Image completion with structure propagation,

    J. Sun, L. Yuan, J. Jia, and H.-Y . Shum, “Image completion with structure propagation,” inACM Siggraph 2005 Papers, 2005, pp. 861– 868

  41. [49]

    Auto-encoding variational bayes,

    D. P. Kingma and M. Welling, “Auto-encoding variational bayes,”arXiv preprint arXiv:1312.6114, 2013

  42. [50]

    Generative adversarial nets,

    I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, “Generative adversarial nets,” Advances in neural information processing systems, vol. 27, 2014

  43. [51]

    A tutorial on hidden markov models and selected applica- tions in speech recognition,

    L. Rabiner, “A tutorial on hidden markov models and selected applica- tions in speech recognition,”Proceedings of the IEEE, vol. 77, no. 2, pp. 257–286, 1989. 30

  44. [52]

    Suppression of acoustic noise in speech using spectral subtraction,

    S. Boll, “Suppression of acoustic noise in speech using spectral subtraction,”IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 27, no. 2, pp. 113–120, 1979

  45. [53]

    Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,

    W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016, pp. 4960–4964

  46. [54]

    X-vectors: Robust dnn embeddings for speaker recognition,

    D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudanpur, “X-vectors: Robust dnn embeddings for speaker recognition,” in2018 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP), 2018, pp. 5329–5333

  47. [55]

    Emotional speech synthesis: a review

    M. Schr ¨oder, “Emotional speech synthesis: a review.” inInterspeech, vol. 2001, 2001, pp. 561–564

  48. [56]

    Speech parameter generation algorithms for hmm-based speech syn- thesis,

    K. Tokuda, T. Yoshimura, T. Masuko, T. Kobayashi, and T. Kitamura, “Speech parameter generation algorithms for hmm-based speech syn- thesis,” in2000 IEEE international conference on acoustics, speech, and signal processing. Proceedings (Cat. No. 00CH37100), vol. 3. IEEE, 2000,...

  49. [57]

    Wavenet: A generative model for raw audio,

    A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, K. Kavukcuogluet al., “Wavenet: A generative model for raw audio,”arXiv preprint arXiv:1609.03499, vol. 12, no. 1, 2016

  50. [58]

    Tacotron: Towards end- to-end speech synthesis,

    Y . Wang, R. Skerry-Ryan, D. Stanton, Y . Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y . Xiao, Z. Chen, S. Bengioet al., “Tacotron: Towards end- to-end speech synthesis,”arXiv preprint arXiv:1703.10135, 2017

  51. [59]

    Imagebind: One embedding space to bind them all,

    R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V . Alwala, A. Joulin, and I. Misra, “Imagebind: One embedding space to bind them all,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 15 180–15 190

  52. [60]

    Making computers laugh: Inves- tigations in automatic humor recognition,

    R. Mihalcea and C. Strapparava, “Making computers laugh: Inves- tigations in automatic humor recognition,” inProceedings of human language technology conference and conference on empirical methods in natural language processing, 2005, pp. 531–538

  53. [61]

    SemEval-2017 task 6: #HashtagWars: Learning a sense of humor,

    P. Potash, A. Romanov, and A. Rumshisky, “SemEval-2017 task 6: #HashtagWars: Learning a sense of humor,” inProceedings of the 11th International Workshop on Semantic Evaluation (SemEval- 2017), S. Bethard, M. Carpuat, M. Apidianaki, S. M. Mohammad, D. Cer, and D. Jurgens, Eds....

  54. [62]

    “president vows to cut <taxes>hair

    N. Hossain, J. Krumm, and M. Gamon, ““president vows to cut <taxes>hair”: Dataset and analysis of creative text editing for humorous headlines,” inProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language ...

  55. [63]

    Cards against AI: Predicting humor in a fill-in-the-blank party game,

    D. Ofer and D. Shahaf, “Cards against AI: Predicting humor in a fill-in-the-blank party game,” inFindings of the Association for Computational Linguistics: EMNLP 2022, Y . Goldberg, Z. Kozareva, and Y . Zhang, Eds. Abu Dhabi, United Arab Emirates: Association for Computational...

  56. [64]

    Can language models make fun? a case study in Chinese comical crosstalk,

    J. Li, X. Wu, X. Liu, Q. Xie, P. Tiwari, and B. Wang, “Can language models make fun? a case study in Chinese comical crosstalk,” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J. Boyd-Graber, and N...

  57. [65]

    Talk funny! a large-scale humor response dataset with chain-of-humor interpretation,

    Y . Chen, Y . Yuan, P. Liu, D. Liu, Q. Guan, M. Guo, H. Peng, B. Liu, Z. Li, and Y . Xiao, “Talk funny! a large-scale humor response dataset with chain-of-humor interpretation,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 16, p. 17826–17834, Mar....

  58. [66]

    Chumor 2.0: Towards better benchmarking Chinese humor understanding from (ruo zhi ba),

    R. He, Y . He, L. Bai, J. Liu, Z. Sun, Z. Tang, H. Wang, H. Xia, R. Mihalcea, and N. Deng, “Chumor 2.0: Towards better benchmarking Chinese humor understanding from (ruo zhi ba),” inFindings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shu...

  59. [67]

    “what do you call a dog that is incontrovertibly true? dogma

    A. Cocchieri, L. Ragazzi, P. Italiani, G. Tagliavini, and G. Moro, ““what do you call a dog that is incontrovertibly true? dogma”: Testing LLM generalization through humor,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Lo...

  60. [68]

    Cfunmodel: A

    Z. Yu, X. Hu, and X. Wan, “Cfunmodel: A” funny” language model capable of chinese humor generation and processing,”arXiv preprint arXiv:2503.20417, 2025

  61. [69]

    Comparing apples to oranges: A dataset & analysis of LLM humour understanding from traditional puns to topical jokes,

    T. Loakman, W. Thorne, and C. Lin, “Comparing apples to oranges: A dataset & analysis of LLM humour understanding from traditional puns to topical jokes,” inFindings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, a...

  62. [70]

    Drivel-ology: Challenging llms with interpreting nonsense with depth,

    Y . Wang, C. Xiao, C.-Y . Hsiao, Z. Y . Chang, C.-L. Chen, T. Loakman, and C. Lin, “Drivel-ology: Challenging llms with interpreting nonsense with depth,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 23 085–23 107

  63. [71]

    Do androids laugh at electric sheep? humor “understanding

    J. Hessel, A. Marasovic, J. D. Hwang, L. Lee, J. Da, R. Zellers, R. Mankoff, and Y . Choi, “Do androids laugh at electric sheep? humor “understanding” benchmarks from the new yorker caption contest,” inProceedings of the 61st Annual Meeting of the Association for Computational...

  64. [72]

    Humor in ai: Massive scale crowd-sourced preferences and benchmarks for cartoon caption- ing,

    J. Zhang, L. Jain, Y . Guo, J. Chen, K. L. Zhou, S. Suresh, A. Wagen- maker, S. Sievert, T. Rogers, K. Jamiesonet al., “Humor in ai: Massive scale crowd-sourced preferences and benchmarks for cartoon caption- ing,”Advances in Neural Information Processing Systems, vol. 37, pp....

  65. [73]

    Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation,

    S. Zhong, Z. Huang, S. Gao, W. Wen, L. Lin, M. Zitnik, and P. Zhou, “Let’s think outside the box: Exploring leap-of-thought in large language models with creative humor generation,” inProceed- ings - 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 202...

  66. [74]

    Cracking the code of juxtaposition: can ai models understand the humorous contradictions,

    Z. Hu, T. Liang, J. Li, Y . Lu, Y . Zhou, Y . Qiao, J. Ma, and Y . Yin, “Cracking the code of juxtaposition: can ai models understand the humorous contradictions,” inProceedings of the 38th International Conference on Neural Information Processing Systems, ser. NIPS ’24. Red H...

  67. [75]

    Visionarena: 230k real world user-vlm conversations with preference labels,

    C. Chou, L. Dunlap, K. Mashita, K. Mandal, T. Darrell, I. Stoica, J. E. Gonzalez, and W.-L. Chiang, “Visionarena: 230k real world user-vlm conversations with preference labels,” in2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, pp. 3877– 3887

  68. [76]

    D-humor: Dark humor understanding via multimodal open-ended reasoning,

    S. K. R. Kasu, M. Z. Ur Rehman, S. S. Dar, R. B. Junghare, D. S. Namboodiri, and N. Kumar, “D-humor: Dark humor understanding via multimodal open-ended reasoning,” in2025 IEEE International Conference on Data Mining (ICDM), 2025, pp. 377–386

  69. [77]

    Humor in pixels: Benchmarking large multimodal models understanding of online comics,

    Y . Ryan, R. Y . Tan, K. T. W. Choo, and R. K.-W. Lee, “Humor in pixels: Benchmarking large multimodal models understanding of online comics,” inFindings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V . Peng,...

  70. [78]

    HumorDB: Can AI understand graphical humor?

    V . V . Jain, F. d. S. A. Feitosa, and G. Kreiman, “HumorDB: Can AI understand graphical humor?” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025. [Online]. Available: https: //openaccess.thecvf.com/content/ICCV2025/papers/Jain HumorDB Can AI un...

  71. [79]

    V-hub: A visual-centric humor understanding benchmark for video llms,

    Z. Shi, H. Li, Y . Zhao, J. Zhou, Y . Wang, Q. Cui, W. Bi, S. Zhu, B. Zhao, and Z. Zheng, “V-hub: A visual-centric humor understanding benchmark for video llms,”arXiv preprint arXiv:2509.25773, 2025

  72. [80]

    GODBench: A benchmark for multimodal large language models in video comment art,

    Y . Lei, C. Zhang, Z. Liu, H. Leng, S. Liu, T. Gao, Q. Liu, and Y . Wang, “GODBench: A benchmark for multimodal large language models in video comment art,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics 31 (Volume 1: Long Papers), W....

  73. [81]

    Can language models laugh at YouTube short-form videos?

    D. Ko, S. Lee, and G. Kim, “Can language models laugh at YouTube short-form videos?” inProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Singapore: Association for Computational Linguistics, 2023, pp. 2897–2916. [Online]. Available: https:...

  74. [82]

    Towards multimodal prediction of spontaneous humor: A novel dataset and first results,

    L. Christ, S. Amiriparian, A. Kathan, N. M ¨uller, A. K ¨onig, and B. W. Schuller, “Towards multimodal prediction of spontaneous humor: A novel dataset and first results,”IEEE Transactions on Affective Computing, vol. 16, no. 2, pp. 844–860, 2025

  75. [83]

    StandUp4AI: A new multilingual dataset for humor detection in stand- up comedy videos,

    V . Barriere, N. Gomez, L. Hemamou, S. Callejas, and B. Ravenet, “StandUp4AI: A new multilingual dataset for humor detection in stand- up comedy videos,” inFindings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, an...

  76. [84]

    Testing the ability of language models to interpret figurative language,

    E. Liu, C. Cui, K. Zheng, and G. Neubig, “Testing the ability of language models to interpret figurative language,” inProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, M. Carpuat, M.-...

  77. [85]

    FLUTE: Figurative language understanding through textual explanations,

    T. Chakrabarty, A. Saakyan, D. Ghosh, and S. Muresan, “FLUTE: Figurative language understanding through textual explanations,” inProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y . Goldberg, Z. Kozareva, and Y . Zhang, Eds. Abu Dhabi, Un...

  78. [86]

    Chinese metaphorical relation extraction,

    G. Chen, T. Wu, M. Cheng, X. Han, J. Gong, S. Wang, and W. Song, “Chinese metaphorical relation extraction,” inFindings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational Linguistics, De...

  79. [87]

    Chinese idiom paraphrasing,

    J. Qiang, Y . Li, C. Zhang, Y . Li, Y . Zhu, Y . Yuan, and X. Wu, “Chinese idiom paraphrasing,”Transactions of the Association for Computational Linguistics, vol. 11, pp. 740–754, 2023

  80. [88]

    Multi-lingual and multi- cultural figurative language understanding,

    A. Kabra, E. Liu, S. Khanuja, A. F. Aji, G. Winata, S. Cahyawijaya, A. Aremu, P. Ogayo, and G. Neubig, “Multi-lingual and multi- cultural figurative language understanding,” inFindings of the Association for Computational Linguistics: ACL 2023, A. Rogers, J. Boyd-Graber, and N...

  81. [89]

    NewsMet : A ‘do it all’ dataset of contemporary metaphors in news headlines,

    R. Joseph, T. Liu, A. B. Ng, S. See, and S. Rai, “NewsMet : A ‘do it all’ dataset of contemporary metaphors in news headlines,” inFindings of the Association for Computational Linguistics: ACL 2023, A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds. Toronto, Canada: Association f...

  82. [90]

    MMTE: Corpus and metrics for evaluating machine translation quality of metaphorical language,

    S. Wang, G. Zhang, H. Wu, T. Loakman, W. Huang, and C. Lin, “MMTE: Corpus and metrics for evaluating machine translation quality of metaphorical language,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and ...

  83. [91]

    MetaPro 2.0: Computational metaphor processing on the effectiveness of anomalous language modeling,

    R. Mao, K. He, C. Ong, Q. Liu, and E. Cambria, “MetaPro 2.0: Computational metaphor processing on the effectiveness of anomalous language modeling,” inFindings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins, and V . Srikumar, Eds. Bangkok, Tha...

  84. [92]

    Metaphor understanding challenge dataset for LLMs,

    X. Tong, R. Choenni, M. Lewis, and E. Shutova, “Metaphor understanding challenge dataset for LLMs,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L.-W. Ku, A. Martins, and V . Srikumar, Eds. Bangkok, Thailand...

  85. [93]

    Enhancing information extraction with metorie: A metaphor and trap-based dataset for cross-domain fine-tuning,

    Z. Pan, Y . Peng, Z. Jian, Y . Chen, W. Qiu, H. Ma, J. Yao, M. Wang, and Q. Wu, “Enhancing information extraction with metorie: A metaphor and trap-based dataset for cross-domain fine-tuning,” inICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal P...

  86. [94]

    Investigating the impact of conceptual metaphors on llm-based nli through shapley interactions,

    M. Sengupta, M. Muschalik, F. Fumagalli, B. Hammer, E. H ¨ullermeier, D. Ghosh, and H. Wachsmuth, “Investigating the impact of conceptual metaphors on llm-based nli through shapley interactions,” inFindings of the Association for Computational Linguistics: EMNLP 2025, 2025, pp...

  87. [95]

    Automatic extraction of metaphoric analogies from literary texts: Task formulation, dataset construction, and evaluation,

    J. Boisson, Z. Siddique, H. Borkakoty, D. Antypas, L. Espinosa Anke, and J. Camacho-Collados, “Automatic extraction of metaphoric analogies from literary texts: Task formulation, dataset construction, and evaluation,” inProceedings of the 31st International Conference on Compu...

  88. [96]

    Comparative study of multilingual idioms and similes in large language models,

    P. Khoshtab, D. Namazifard, M. Masoudi, A. Akhgary, S. Mahdizadeh Sani, and Y . Yaghoobzadeh, “Comparative study of multilingual idioms and similes in large language models,” in Proceedings of the 31st International Conference on Computational Linguistics, O. Rambow, L. Wanner...

  89. [97]

    I spy a metaphor: Large language models and diffusion models co-create visual metaphors,

    T. Chakrabarty, A. Saakyan, O. Winn, A. Panagopoulou, Y . Yang, M. Apidianaki, and S. Muresan, “I spy a metaphor: Large language models and diffusion models co-create visual metaphors,” inFindings of the Association for Computational Linguistics: ACL 2023, A. Rogers, J. Boyd-G...

  90. [98]

    MultiCMET: A novel Chinese benchmark for understanding multimodal metaphor,

    D. Zhang, J. Yu, S. Jin, L. Yang, and H. Lin, “MultiCMET: A novel Chinese benchmark for understanding multimodal metaphor,” in Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational...

  91. [99]

    MemeCap: A dataset for captioning and interpreting memes,

    E. Hwang and V . Shwartz, “MemeCap: A dataset for captioning and interpreting memes,” inProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational Linguistics, Dec. 202...

  92. [100]

    Cultural bias matters: A cross-cultural benchmark dataset and sentiment-enriched model for understanding multimodal metaphors,

    S. Yang, D. Zhang, J. Ren, Z. Xu, X. Zhang, Y . Song, H. Lin, and F. Xia, “Cultural bias matters: A cross-cultural benchmark dataset and sentiment-enriched model for understanding multimodal metaphors,” inProceedings of the 63rd Annual Meeting of the Association for Computatio...

  93. [101]

    Infochartqa: A benchmark for multimodal question answering on infographic charts,

    T. Xie, M. Lin, M. Liu, Y . Ye, C. Chen, and S. Liu, “Infochartqa: A benchmark for multimodal question answering on infographic charts,” inAdvances in Neural Information Processing Systems, D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, Eds., ...

  94. [102]

    Emometa: A multimodal dataset for fine-grained emotion classification in chinese metaphors,

    X. Lu, Y . Liu, D. Zhang, Z. Wu, J. Ren, and F. Xia, “Emometa: A multimodal dataset for fine-grained emotion classification in chinese metaphors,” inCompanion Proceedings of the ACM on Web Conference 2025, 2025, pp. 3080–3083

  95. [103]

    Towards multimodal metaphor understanding: a chinese dataset and model for metaphor mapping identification,

    D. Zhang, S. Yin, J. Yu, Z. Wu, Z. Li, C. Xu, X. Wang, and F. Xia, “Towards multimodal metaphor understanding: a chinese dataset and model for metaphor mapping identification,”ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 24, no. 12, pp. 1–25, 2025

  96. [104]

    Puzzled by puzzles: When vision-language models can’t take a hint,

    H. Lee, J. Ge, T.-H. Wu, M. Kang, T. Darrell, and D. M. Chan, “Puzzled by puzzles: When vision-language models can’t take a hint,” inProceedings of the 2025 Conference on Empirical Methods in 32 Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V...

  97. [105]

    BANMIME : Misogyny detection with metaphor explanation on Bangla memes,

    M. A. Mia, A. M. R. Mazumder, K. S. Sayma, M. Fahim, M. T. H. Fuad, M. I. Khan, and A. Rahman, “BANMIME : Misogyny detection with metaphor explanation on Bangla memes,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopou...

  98. [106]

    Looking beyond the pixels: Evaluating visual metaphor understanding in vlms,

    M. Kundu, S. Shekhar, and P. Bhattacharyya, “Looking beyond the pixels: Evaluating visual metaphor understanding in vlms,”Findings of the Association for Computational Linguistics: EMNLP, pp. 23 137– 23 158, 2025

  99. [107]

    Can audio language models listen between the lines? a study on metaphorical reasoning via unspoken,

    H. Xiao, X. Li, D. Pan, L. Zhang, Z. ZhixueSong, J. Han, S. Lai, W. Chen, J. Tang, and B. Wang, “Can audio language models listen between the lines? a study on metaphorical reasoning via unspoken,” in Proceedings of the 33rd ACM International Conference on Multimedia, 2025, pp...

  100. [108]

    SemEval-2022 task 6: iSarcasmEval, intended sarcasm detection in English and Arabic,

    I. Abu Farha, S. V . Oprea, S. Wilson, and W. Magdy, “SemEval-2022 task 6: iSarcasmEval, intended sarcasm detection in English and Arabic,” inProceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022), G. Emerson, N. Schluter, G. Stanovsky, R. Kumar, ...

  101. [109]

    Generalizable sarcasm detection is just around the corner, of course!

    H. Jang and D. Frassinelli, “Generalizable sarcasm detection is just around the corner, of course!” inProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), K. Duh,...

  102. [110]

    Sarcasmbench: Towards evaluating large language models on sarcasm understanding,

    Y . Zhang, C. Zou, Z. Lian, P. Tiwari, and J. Qin, “Sarcasmbench: Towards evaluating large language models on sarcasm understanding,” IEEE Transactions on Affective Computing, vol. 16, no. 4, pp. 2560– 2578, 2025

  103. [111]

    MultiPICo: Multilingual perspectivist irony corpus,

    S. Casola, S. Frenda, S. M. Lo, E. Sezerer, A. Uva, V . Basile, C. Bosco, A. Pedrani, C. Rubagotti, V . Patti, and D. Bernardi, “MultiPICo: Multilingual perspectivist irony corpus,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volu...

  104. [112]

    Rhetorical device-aware sarcasm detection with counterfactual data augmentation,

    Q. Hong, D. Zhang, J. Lin, D. Yin, S. Zhu, and J. Wang, “Rhetorical device-aware sarcasm detection with counterfactual data augmentation,” inFindings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna, Au...

  105. [113]

    FanChuan: A multilingual and graph-structured benchmark for parody detection and analysis,

    Y . Zheng, S. Li, F. Wu, Y . Ziyi, L. Hongchao, Z. Hu, C. Xinjun, Z. Wang, J. Chen, S. Luan, J. Xu, and L. Chen, “FanChuan: A multilingual and graph-structured benchmark for parody detection and analysis,” inFindings of the Association for Computational Linguistics: ACL 2025, ...

  106. [114]

    Multi-modal sarcasm generation: Dataset and solution,

    W. Zhao, Q. Huang, D. Xu, and P. Zhao, “Multi-modal sarcasm generation: Dataset and solution,” inFindings of the Association for Computational Linguistics: ACL 2023, A. Rogers, J. Boyd- Graber, and N. Okazaki, Eds. Toronto, Canada: Association for Computational Linguistics, Ju...

  107. [115]

    SarcNet: A multilingual multimodal sarcasm detection dataset,

    T. Yue, X. Shi, R. Mao, Z. Hu, and E. Cambria, “SarcNet: A multilingual multimodal sarcasm detection dataset,” inProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), N. Calzolari, M.-Y . Kan,...

  108. [116]

    Docmsu: A comprehensive benchmark for document-level multimodal sarcasm understanding,

    H. Du, G. Nan, S. Zhang, B. Xie, J. Xu, H. Fan, Q. Cui, X. Tao, and X. Jiang, “Docmsu: A comprehensive benchmark for document-level multimodal sarcasm understanding,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 16, p. 17933–17941, Mar. 2024. [Onl...

  109. [117]

    A multimodal framework to detect target aware aggression in memes,

    S. Ahsan, E. Hossain, O. Sharif, A. Das, M. M. Hoque, and M. Dewan, “A multimodal framework to detect target aware aggression in memes,” inProceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), Y . G...

  110. [118]

    ***YesBut***: A high-quality annotated multimodal dataset for evaluating satire comprehension capability of vision-language models,

    A. Nandy, Y . Agarwal, A. Patwa, M. M. Das, A. Bansal, A. Raj, P. Goyal, and N. Ganguly, “***YesBut***: A high-quality annotated multimodal dataset for evaluating satire comprehension capability of vision-language models,” inProceedings of the 2024 Conference on Empirical Meth...

  111. [119]

    Integrating stickers into multimodal dialogue summarization: A novel dataset and approach for enhancing social media interaction,

    Y . Shi and F. Kong, “Integrating stickers into multimodal dialogue summarization: A novel dataset and approach for enhancing social media interaction,” inProceedings of the 32nd ACM International Conference on Multimedia, ser. MM ’24. New York, NY , USA: Association for Compu...

  112. [120]

    Deciphering hate: Identifying hateful memes and their targets,

    E. Hossain, O. Sharif, M. M. Hoque, and S. M. Preum, “Deciphering hate: Identifying hateful memes and their targets,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L.-W. Ku, A. Martins, and V . Srikumar, Eds....

  113. [121]

    MMSD3.0: A multi-image benchmark for real-world multimodal sarcasm detection,

    H. Zhao, Y . Kong, Y . Xu, G. Gou, H. Xu, Y . Wang, and H. Zhang, “MMSD3.0: A multi-image benchmark for real-world multimodal sarcasm detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026. [Online]. Available: https://openaccess....

  114. [122]

    Towards probing speech- specific risks in large multimodal models: A taxonomy, benchmark, and insights,

    H. Yang, L. Qu, E. Shareghi, and R. Haf, “Towards probing speech- specific risks in large multimodal models: A taxonomy, benchmark, and insights,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Ch...

  115. [123]

    Toxictone: A mandarin audio dataset annotated for toxicity and toxic utterance tonality,

    Y .-X. Luo, Y .-C. Lin, M.-T. Chuang, J.-H. Chen, I.-N. Tsai, P. X. Kiew, Y .-H. Huang, C.-F. Liu, Y .-C. Chen, B.-H. Feng, W. Ren, and H. yi Lee, “Toxictone: A mandarin audio dataset annotated for toxicity and toxic utterance tonality,” 2025. [Online]. Available: https://arxi...

  116. [124]

    Optimism, expectation, or sarcasm? multi-class hope speech detection in spanish and english,

    S. Butt, F. Balouchzahi, A. I. Amjad, M. Amjad, H. G. Ceballos, and S. M. Jimenez-Zafra, “Optimism, expectation, or sarcasm? multi-class hope speech detection in spanish and english,” 2025. [Online]. Available: https://arxiv.org/abs/2504.17974

  117. [125]

    Leveraging large language models for sarcastic speech annotation in sarcasm detection,

    Z. Li, Y . Zhang, X. Gao, S. Nayak, and M. Coler, “Leveraging large language models for sarcastic speech annotation in sarcasm detection,”

  118. [126]

    Towards multimodal sarcasm detection (anObviously perfect paper),

    S. Castro, D. Hazarika, V . P ´erez-Rosas, R. Zimmermann, R. Mihalcea, and S. Poria, “Towards multimodal sarcasm detection (anObviously perfect paper),” inProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, A. Korhonen, D. Traum, and L. M `...

  119. [127]

    When did you become so smart, oh wise one?! sarcasm explanation in multi-modal multi-party dialogues,

    S. Kumar, A. Kulkarni, M. S. Akhtar, and T. Chakraborty, “When did you become so smart, oh wise one?! sarcasm explanation in multi-modal multi-party dialogues,” inProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S...

  120. [128]

    A Multimodal Chinese Dataset for Cross-lingual Sarcasm Detection,

    X. Gao, B. X. Wang, M. Zhang, S. Huang, Z. Li, S. Nayak, and 33 M. Coler, “A Multimodal Chinese Dataset for Cross-lingual Sarcasm Detection,” inInterspeech 2025, 2025, pp. 3968–3972

  121. [129]

    Multi- party empathetic dialogue generation: A new task for dialog systems,

    L. Zhu, Z. Zhang, J. Wang, H. Wang, H. Wu, and Z. Yang, “Multi- party empathetic dialogue generation: A new task for dialog systems,” inProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A....

  122. [130]

    A taxonomy of empathetic questions in social dialogs,

    E. Svikhnushina, I. V oinea, A. Welivita, and P. Pu, “A taxonomy of empathetic questions in social dialogs,” inProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), S. Muresan, P. Nakov, and A. Villavicencio, Eds. Dubl...

  123. [131]

    PAL to lend a helping hand: Towards building an emotion adaptive polite and empathetic counseling conversational agent,

    K. Mishra, P. Priya, and A. Ekbal, “PAL to lend a helping hand: Towards building an emotion adaptive polite and empathetic counseling conversational agent,” inProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Ro...

  124. [132]

    Apathetic or empathetic? evaluating llms’ emotional alignments with humans,

    J.-t. Huang, M. H. Lam, E. J. Li, S. Ren, W. Wang, W. Jiao, Z. Tu, and M. R. Lyu, “Apathetic or empathetic? evaluating llms’ emotional alignments with humans,”Advances in Neural Information Processing Systems, vol. 37, pp. 97 053–97 087, 2024

  125. [133]

    HEART-felt narratives: Tracing empathy and narrative style in personal stories with LLMs,

    J. Shen, J. Mire, H. W. Park, C. Breazeal, and M. Sap, “HEART-felt narratives: Tracing empathy and narrative style in personal stories with LLMs,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Ch...

  126. [134]

    Modeling empathetic alignment in conversation,

    J. Yang and D. Jurgens, “Modeling empathetic alignment in conversation,” inProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), K. Duh, H. Gomez, and S. Bethard, ...

  127. [135]

    EmotionQueen: A benchmark for evaluating empathy of large language models,

    Y . Chen, S. Yan, S. Liu, Y . Li, and Y . Xiao, “EmotionQueen: A benchmark for evaluating empathy of large language models,” in Findings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins, and V . Srikumar, Eds. Bangkok, Thailand: Association for ...

  128. [136]

    SYNTHEMPATHY: A scalable empathy corpus generated using LLMs without any crowdsourcing,

    R. Chen, J. Shin, and J. Hirschberg, “SYNTHEMPATHY: A scalable empathy corpus generated using LLMs without any crowdsourcing,”

  129. [137]

    Ecc: An emotion- cause conversation dataset for empathy response,

    Y . He, Y . Pan, W. Li, J. You, J. Deng, and F. Ren, “Ecc: An emotion- cause conversation dataset for empathy response,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp. 6011–6028

  130. [138]

    Empathy prediction from diverse perspectives,

    F. Chen, S. Carter, T. Lau, N. S. Bravo, S. Bhattacharyya, K. Sieck, and C. C. Wu, “Empathy prediction from diverse perspectives,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova,...

  131. [139]

    The pursuit of empathy: Evaluating small language models for ptsd dialogue support,

    S. Bn, Y . Mahajan, D. O. Mattioli, A. M. Sherrill, R. I. Arriaga, C. Wiese, and S. Abdullah, “The pursuit of empathy: Evaluating small language models for ptsd dialogue support,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, p...

  132. [140]

    Empathy-r1: A chain-of-empathy and reinforcement learn- ing framework for long-form mental health support,

    X. Yao, D. She, C. Zhang, Y . Zhang, Y . Sun, N. Ahmed, Y . Gao, and Z. Jin, “Empathy-r1: A chain-of-empathy and reinforcement learn- ing framework for long-form mental health support,”arXiv preprint arXiv:2509.14851, 2025

  133. [141]

    Beyond coarse labels: Fine- grained problem augmentation and multi-dimensional feedback for emotional support conversation,

    Y . Shi, J. Hao, and F. Kong, “Beyond coarse labels: Fine- grained problem augmentation and multi-dimensional feedback for emotional support conversation,” inFindings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, ...

  134. [142]

    TactfulToM: Do LLMs have the theory of mind ability to understand white lies?

    Y . Liu, E. J. Pretty, J. Huang, and S. Sugawara, “TactfulToM: Do LLMs have the theory of mind ability to understand white lies?” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V . P...

  135. [143]

    Sense-7: Taxonomy and dataset for measuring user perceptions of empathy in sustained human-ai conversations,

    J. Suh, L. Le, E. Shayegani, G. Ramos, J. Amores, D. C. Ong, M. Czerwinski, and J. Hernandez, “Sense-7: Taxonomy and dataset for measuring user perceptions of empathy in sustained human-ai conversations,”IEEE Transactions on Affective Computing, 2026

  136. [144]

    Kardia-r1: Unleashing llms to reason toward understanding and empathy for emotional support via rubric-as-judge reinforcement learning,

    J. Yuan, Z. Cui, H. Wang, Y . Gao, Y . Zhou, and U. Naseem, “Kardia-r1: Unleashing llms to reason toward understanding and empathy for emotional support via rubric-as-judge reinforcement learning,” inProceedings of the ACM Web Conference 2026, ser. WWW ’26. New York, NY , USA:...

  137. [145]

    STICKERCONV: Generating multimodal empathetic responses from scratch,

    Y . Zhang, F. Kong, P. Wang, S. Sun, L. Wang, S. Feng, D. Wang, Y . Zhang, and K. Song, “STICKERCONV: Generating multimodal empathetic responses from scratch,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L....

  138. [146]

    Empathyagent: Can embodied agents conduct empathetic actions?

    X. Chen, J. Ge, H. Dai, Q. Zhou, Q. Feng, J. Hu, Y . Wang, J. Liu, and S. Zhang, “Empathyagent: Can embodied agents conduct empathetic actions?”arXiv preprint arXiv:2503.16545, 2025

  139. [147]

    Generative expressive conversational speech synthesis,

    R. Liu, Y . Hu, Y . Ren, X. Yin, and H. Li, “Generative expressive conversational speech synthesis,” inProceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 4187–4196

  140. [148]

    Osum-echat: Enhancing end-to-end empathetic spoken chatbot via understanding-driven spoken dialogue,

    X. Geng, Q. Shao, H. Xue, S. Wang, H. Xie, Z. Guo, Y . Zhao, G. Li, W. Tian, C. Wanget al., “Osum-echat: Enhancing end-to-end empathetic spoken chatbot via understanding-driven spoken dialogue,” arXiv preprint arXiv:2508.09600, 2025

  141. [149]

    Echomind: An interrelated multi-level benchmark for evaluating empathetic speech language models,

    L. Zhou, L. Yu, Y . Lyu, Y . Lin, Z. Zhao, J. Ao, Y . Zhang, B. Wang, and H. Li, “Echomind: An interrelated multi-level benchmark for evaluating empathetic speech language models,” 2025. [Online]. Available: https://arxiv.org/abs/2510.22758

  142. [150]

    Aeq-bench: Measuring empathy of omni-modal large models,

    X. Luo, L. Yao, L. Zhao, L. Hong, K. Chen, D. Tao, D. Tan, R. Xu, and J. Li, “Aeq-bench: Measuring empathy of omni-modal large models,” arXiv preprint arXiv:2601.10513, 2026

  143. [151]

    MEDIC: A multimodal empathy dataset in counseling,

    Z. Zhu, X. Li, J. Pan, Y . Xiao, Y . Chang, F. Zheng, and S. Wang, “MEDIC: A multimodal empathy dataset in counseling,” in Proceedings of the 31st ACM International Conference on Multimedia. Association for Computing Machinery, 2023, pp. 6054–6062. [Online]. Available: https:/...

  144. [152]

    EmpathicStories++: A multimodal dataset for empathy towards personal experiences,

    J. Shen, Y . Kim, M. Hulse, W. Zulfikar, S. Alghowinem, C. Breazeal, and H. Park, “EmpathicStories++: A multimodal dataset for empathy towards personal experiences,” inFindings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins, and V . Srikumar, ...

  145. [153]

    Persuading across diverse domains: a dataset and persuasion large language model,

    C. Jin, K. Ren, L. Kong, X. Wang, R. Song, and H. Chen, “Persuading across diverse domains: a dataset and persuasion large language model,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L.-W. Ku, A. Martins, ...

  146. [154]

    Measuring bargaining abilities of LLMs: A benchmark and a buyer-enhancement method,

    T. Xia, Z. He, T. Ren, Y . Miao, Z. Zhang, Y . Yang, and R. Wang, “Measuring bargaining abilities of LLMs: A benchmark and a buyer-enhancement method,” inFindings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins, and V . Srikumar, Eds. Bangkok, ...

  147. [155]

    NegotiationToM: A benchmark for stress-testing machine theory of mind on negotiation surrounding,

    C. Chan, C. Jiayang, Y . Yim, Z. Deng, W. Fan, H. Li, X. Liu, H. Zhang, W. Wang, and Y . Song, “NegotiationToM: A benchmark for stress-testing machine theory of mind on negotiation surrounding,” inFindings of the Association for Computational Linguistics: EMNLP 2024, Y . Al-On...

  148. [156]

    SafePersuasion: A dataset, taxonomy, and baselines for analysis of rational persuasion and manipulation,

    H. Kong, A. M. M. Rahman, R. Tang, and V . Singh, “SafePersuasion: A dataset, taxonomy, and baselines for analysis of rational persuasion and manipulation,” inProceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the As...

  149. [157]

    PersuasiveToM: A benchmark for evaluating machine theory of mind in persuasive dialogues,

    F. Yu, L. Jiang, S. Huang, Z. Wu, and X. Dai, “PersuasiveToM: A benchmark for evaluating machine theory of mind in persuasive dialogues,” 2025. [Online]. Available: https://arxiv.org/abs/2502.21017

  150. [158]

    Debt collection negotiations with large language models: An evaluation system and optimizing decision making with multi-agent,

    X. Wang, Z. Zhang, J. G. Zheng, Y . Ai, and R. Wang, “Debt collection negotiations with large language models: An evaluation system and optimizing decision making with multi-agent,” inFindings of the Association for Computational Linguistics: ACL 2025, 2025, pp. 9555– 9577

  151. [159]

    Persuade me if you can: Evaluating ai agent influence on safety monitors,

    J. Za, J. Bainiaksina, T. Chopra, N. Ostrovsky, and V . Krakovna, “Persuade me if you can: Evaluating ai agent influence on safety monitors,” inICML 2025 Workshop on Reliable and Responsible Foundation Models, 2025

  152. [160]

    Persuasion should be double-blind: A multi-domain dialogue dataset with faith- fulness based on causal theory of mind,

    D. Zhang, L. Zhang, F. Qu, Z. Zhuang, and D. Zhou, “Persuasion should be double-blind: A multi-domain dialogue dataset with faith- fulness based on causal theory of mind,” inICASSP 2026-2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEE...

  153. [161]

    Persuasion strategies in advertisements,

    Y . Singla, R. Jha, A. Gupta, M. Aggarwal, A. Garg, T. Malyan, A. Bhardwaj, R. R. Shah, B. Krishnamurthy, and C. Chen, “Persuasion strategies in advertisements,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, pp. 57–66, 06 2023

  154. [162]

    SemEval-2024 task 4: Multilingual detection of persuasion techniques in memes,

    D. Dimitrov, F. Alam, M. Hasanain, A. Hasnat, F. Silvestri, P. Nakov, and G. Da San Martino, “SemEval-2024 task 4: Multilingual detection of persuasion techniques in memes,” inProceedings of the 18th International Workshop on Semantic Evaluation (SemEval-2024), A. K. Ojha, A. ...

  155. [163]

    PVP: An image dataset for personalized visual persuasion with persuasion strategies, viewer characteristics, and persuasiveness ratings,

    J. Kim, J. Han, D. Choi, J. Yoon, E.-J. Lee, and Y . Jo, “PVP: An image dataset for personalized visual persuasion with persuasion strategies, viewer characteristics, and persuasiveness ratings,” inProceedings of the 63rd Annual Meeting of the Association for Computational Lin...

  156. [164]

    A corpus and cloze evaluation for deeper understanding of commonsense stories,

    N. Mostafazadeh, N. Chambers, X. He, D. Parikh, D. Batra, L. Vanderwende, P. Kohli, and J. Allen, “A corpus and cloze evaluation for deeper understanding of commonsense stories,” inProceedings of the 2016 Conference of the North American Chapter of the Association for Computat...

  157. [165]

    Event2Mind: Commonsense inference on events, intents, and reactions,

    H. Rashkin, M. Sap, E. Allaway, N. A. Smith, and Y . Choi, “Event2Mind: Commonsense inference on events, intents, and reactions,” inProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), I. Gurevych and Y . Miyao, Eds. ...

  158. [166]

    The NarrativeQA reading comprehension challenge,

    T. Ko ˇcisk´y, J. Schwarz, P. Blunsom, C. Dyer, K. M. Hermann, G. Melis, and E. Grefenstette, “The NarrativeQA reading comprehension challenge,”Transactions of the Association for Computational Linguistics, vol. 6, pp. 317–328, 2018. [Online]. Available: https: //aclanthology....

  159. [167]

    Modeling naive psychology of characters in simple commonsense stories,

    H. Rashkin, A. Bosselut, M. Sap, K. Knight, and Y . Choi, “Modeling naive psychology of characters in simple commonsense stories,” inProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), I. Gurevych and Y . Miyao, Eds....

  160. [168]

    “let your characters tell their story

    F. Brahman, M. Huang, O. Tafjord, C. Zhao, M. Sachan, and S. Chaturvedi, ““let your characters tell their story”: A dataset for character-centric narrative understanding,” inFindings of the Association for Computational Linguistics: EMNLP 2021, M.-F. Moens, X. Huang, L. Specia...

  161. [169]

    Fantastic questions and where to find them: FairytaleQA – an authentic dataset for narrative comprehension,

    Y . Xu, D. Wang, M. Yu, D. Ritchie, B. Yao, T. Wu, Z. Zhang, T. J.-J. Li, N. Bradford, B. Sun, T. B. Hoang, Y . Sang, Y . Hou, X. Ma, D. Yang, N. Peng, Z. Yu, and M. Warschauer, “Fantastic questions and where to find them: FairytaleQA – an authentic dataset for narrative compr...

  162. [170]

    LOT: A story-centric benchmark for evaluating Chinese long text understanding and generation,

    J. Guan, Z. Feng, Y . Chen, R. He, X. Mao, C. Fan, and M. Huang, “LOT: A story-centric benchmark for evaluating Chinese long text understanding and generation,”Transactions of the Association for Computational Linguistics, vol. 10, pp. 434–451, 2022. [Online]. Available: https...

  163. [171]

    One thousand and one pairs: A “novel

    M. Karpinska, K. Thai, K. Lo, T. Goyal, and M. Iyyer, “One thousand and one pairs: A “novel” challenge for long-context language models,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds. ...

  164. [172]

    DetectBench: Can large language model detect and piece together implicit evidence?

    Z. Gu, L. Zhang, X. Zhu, J. Chen, W. Huang, Y . Zhang, S. Wang, Z. Ye, Y . Gao, H. Feng, and Y . Xiao, “DetectBench: Can large language model detect and piece together implicit evidence?” inFindings of the Association for Computational Linguistics: EMNLP 2024, Y . Al- Onaizan,...

  165. [173]

    Where do people tell stories online? story detection across online communities,

    M. Antoniak, J. Mire, M. Sap, E. Ash, and A. Piper, “Where do people tell stories online? story detection across online communities,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L.-W. Ku, A. Martins, and V ...

  166. [174]

    Detectiveqa: Evaluating long- context reasoning on detective novels,

    Z. Xu, J. Ye, X. Liu, X. Liu, T. Sun, Z. Liu, Q. Guo, L. Li, Q. Liu, X. Huang, and X. Qiu, “Detectiveqa: Evaluating long- context reasoning on detective novels,” 2025. [Online]. Available: https://arxiv.org/abs/2409.02465

  167. [175]

    Prelude: A benchmark designed to require global comprehension and reasoning over long contexts,

    M. Yu, T. T. Chung, C. Zhou, T. Li, R. Lu, J. Li, L. Xu, H. Lu, N. Zhang, J. Li, and J. Zhou, “Prelude: A benchmark designed to require global comprehension and reasoning over long contexts,”

  168. [176]

    NovelHopQA: Diagnosing multi-hop reasoning failures in long narrative contexts,

    A. Gupta, K. Zhu, V . Sharma, S. O’Brien, and M. Lu, “NovelHopQA: Diagnosing multi-hop reasoning failures in long narrative contexts,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and ...

  169. [177]

    CHATTER: A character-attribution dataset for narrative understanding,

    S. Baruah and S. Narayanan, “CHATTER: A character-attribution dataset for narrative understanding,” inProceedings of the The 7th Workshop on Narrative Understanding, E. Clark, Y . K. Lal, S. Chaturvedi, M. Iyyer, A. Brei, A. Modi, and K. R. Chandu, Eds. Albuquerque, New Mexico...

  170. [178]

    Whodunit: Evaluation benchmark for culprit detection in mystery stories,

    K. Gupta, “Whodunit: Evaluation benchmark for culprit detection in mystery stories,” 2025. [Online]. Available: https://arxiv.org/abs/2502. 07747

  171. [179]

    Available: https://arxiv.org/abs/2508.09848

    [Online]. Available: https://arxiv.org/abs/2508.09848

  172. [180]

    TurnaboutLLM: A deductive reasoning benchmark from detective games,

    Y . Yuan, M. He, M. A. Shahid, Z. Li, J. Huang, and L. Zhang, “TurnaboutLLM: A deductive reasoning benchmark from detective games,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V ....

  173. [181]

    Beyond single frames: Can lmms comprehend 35 implicit narratives in comic strip?

    X. Wang, H. Xia, J. Song, L. Guan, Q. Dong, R. Li, Y . Yang, W. Luo, Y . Wang, Y . Puet al., “Beyond single frames: Can lmms comprehend 35 implicit narratives in comic strip?” inFindings of the Association for Computational Linguistics: EMNLP 2025, 2025, pp. 6436–6452

  174. [182]

    From panels to prose: Generating literary narratives from comics,

    R. Sachdeva and A. Zisserman, “From panels to prose: Generating literary narratives from comics,” inProceedings of the IEEE/CVF In- ternational Conference on Computer Vision, 2025, pp. 21 864–21 873

  175. [183]

    Finding flawed fictions: Evalu- ating complex reasoning in language models via plot hole detection,

    K. Ahuja, M. Sclar, and Y . Tsvetkov, “Finding flawed fictions: Evalu- ating complex reasoning in language models via plot hole detection,” arXiv preprint arXiv:2504.11900, 2025

  176. [184]

    Is your image a good storyteller?

    X. Song, X. Pang, H. Tang, M. Wu, and K. Q. Zhu, “Is your image a good storyteller?”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 24, pp. 25 165–25 173, Apr. 2025. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/34702

  177. [185]

    Rˆ 3-vqa:

    L. Niu, J. Li, X. Yu, S. Wang, R. Feng, B. Wu, P. Wei, Y . Wang, and L. Fan, “Rˆ 3-vqa:” read the room” by video social reasoning,”arXiv preprint arXiv:2505.04147, 2025

  178. [186]

    V-ALPHASOCIAL: Benchmark and self-reflective chain-of- thought generation for visual social commonsense reasoning,

    Z. Lin, Z. Xu, X. Song, Y . Wan, X. Yao, T.-H. Lin, S. Song, P. Subbaraman, B. Zhou, K.-W. Chang, and Y . Sun, “V-ALPHASOCIAL: Benchmark and self-reflective chain-of- thought generation for visual social commonsense reasoning,” inFindings of the Association for Computational L...

  179. [187]

    A cognitive evaluation benchmark of image reasoning and description for large vision-language models,

    X. Song, M. Wu, K. Q. Zhu, C. Zhang, and Y . Chen, “A cognitive evaluation benchmark of image reasoning and description for large vision-language models,” inProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguisti...

  180. [188]

    Seriesbench: A benchmark for narrative-driven drama series under- standing,

    C. Zhang, Y . Lei, Z. Liu, H. Leng, S. Liu, T. Gao, Q. Liu, and Y . Wang, “Seriesbench: A benchmark for narrative-driven drama series under- standing,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2025, pp. 28 995–29 004

  181. [189]

    Vrbench: A benchmark for multi-step reasoning in long narrative videos,

    J. Yu, Y . Wu, M. Chu, Z. Ren, Z. Huang, P. Chu, R. Zhang, Y . He, Q. Li, S. Liet al., “Vrbench: A benchmark for multi-step reasoning in long narrative videos,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 21 655–21 666

  182. [190]

    Sagaqa: A multi-hop reasoning benchmark for long-form narrative understanding in tv series,

    G. Pennec, Z. Liu, N. Asher, P. Muller, and N. F. Chen, “Sagaqa: A multi-hop reasoning benchmark for long-form narrative understanding in tv series,” 2026. [Online]. Available: https: //arxiv.org/abs/2606.03301

  183. [191]

    MoMentS: A comprehensive multimodal benchmark for theory of mind,

    E. Villa-Cueva, S. M. M. Ahmed, R. Chevi, J. C. B. Cruz, K. Elzeky, F. Cristobal, A. F. Aji, S. Wang, R. Mihalcea, and T. Solorio, “MoMentS: A comprehensive multimodal benchmark for theory of mind,” inFindings of the Association for Computational Linguistics: EMNLP 2025, C. Ch...

  184. [192]

    ComicVQA: A benchmark for visual reasoning in multimodal LLMs,

    E. Gan, H. Brown, D. Herel, K. Kawaguchi, M.-Y . Kan, and M. Q. Shieh, “ComicVQA: A benchmark for visual reasoning in multimodal LLMs,” inFindings of the Association for Computational Linguistics: ACL 2026, M. Liakata, V . P. Moreira, J. Zhang, and D. Jurgens, Eds. San Diego, ...

  185. [193]

    Movie101v2: Improved movie narration benchmark,

    Z. Yue, Y . Zhang, Z. Wang, and Q. Jin, “Movie101v2: Improved movie narration benchmark,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna, Austria...

  186. [194]

    Context-situated pun generation,

    J. Sun, A. Narayan-Chen, S. Oraby, S. Gao, T. Chung, J. Huang, Y . Liu, and N. Peng, “Context-situated pun generation,” inProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022, pp. 4635–4648

  187. [195]

    Narrativetrack: Evaluating entity-centric reasoning for narrative understanding,

    H. Ha, J. Ge, B. Feng, K. Ma, and G. Chakraborty, “Narrativetrack: Evaluating entity-centric reasoning for narrative understanding,” 2026. [Online]. Available: https://arxiv.org/abs/2601.01095

  188. [196]

    Potential idiomatic expression (PIE)-English: Corpus for classes of idioms,

    T. Adewumi, R. Vadoodi, A. Tripathy, K. Nikolaido, F. Liwicki, and M. Liwicki, “Potential idiomatic expression (PIE)-English: Corpus for classes of idioms,” inProceedings of the Thirteenth Language Resources and Evaluation Conference, N. Calzolari, F. B ´echet, P. Blache, K. C...

  189. [197]

    ANALOGICAL - a novel benchmark for long text analogy evaluation in large language models,

    T. Wijesiriwardene, R. Wickramarachchi, B. Gajera, S. Gowaikar, C. Gupta, A. Chadha, A. N. Reganti, A. Sheth, and A. Das, “ANALOGICAL - a novel benchmark for long text analogy evaluation in large language models,” inFindings of the Association for Computational Linguistics: AC...

  190. [198]

    Multilingual multi-figurative language detection,

    H. Lai, A. Toral, and M. Nissim, “Multilingual multi-figurative language detection,” inFindings of the Association for Computational Linguistics: ACL 2023, A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds. Toronto, Canada: Association for Computational Linguistics, Jul. 2023, pp...

  191. [199]

    ExPUNations: Augmenting puns with keywords and explanations,

    J. Sun, A. Narayan-Chen, S. Oraby, A. Cervone, T. Chung, J. Huang, Y . Liu, and N. Peng, “ExPUNations: Augmenting puns with keywords and explanations,” inProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y . Goldberg, Z. Kozareva, and Y . ...

  192. [200]

    Do LLMs understand social knowledge? evaluating the sociability of large language models with SocKET benchmark,

    M. Choi, J. Pei, S. Kumar, C. Shu, and D. Jurgens, “Do LLMs understand social knowledge? evaluating the sociability of large language models with SocKET benchmark,” inProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, a...

  193. [201]

    Think before you speak: Cultivating communication skills of large language models via inner monologue,

    J. Zhou, L. Pang, H. Shen, and X. Cheng, “Think before you speak: Cultivating communication skills of large language models via inner monologue,” inFindings of the Association for Computational Linguistics: NAACL 2024, K. Duh, H. Gomez, and S. Bethard, Eds. Mexico City, Mexico...

  194. [202]

    Are U a joke master? pun generation via multi-stage curriculum learning towards a humor LLM,

    Y . Chen, C. Yang, T. Hu, X. Chen, M. Lan, L. Cai, X. Zhuang, X. Lin, X. Lu, and A. Zhou, “Are U a joke master? pun generation via multi-stage curriculum learning towards a humor LLM,” in Findings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins...

  195. [203]

    Diplomat: A dialogue dataset for situated pragmatic reasoning,

    H. Li, S.-C. Zhu, and Z. Zheng, “Diplomat: A dialogue dataset for situated pragmatic reasoning,”Advances in Neural Information Processing Systems, vol. 36, pp. 46 856–46 884, 2023

  196. [204]

    When llms meet cunning texts: A fallacy understanding benchmark for large language models,

    Y . Li, Q. Zhou, Y . Luo, S. Ma, Y . Li, H.-T. Zheng, X. Hu, and P. S. Yu, “When llms meet cunning texts: A fallacy understanding benchmark for large language models,”Advances in Neural Information Processing Systems, vol. 37, pp. 112 433–112 458, 2024

  197. [205]

    Proverbs run in pairs: Evaluating proverb translation capability of large language model,

    M. Wang, V . T. Pham, F. Moghimifar, and T.-T. Vu, “Proverbs run in pairs: Evaluating proverb translation capability of large language model,” inFindings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna...

  198. [206]

    Fluid qa: A multilingual benchmark for figurative language usage in dialogue across english, chinese, and korean,

    S. Park, H. Choi, M. Kim, S. An, X. Wang, G. Choi, and H. Kim, “Fluid qa: A multilingual benchmark for figurative language usage in dialogue across english, chinese, and korean,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 2025, pp...

  199. [207]

    SLANG: New concept comprehension of large language models,

    L. Mei, S. Liu, Y . Wang, B. Bi, and X. Cheng, “SLANG: New concept comprehension of large language models,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds. Miami, Florida, USA: Associati...

  200. [208]

    Being kind isn’t always being safe: Diagnosing affective hallucination in LLMs,

    S. Kim, J. Kim, S. Shin, H. Chung, D. Moon, Y . Kwon, and H. Yoon, “Being kind isn’t always being safe: Diagnosing affective hallucination in LLMs,” inFindings of the Association for 36 Computational Linguistics: EACL 2026, V . Demberg, K. Inui, and L. Marquez, Eds. Rabat, Mor...

  201. [209]

    Redefining machine translation on social network services with large language models,

    H. Guo, F. Zhao, S. Cao, X. Lyu, Z. Liu, Y . Wang, B. Wang, Z. Li, C. Lu, Z. Xuet al., “Redefining machine translation on social network services with large language models,”arXiv preprint arXiv:2504.07901, 2025

  202. [210]

    Pun2Pun: Benchmarking LLMs on textual-visual Chinese-English pun translation via pragmatics model and linguistic reasoning,

    Y . R. Ma, S. Huang, Y . Xu, Z. Zhou, and Y . Wei, “Pun2Pun: Benchmarking LLMs on textual-visual Chinese-English pun translation via pragmatics model and linguistic reasoning,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4:...

  203. [211]

    Beyond context to cognitive appraisal: Emotion reasoning as a theory of mind benchmark for large language models,

    G. C. Yeo and K. Jaidka, “Beyond context to cognitive appraisal: Emotion reasoning as a theory of mind benchmark for large language models,” inFindings of the Association for Computational Linguistics: ACL 2025, 2025, pp. 26 517–26 525

  204. [212]

    Can large language models understand Internet buzzwords through user-generated content,

    C. Huang, J. Luo, X. Wang, W. Lei, and J. Lv, “Can large language models understand Internet buzzwords through user-generated content,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shu...

  205. [213]

    FigMemes: A dataset for figurative language identification in politically-opinionated memes,

    C. Liu, G. Geigle, R. Krebs, and I. Gurevych, “FigMemes: A dataset for figurative language identification in politically-opinionated memes,” inProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y . Goldberg, Z. Kozareva, and Y . Zhang, Eds....

  206. [214]

    SMILE: Multimodal dataset for understanding laughter in video with language models,

    L. Hyun, K. Sung-Bin, S. Han, Y . Yu, and T.-H. Oh, “SMILE: Multimodal dataset for understanding laughter in video with language models,” inFindings of the Association for Computational Linguistics: NAACL 2024, K. Duh, H. Gomez, and S. Bethard, Eds. Mexico City, Mexico: Associ...

  207. [215]

    Beyond literal mapping: Benchmarking and improving non-literal translation evaluation,

    Y . Tian, C. Wang, Z. Liu, H. Huang, W. Yu, D. Song, J. Tang, and Y . Guo, “Beyond literal mapping: Benchmarking and improving non-literal translation evaluation,” 2026. [Online]. Available: https://arxiv.org/abs/2601.07338

  208. [216]

    Ii-bench: An image implication understanding benchmark for multimodal large language models,

    Z. Liu, F. Fang, X. Feng, X. Du, C. Zhang, N. Wang, Q. Zhao, L. Fan, C. GAN, H. Linet al., “Ii-bench: An image implication understanding benchmark for multimodal large language models,”Advances in Neural Information Processing Systems, vol. 37, pp. 46 378–46 480, 2024

  209. [217]

    Understanding figurative meaning through explainable visual entailment,

    A. Saakyan, S. Kulkarni, T. Chakrabarty, and S. Muresan, “Understanding figurative meaning through explainable visual entailment,” inProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Techn...

  210. [218]

    Insightvision: A comprehensive, multi-level chinese-based benchmark for evaluating implicit visual semantics in large vision language models,

    X. Yin, Y . Hong, Y . Guo, Y . Tu, W. Wang, G. Liuet al., “Insightvision: A comprehensive, multi-level chinese-based benchmark for evaluating implicit visual semantics in large vision language models,”arXiv preprint arXiv:2502.15812, 2025

  211. [219]

    Can large multimodal models uncover deep semantics behind images?

    Y . Yang, Z. Li, Q. Dong, H. Xia, and Z. Sui, “Can large multimodal models uncover deep semantics behind images?” inFindings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins, and V . Srikumar, Eds. Bangkok, Thailand: Association for Computationa...

  212. [220]

    Impact of stickers on multimodal sen- timent and intent in social media: A new task, dataset and baseline,

    Y . Shi, F. Kong, and L. Zhang, “Impact of stickers on multimodal sen- timent and intent in social media: A new task, dataset and baseline,” in Proceedings of the 33rd ACM International Conference on Multimedia, 2025, pp. 5637–5646

  213. [221]

    Svbench: Evaluation of video generation models on social reasoning,

    W. Peng, G. Wang, T. Yang, C. Li, X. Xu, H. He, and K. Zhang, “Svbench: Evaluation of video generation models on social reasoning,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2026, pp. 32 872–32 881

  214. [222]

    Can MLLMs understand the deep implication behind Chinese images?

    C. Zhang, X. Feng, Y . Bai, X. Du, J. Hou, K. Deng, G. Han, Q. Li, B. Wang, J. Liu, X. Qu, Y . Zhang, Q. Zhao, Y . Liang, Z. Liu, F. Fang, M. Yang, W. Huang, C. Lin, G. Zhang, and S. Ni, “Can MLLMs understand the deep implication behind Chinese images?” inProceedings of the 63...

  215. [223]

    Figurative-cum- commonsense knowledge infusion for multimodal mental health meme classification,

    A. Mazhar, Z. H. Shaik, A. Srivastava, P. Ruhnke, L. Vaddavalli, S. K. Katragadda, S. Yadav, and M. S. Akhtar, “Figurative-cum- commonsense knowledge infusion for multimodal mental health meme classification,” inProceedings of the ACM on Web Conference 2025, ser. WWW ’25. New ...

  216. [224]

    The hateful memes challenge: Detecting hate speech in multimodal memes,

    D. Kiela, H. Firooz, A. Mohan, V . Goswami, A. Singh, P. Ringshia, and D. Testuggine, “The hateful memes challenge: Detecting hate speech in multimodal memes,” inAdvances in Neural Information Processing Systems, vol. 33, 2020. [Online]. Available: https://proceedings.neurips....

  217. [225]

    PunMemeCN: A benchmark to explore vision-language models’ understanding of Chinese pun memes,

    Z. Xu, S. Yuan, Y . Zhang, J. Sun, T. Zheng, and D. Yang, “PunMemeCN: A benchmark to explore vision-language models’ understanding of Chinese pun memes,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakrab...

  218. [226]

    PunchBench: Benchmarking MLLMs in multimodal punchline comprehension,

    K. Ouyang, Y . Liu, S. Li, Y . Liu, H. Zhou, F. Meng, J. Zhou, and X. Sun, “PunchBench: Benchmarking MLLMs in multimodal punchline comprehension,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for...

  219. [227]

    VIBE: Can a VLM read the room?

    T. Chakraborty, E. Caplan, and D. Goldwasser, “VIBE: Can a VLM read the room?” inFindings of the Association for Computational Linguistics: EMNLP 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V . Peng, Eds. Suzhou, China: Association for Computational Linguistics, ...

  220. [228]

    Goat-bench: Safety insights to large multimodal models through meme-based social abuse,

    H. Lin, Z. Luo, B. Wang, R. Yang, and J. Ma, “Goat-bench: Safety insights to large multimodal models through meme-based social abuse,”ACM Trans. Intell. Syst. Technol., vol. 17, no. 4, Apr. 2026. [Online]. Available: https://doi.org/10.1145/3729239

  221. [229]

    Discrimination of the dif- ferent intents carried by the same text through integrating multimodal information

    Z. Li, G. Zhang, L. Wang, and J. Dang, “Discrimination of the dif- ferent intents carried by the same text through integrating multimodal information.” inINTERSPEECH, 2023, pp. 2423–2427

  222. [230]

    Mmar: A challenging benchmark for deep reasoning in speech, audio, music, and their mix,

    Z. Ma, Y . Ma, Y . Zhu, C. Yang, Y .-W. Chao, R. Xu, W. Chen, Y . Chen, Z. Chen, J. Cong, K. Li, K. Li, S. Li, X. Li, X. Li, Z. Lian, Y . Liang, M. Liu, Z. Niu, T. Wang, W. Yuping, Y . Wang, Y . Wu, G. Yang, J. Yu, R. Yuan, Z. Zheng, Z. Zhou, H. Zhu, W. Xue, E. Benetos, K. Yu,...

  223. [231]

    Paras2s: Benchmarking and aligning spoken language models for paralinguistic-aware speech-to-speech interaction,

    S.-w. Yang, M. Tu, A. T. Liu, X. Qu, H.-y. Lee, L. Lu, Y . Wang, and Y . Wu, “Paras2s: Benchmarking and aligning spoken language models for paralinguistic-aware speech-to-speech interaction,”arXiv preprint arXiv:2511.08723, 2025

  224. [232]

    Are large language models chronically online surfers? a dataset for Chinese Internet meme explanation,

    Y . Xie, C. Wang, Z. Ma, and F. Miao, “Are large language models chronically online surfers? a dataset for Chinese Internet meme explanation,” inProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Ro...

  225. [233]

    Genesis: A large-scale benchmark for mul- timodal large language model in emotional causality analysis,

    Y . Li, Y . Zhang, R. Chen, F. Tang, Z. Lu, M. Hu, J. Wu, H. Xue, M. Zhou, C. Liet al., “Genesis: A large-scale benchmark for mul- timodal large language model in emotional causality analysis,” in Proceedings of the 33rd ACM International Conference on Multimedia, 2025, pp. 12...

  226. [234]

    DRT: Deep reasoning translation via long chain-of-thought,

    J. Wang, F. Meng, Y . Liang, and J. Zhou, “DRT: Deep reasoning translation via long chain-of-thought,” inFindings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna, Austria: Association for Computational...

  227. [235]

    Generating storytelling images with rich chains-of-reasoning,

    X. Song, Q. Jia, S. Watanabe, X. Pang, R. Chen, M. Wu, and K. Q. Zhu, “Generating storytelling images with rich chains-of-reasoning,” arXiv preprint arXiv:2512.07198, 2025

  228. [236]

    Boss: Beyond-semantic speech,

    Q. Wang, Z. Li, H. Lv, H. Chen, Y . Song, J. Kang, J. Lian, J. Li, Y . Li, Z. He, and X. Li, “Boss: Beyond-semantic speech,” 2025. [Online]. Available: https://arxiv.org/abs/2507.17563

  229. [237]

    Plast: Towards paralinguistic-aware speech translation,

    Y . Li, R. Zhao, R. Zhang, J. Su, D. Wei, M. Zhang, and Y . Chen, “Plast: Towards paralinguistic-aware speech translation,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 38, 2026, pp. 31 805–31 813

  230. [238]

    “a good pun is its own reword

    Z. Xu, S. Yuan, L. Chen, and D. Yang, ““a good pun is its own reword”: Can large language models understand puns?” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds. Miami, Florida, USA: A...

  231. [239]

    Interarm: Interpretable affective reasoning model for multimodal sarcasm detection,

    T. Yue, R. Mao, X. Shi, and E. Cambria, “Interarm: Interpretable affective reasoning model for multimodal sarcasm detection,”IEEE Transactions on Affective Computing, pp. 1–12, 2026

  232. [240]

    Getting serious about humor: Crafting humor datasets with unfunny large language models,

    Z. Horvitz, J. Chen, R. Aditya, H. Srivastava, R. West, Z. Yu, and K. McKeown, “Getting serious about humor: Crafting humor datasets with unfunny large language models,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short ...

  233. [241]

    Collaborative multi-agent scripts generation for enhancing imperfect-information reasoning in murder mystery games,

    K. Zhong, J. Xie, H. Wu, H. Li, and G. Li, “Collaborative multi-agent scripts generation for enhancing imperfect-information reasoning in murder mystery games,” inFindings of the Association for Computational Linguistics: ACL 2026, M. Liakata, V . P. Moreira, J. Zhang, and D. ...

  234. [242]

    Paralinguistics- enhanced large language modeling of spoken dialogue,

    G.-T. Lin, P. G. Shivakumar, A. Gandhe, C.-H. H. Yang, Y . Gu, S. Ghosh, A. Stolcke, H.-y. Lee, and I. Bulyko, “Paralinguistics- enhanced large language modeling of spoken dialogue,” inICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (I...

  235. [243]

    Uplme: Uncertainty-aware probabilistic language modelling for robust empathy regression,

    M. R. Hasan, M. Z. Hossain, A. Krishna, S. Rahman, and T. Gedeon, “Uplme: Uncertainty-aware probabilistic language modelling for robust empathy regression,”arXiv preprint arXiv:2508.03520, 2025

  236. [244]

    Analyzing offensive language dataset insights from training dynamics and human agreement level,

    D.-K. Kim, H. Ahn, Y . Kim, and Y .-S. Han, “Analyzing offensive language dataset insights from training dynamics and human agreement level,” inProceedings of the 31st International Conference on Computational Linguistics, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B....

  237. [245]

    Metaphor and large language models: When surface features matter more than deep understanding,

    E. Sanchez-Bayona and R. Agerri, “Metaphor and large language models: When surface features matter more than deep understanding,” inFindings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vienna, Austria: As...

  238. [246]

    When to laugh and how hard? a multimodal approach to detecting humor and its intensity,

    K. Alnajjar, M. H ¨am¨al¨ainen, J. Tiedemann, J. Laaksonen, and M. Kurimo, “When to laugh and how hard? a multimodal approach to detecting humor and its intensity,” inProceedings of the 29th International Conference on Computational Linguistics, N. Calzolari, C.-R. Huang, H. K...

  239. [247]

    Who laughs with whom? disentangling influential factors in humor preferences across user clusters and LLMs,

    S. Murakami, H. Kamigaito, H. Takamura, and M. Okumura, “Who laughs with whom? disentangling influential factors in humor preferences across user clusters and LLMs,” 2026. [Online]. Available: https://arxiv.org/abs/2601.03103

  240. [248]

    Multi-dimensional evaluation of empathetic dialogue responses,

    Z. Xu and J. Jiang, “Multi-dimensional evaluation of empathetic dialogue responses,” inFindings of the Association for Computational Linguistics: EMNLP 2024. Association for Computational Linguistics, 2024, pp. 2066–2087. [Online]. Available: https://aclanthology.org/ 2024.fin...

  241. [249]

    Metaphor detection via linguistics enhanced siamese network,

    S. Zhang, Y . Liu, J. Chen, Y . Liu, and M. Sun, “Metaphor detection via linguistics enhanced siamese network,” inProceedings of the 29th International Conference on Computational Linguistics. International Committee on Computational Linguistics, 2022, pp. 4149–4159. [Online]....

  242. [250]

    Hatred stems from ignorance! distillation of the persuasion modes in countering conversational hate speech,

    G. Alyahya and A. Aldayel, “Hatred stems from ignorance! distillation of the persuasion modes in countering conversational hate speech,” in Proceedings of the International AAAI Conference on Web and Social Media, vol. 19, 2025, pp. 52–67

  243. [251]

    Modeling conceptual attribute likeness and domain inconsistency for metaphor detection,

    Y . Tian, N. Xu, W. Mao, and D. Zeng, “Modeling conceptual attribute likeness and domain inconsistency for metaphor detection,” inProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Associati...

  244. [252]

    Empathy intent drives empathy detection,

    L. Jiang, D. Wu, B. Mao, Y . Li, and W. Slamu, “Empathy intent drives empathy detection,” inProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 2023, pp. 6279–6290. [Online]. Available: https://acla...

  245. [253]

    Enhancing semantic awareness by sentimental constraint with automatic outlier masking for multimodal sarcasm detection,

    S. Yuan, Y . Wei, H. Zhou, Q. Xu, M. Chen, and X. He, “Enhancing semantic awareness by sentimental constraint with automatic outlier masking for multimodal sarcasm detection,”IEEE Transactions on Multimedia, vol. 27, pp. 5376–5386, 2025

  246. [254]

    Clcl: Non-compositional expres- sion detection with contrastive learning and curriculum learning,

    J. Zhou, Z. Zeng, and S. Bhat, “Clcl: Non-compositional expres- sion detection with contrastive learning and curriculum learning,” inProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2023, pp. 730– 743

  247. [255]

    Learn from failure: Causality-guided contrastive learning for generalizable implicit hate speech detection,

    T. Jiang, “Learn from failure: Causality-guided contrastive learning for generalizable implicit hate speech detection,” inProceedings of the 31st International Conference on Computational Linguistics, O. Rambow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, and S. Sc...

  248. [256]

    Multimodal propaganda detection via anti-persuasion prompt enhanced contrastive learning,

    J. Cui, L. Li, X. Zhang, and J. Yuan, “Multimodal propaganda detection via anti-persuasion prompt enhanced contrastive learning,” inICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5

  249. [257]

    Modeling highlighting of metaphors in multitask contrastive learning paradigms,

    M. Sengupta, M. Alshomary, I. Scharlau, and H. Wachsmuth, “Modeling highlighting of metaphors in multitask contrastive learning paradigms,” inFindings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association fo...

  250. [258]

    Innovative thinking, infinite humor: Humor research of large language models through structured thought leaps,

    H. Wang, Y . Zhao, D. Li, X. Wang, sinbadliu, X. Lan, and H. Wang, “Innovative thinking, infinite humor: Humor research of large language models through structured thought leaps,” inThe Thirteenth International Conference on Learning Representations, 2025. [Online]. Available:...

  251. [259]

    Sarcasm-R1: Enhancing sarcasm detection through focused reasoning,

    Q. Yang, J. Zeng, L. Yang, K. Ma, and H. Lin, “Sarcasm-R1: Enhancing sarcasm detection through focused reasoning,” inFindings of the Association for Computational Linguistics: EMNLP 2025. Association for Computational Linguistics, 2025, pp. 10 773–10 785. [Online]. Available: ...

  252. [260]

    Emotion-o1: Adaptive long reasoning for emotion understanding in llms,

    C. Song, Y . Zhang, H. Gao, K. Huang, and P. Zhang, “Emotion-o1: Adaptive long reasoning for emotion understanding in llms,”arXiv preprint arXiv:2505.22548, 2025

  253. [261]

    Exploring chain-of-thought for multi-modal metaphor detection,

    Y . Xu, Y . Hua, S. Li, and Z. Wang, “Exploring chain-of-thought for multi-modal metaphor detection,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 2024, pp. 91–101....

  254. [262]

    Bridging the creativity understanding gap: Small-scale human alignment enables expert-level humor ranking in LLMs,

    K. L. Zhou, J. Chen, S. Suresh, R. Narad, T. T. Rogers, L. K. Jain, R. D. Nowak, B. Mankoff, and J. Zhang, “Bridging the creativity understanding gap: Small-scale human alignment enables expert-level humor ranking in LLMs,” inFindings of the Association for Computational Lingu...

  255. [263]

    The glasgow norms: Ratings of 5,500 words on nine scales,

    G. G. Scott, A. Keitel, M. Becirspahic, B. Yao, and S. C. Sereno, “The glasgow norms: Ratings of 5,500 words on nine scales,”Behavior research methods, vol. 51, no. 3, pp. 1258–1270, 2019

  256. [264]

    How deep is love in llms’ hearts? exploring semantic size in human- like cognition,

    Y . Yao, Y . Yang, X. Ma, D. Yang, Z. Zhang, Z. Li, and H. Zhao, “How deep is love in llms’ hearts? exploring semantic size in human- like cognition,”arXiv preprint arXiv:2503.00330, 2025

  257. [265]

    Semantic-aware logical reasoning via a semiotic framework,

    Y . Zhang, X. Zhang, J. Sheng, W. Li, J. Yu, Y .-P. P. Chen, W. Yang, and Z. Song, “Semantic-aware logical reasoning via a semiotic framework,” Preprint, 2026

  258. [266]

    Semantic complexity in end-to-end spoken language understanding,

    J. P. McKenna, S. Choudhary, M. Saxon, G. P. Strimel, and A. Mouchtaris, “Semantic complexity in end-to-end spoken language understanding,” in21st Annual Conference of the International Speech Communication Association, Interspeech 2020, Virtual Event, Shanghai, China, October...

  259. [267]

    Yes FLoReNce, i will do better next time! agentic feedback reasoning for humorous meme detection,

    O. S. Liu, P. C. Ng, D. W. Soh, and K. N. Plataniotis, “Yes FLoReNce, i will do better next time! agentic feedback reasoning for humorous meme detection,” 2026. [Online]. Available: https://arxiv.org/abs/2601.07232

  260. [268]

    Measuring psychological depth in language models,

    F. Y . Harel-Canada, H. Zhou, S. Muppalla, Z. S. Yildiz, M. Kim, A. Sahai, and N. Peng, “Measuring psychological depth in language models,” inProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, Eds...

  261. [269]

    Modeling empathy and distress in reaction to news stories,

    S. Buechel, A. Buffone, B. Slaff, L. Ungar, and J. Sedoc, “Modeling empathy and distress in reaction to news stories,” inProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2018, pp. 4758–4765

  262. [270]

    Empathy level alignment via reinforcement learning for empathetic response generation,

    H. Ma, B. Zhang, B. Xu, J. Wang, H. Lin, and X. Sun, “Empathy level alignment via reinforcement learning for empathetic response generation,”IEEE Transactions on Affective Computing, pp. 1–12, 2025

  263. [271]

    Empathic conversations: A multi-level dataset of contextualized conversations,

    D. Omitaomu, S. Tafreshi, T. Liu, S. Buechel, C. Callison-Burch, J. Eichstaedt, L. Ungar, and J. Sedoc, “Empathic conversations: A multi-level dataset of contextualized conversations,”arXiv preprint arXiv:2205.12698, 2022

  264. [272]

    Ic9600: A benchmark dataset for automatic image complexity assessment,

    T. Feng, Y . Zhai, J. Yang, J. Liang, D.-P. Fan, J. Zhang, L. Shao, and D. Tao, “Ic9600: A benchmark dataset for automatic image complexity assessment,”IEEE Transactions on Pattern Analysis and Machine Intelligence, no. 01, pp. 1–17, 2023

  265. [273]

    Distribution-based measures of surprise for creative language: Experiments with humor and metaphor,

    R. C. Bunescu and O. O. Uduehi, “Distribution-based measures of surprise for creative language: Experiments with humor and metaphor,” inProceedings of the 3rd Workshop on Figurative Language Processing (FLP), D. Ghosh, B. Beigman Klebanov, S. Muresan, A. Feldman, S. Poria, and...

  266. [274]

    M2p2: Multimodal persuasion prediction using adaptive fusion,

    C. Bai, H. Chen, S. Kumar, J. Leskovec, and V . Subrahmanian, “M2p2: Multimodal persuasion prediction using adaptive fusion,”IEEE Transactions on Multimedia, vol. 25, pp. 942–952, 2021

  267. [275]

    Bridging visual dynamics and narrative reasoning: Multimodal large language models for short drama quality assessment,

    Q. Liu, J. Li, Z. Peng, S. Wang, Z. Liao, S. Chang, B. Gao, H. Zhao, M. Liu, J. Jianget al., “Bridging visual dynamics and narrative reasoning: Multimodal large language models for short drama quality assessment,” inProceedings of the ACM Web Conference 2026, 2026, pp. 7890–7901

  268. [276]

    Dip: Dual incongruity perceiving network for sarcasm detection,

    C. Wen, G. Jia, and J. Yang, “Dip: Dual incongruity perceiving network for sarcasm detection,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 2540–2550

  269. [277]

    LLM-GEm: Large language model-guided prediction of people’s empathy levels towards newspaper article,

    M. R. Hasan, M. Z. Hossain, T. Gedeon, and S. Rahman, “LLM-GEm: Large language model-guided prediction of people’s empathy levels towards newspaper article,” inFindings of the Association for Computational Linguistics: EACL 2024, Y . Graham and M. Purver, Eds. St. Julian’s, Ma...

  270. [278]

    Metaphor detection via explicit basic meanings modelling,

    Y . Li, S. Wang, C. Lin, and F. Guerin, “Metaphor detection via explicit basic meanings modelling,” inProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). Association for Computational Linguistics, 2023, pp. 91–100. ...

  271. [279]

    (RSA)²: A rhetorical-strategy-aware rational speech act framework for figurative language understanding,

    C. Spinoso-Di Piano, D. E. Austin, P. Piantanida, and J. C. Cheung, “(RSA)²: A rhetorical-strategy-aware rational speech act framework for figurative language understanding,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: L...

  272. [280]

    Mitigating idiom inconsistency: A multi-semantic contrastive learning method for chinese idiom reading comprehension,

    M. Wu, Y . Hu, Y . Zhang, Z. Zhi, G. Su, and Y . Sha, “Mitigating idiom inconsistency: A multi-semantic contrastive learning method for chinese idiom reading comprehension,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 17, p. 19243–19251, Mar. 202...

  273. [281]

    Tomap: Training opponent-aware llm persuaders with theory of mind,

    P. Han, Z. Liu, and J. You, “Tomap: Training opponent-aware llm persuaders with theory of mind,”arXiv preprint arXiv:2505.22961, 2025

  274. [282]

    Incongruity-aware tension field network for multi-modal sarcasm detection,

    J. Zhang, C. Chen, S. Li, and T. Zhang, “Incongruity-aware tension field network for multi-modal sarcasm detection,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pil...

  275. [283]

    Commonality and individuality! integrating humor commonality with speaker individuality for humor recognition,

    H. Zhu, J. Lu, Z. Zeng, Z. Bai, X. Zhang, L. Yang, and H. Lin, “Commonality and individuality! integrating humor commonality with speaker individuality for humor recognition,” inProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Co...

  276. [284]

    ImaRA: An imaginative frame augmented method for low-resource multimodal metaphor detection and explanation,

    Y . Tian, M. Wang, N. Xu, and W. Mao, “ImaRA: An imaginative frame augmented method for low-resource multimodal metaphor detection and explanation,” inFindings of the Association for Computational Linguistics: NAACL 2025, L. Chiruzzo, A. Ritter, and L. Wang, Eds. Albuquerque, ...

  277. [285]

    Msme: A multi-stage multi-expert framework for zero-shot stance detection,

    Y . Zhang, A. Li, B. Chen, J. Sun, and X. Zhao, “Msme: A multi-stage multi-expert framework for zero-shot stance detection,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 41, 2026, pp. 34 879–34 887

  278. [286]

    Let androids dream of electric sheep: A human- like image implication understanding and reasoning framework,

    C. Zhang and Y . Niu, “Let androids dream of electric sheep: A human- like image implication understanding and reasoning framework,”arXiv preprint arXiv:2505.17019, 2025

  279. [287]

    Large language models as theory of mind aware generative agents with counterfactual reflection,

    B. Yang, J. Guo, Y . Iwasawa, and Y . Matsuo, “Large language models as theory of mind aware generative agents with counterfactual reflection,”arXiv preprint arXiv:2501.15355, 2025

  280. [288]

    Elevating knowledge-enhanced entity and relationship understanding for sarcasm detection,

    X. Wang, Y . Wang, D. He, Z. Yu, Y . Li, L. Wang, J. Dang, and D. Jin, “Elevating knowledge-enhanced entity and relationship understanding for sarcasm detection,”IEEE Transactions on Knowledge and Data Engineering, 2025

  281. [289]

    Just like a human would, direct access to sarcasm augmented with potential result and reaction,

    C. Min, X. Li, L. Yang, Z. Wang, B. Xu, and H. Lin, “Just like a human would, direct access to sarcasm augmented with potential result and reaction,” inProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), A. Rogers, J...

  282. [290]

    Just KIDDIN’ : Knowledge infusion and distillation for detection of INdecent memes,

    R. Garg, T. Padhi, H. Jain, U. Kursuncu, and P. Kumaraguru, “Just KIDDIN’ : Knowledge infusion and distillation for detection of INdecent memes,” inFindings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds. Vie...

  283. [291]

    Cracking the code: Enhancing implicit hate speech detection through coding classification,

    L. Wei, L. Li, T. Xiang, L. Xiao, and N. Garcia, “Cracking the code: Enhancing implicit hate speech detection through coding classification,” inProceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), T. Cao, A. Das, T. Kumarage, Y . Wan, S. Krishna, N. Mehrabi, J. ...

  284. [292]

    Bridging word-pair and token- level metaphor detection with explainable domain mining,

    Y . Tian, R. Zhang, N. Xu, and W. Mao, “Bridging word-pair and token- level metaphor detection with explainable domain mining,” inProceed- ings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 13 311–13 325

  285. [293]

    Psyche-r1: Towards reliable psychological LLMs through unified empathy, expertise, and reasoning,

    C. Dai, J. Hu, H. Shi, Z. Li, D. Guo, X. Yang, and M. Wang, “Psyche-r1: Towards reliable psychological LLMs through unified empathy, expertise, and reasoning,” inProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M....

  286. [294]

    Learning multitask commonness and uniqueness for multimodal sarcasm detection and sentiment analysis in conversation,

    Y . Zhang, Y . Yu, D. Zhao, Z. Li, B. Wang, Y . Hou, P. Tiwari, and J. Qin, “Learning multitask commonness and uniqueness for multimodal sarcasm detection and sentiment analysis in conversation,” IEEE Transactions on Artificial Intelligence, vol. 5, no. 3, pp. 1349– 1361, 2023

  287. [295]

    Social story frames: Contextual reasoning about narrative intent and reception,

    J. Mire, M. Antoniak, S. R. Wilson, Z. Ma, A. R. Ganti, A. Piper, and M. Sap, “Social story frames: Contextual reasoning about narrative intent and reception,” inProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M....

  288. [297]

    Adversarial multi-task learning for end- to-end metaphor detection,

    S. Zhang and Y . Liu, “Adversarial multi-task learning for end- to-end metaphor detection,” inFindings of the Association for 39 Computational Linguistics: ACL 2023, A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds. Toronto, Canada: Association for Computational Linguistics, Jul...

  289. [1008]

    Available: https://aclanthology.org/2025.acl-long.49/

    [Online]. Available: https://aclanthology.org/2025.acl-long.49/

  290. [2018]

    Available: https://arxiv.org/abs/1810.04805

    [Online]. Available: https://arxiv.org/abs/1810.04805

  291. [2025]

    Available: https://arxiv.org/abs/2502.17857

    [Online]. Available: https://arxiv.org/abs/2502.17857

  292. [2026]

    Available: https://arxiv.org/abs/2506.00955

    [Online]. Available: https://arxiv.org/abs/2506.00955

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.