Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

A Review of Human Emotion Synthesis Based on Generative Technology

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper claims to be the first systematic overview of human emotion synthesis based on generative technology, organizing more than 230 papers into a taxonomy across facial images, speech, and text, and arguing that diffusion models…

desk verdict A useful multi-modal survey whose taxonomy and tables earn it referee time, but the 'first systematic overview' claim and the unauditable selection procedure need fixing before it can be relied on. read the letter →

arxiv 2412.07116 v1 pith:L63UYAFN submitted 2024-12-10 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords emotionsynthesisgenerativemodelsadversarialnetworksdiffusionlargelanguageaffectivecomputingfacialtexttransfer
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper is a systematic review of how generative models produce human emotion in three modalities: facial images, speech, and text. It claims to be the first overview spanning all three modalities under one taxonomy and covering five generative families: autoencoders, GANs, diffusion models, large language models, and sequence-to-sequence models. The review catalogs more than 230 papers, summarizes the main datasets and evaluation metrics, and states a set of field-level findings. The central qualitative claim is that GANs have historically dominated facial emotion synthesis but diffusion models now look like the more promising alternative, while LLMs and Seq2Seq models carry textual emotion synthesis. If this map is right, researchers gain a reliable picture of what has been tried and where the field is moving.

What carries the argument

The central object is the taxonomy in Fig. 2, a three-modality by five-model-family grid that organizes every surveyed paper into a cell. The taxonomy is what makes the review's comparative findings possible: the claim that diffusion models are becoming more promising than GANs for facial emotion synthesis, for example, is read from the distribution of recent diffusion entries in the face reenactment, talking head, and face manipulation rows. The dataset table and the metric tables work as auxiliary machinery that lets a reader check whether a given trend claim rests on common benchmarks and comparable evaluation.

What would settle it

Re-run the stated queries on IEEE Xplore, ScienceDirect, and Google Scholar for 2017-2024 and record a documented exclusion flow; if the resulting pool differs materially from the 230+ papers in Fig. 2 in the per-modality distribution of model families, especially the share of diffusion models in facial emotion synthesis, then the taxonomy and the main trend claim do not hold.

Watch

Extended reading notes

Core claim

On its own terms, the paper's contribution is a claim of scope and structure: it is the first systematic review of human emotion synthesis built on generative technology, and its organizing taxonomy is the right way to see the field. The paper groups facial emotion synthesis into face reenactment, face manipulation, and talking head generation; speech emotion synthesis into voice conversion, text-to-speech, and speech manipulation; and textual emotion synthesis into text emotion transfer and empathetic dialogue generation. Against that grid it arrays roughly 230 papers and reads off trends: diffusion models are emerging as a more promising alternative to GANs for facial emotions; speech synthesis is carried by GAN and Seq2Seq adaptation with AE and DM refinements; and textual emotion synthesis increasingly leans on LLMs and Seq2Seq architectures. The review also asserts that credible evaluation requires both subjective human scoring and objective metrics, and that future progress will come from hybrid model combinations, new modalities, and real-time edge deployment.

Load-bearing premise

The review's completeness and its main trend finding depend on the Section 2 search being unbiased, but the selection steps are not auditable: per-stage exclusion counts are missing, the filter against 'incremental contributions' is subjective, and the stated peer-review criterion contradicts the presence of arXiv preprints.

Editorial extensions

If this is right

  • A newcomer can locate the dominant model family for any emotion-synthesis sub-task through the taxonomy in Fig. 2 rather than searching from scratch.
  • The finding that diffusion models outperform GANs in facial emotion synthesis, if correct, predicts that new facial emotion work will increasingly be diffusion-based and that GAN-only baselines will be compared against DM alternatives.
  • LLM- and Seq2Seq-based text emotion synthesis is mature enough to be a distinct sub-field with its own benchmarks, such as EmpatheticDialogues and the YELP review corpus, which the review's dataset table documents.
  • The review's account of evaluation metrics implies that no single metric captures emotional authenticity, which is why future work should continue combining classifier accuracy, structural or pitch metrics, and human scoring.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A consequence the authors leave implicit: if the taxonomy is accurate, cross-modal emotion synthesis systems can be assembled by slotting a proven DM-based face module and an LLM-based text module under a shared emotion-label space.
  • An audit that reproduces Section 2's database queries would test whether the diffusion-model trend survives a more inclusive search that does not filter out 'incremental' papers.
  • The review's many arXiv preprints despite its stated peer-review criterion suggest that the field's fast-moving results live outside traditional venues; a follow-up review might deliberately track preprints as a separate stratum.
  • A testable extension: assign each paper in Tables 5 to 7 a publication year and model family, then plot the family share over time; the diffusion-upward, GAN-downward trajectory for faces could be quantified rather than asserted.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. This manuscript is a survey of generative-model-based human emotion synthesis across facial images, speech, and text. It claims to be the first systematic overview of this area, analyzes more than 230 papers, and organizes them into a taxonomy (Fig. 2) with separate treatment of face reenactment, face manipulation, talking-head generation, voice conversion, text-to-speech, speech manipulation, text emotion transfer, and empathetic dialogue generation. The paper also presents emotion models, mathematical background for five generative model families (AE, GAN, DM, LLM, Seq2Seq), a dataset summary (Table 4), literature tables with reported performance (Tables 5-7), evaluation metrics, a discussion of major findings, and future research directions.

Significance. If the underlying selection and reporting are reliable, this survey is a useful reference for a growing interdisciplinary area. Its strengths include a broad coverage of roughly 230 works, a clear three-modality taxonomy, dataset tables, a comparison with prior reviews (Table 2), and qualitative field-level findings such as the growing role of diffusion models in facial emotion synthesis. The review contains no new experiments or derivations, so its value rests entirely on the completeness and accuracy of its literature selection and on the faithful transcription of performance numbers. The bookkeeping and methodological transparency problems identified below are therefore load-bearing rather than cosmetic.

major comments (4)
  1. [Section 2 (Review Methodology), Fig. 3] The screening process is not auditable. The text reports only that searches in IEEE Xplore, ScienceDirect, and Google Scholar yielded 'more than 270 papers' and that a 'two-step filtering process' was applied, but it gives no per-stage exclusion counts, query strings, search dates, or protocol for title/abstract/full-text screening. Because the central claim is to be the 'first systematic overview,' the absence of an auditable protocol means the taxonomy in Fig. 2 and the coverage claims cannot be checked. Please provide the full search strings, the number of records at each stage, and a reproducible list of inclusion/exclusion decisions.
  2. [Section 2, inclusion criterion (1)] The stated criterion that only 'Peer-reviewed papers published up to November 2024' are included is contradicted by the reference list, which contains many arXiv preprints without a peer-reviewed venue, including [23], [36], [133], [134], [146], [166], [183], [193], [200], [210], [224], and [225]. The manuscript should either restrict the corpus to peer-reviewed versions where they exist or explicitly revise the inclusion criterion to admit preprints and explain their role; as written, the methodology does not describe the actual corpus.
  3. [Section 2, filtering step] The statement that the authors 'eliminated papers with incremental contributions' is a subjective filter that is not operationalized. If applied without clear rules, this filter can systematically favor well-known model families and thereby shape the field-level finding in Section 11.1 that diffusion models are 'a more promising alternative' for facial emotion synthesis. Please define the exclusion rule (for example, based on citation counts, novelty criteria, or reproducibility status) and, ideally, report a sensitivity analysis of the model-family distribution under alternative inclusion rules.
  4. [Tables 5-7] The performance transcriptions are not sufficiently reliable as presented. Table 5 lists Ding et al. [125] ('ExprGAN') twice; several entries in Tables 5-7 report only 'figures' without numeric values (for example, Table 5 entries for [93], [126], and [127]; Table 6 entries for [142], [143], [148], [150], [153], [162], [173], [176], and [179]; Table 7 entries for [184], [196], [217], and [222]); and some entries lack units or clear metric definitions. Since the review's value depends on faithfully summarizing reported performance, please correct the duplicate and either provide the actual reported numbers with units or clearly mark entries as unavailable.
minor comments (6)
  1. [Section 6] The section heading 'Databases' should be 'Datasets' to match the content and Table 4.
  2. [Fig. 1 caption] In the source list, items (2) and (3) both describe a 'speech emotion synthesis' schematic; one of them should presumably refer to the text emotion synthesis schematic.
  3. [Section 7.2] There is a typo in the sentence about reference [108]: 'Kong te al.' should be 'Kong et al.'
  4. [Section 12] In the first sentence, 'genrative models' should be 'generative models'.
  5. [Abstract] The phrase 'this paper aims to address this gap by providing' is redundant immediately after 'there is a notable lack of comprehensive reviews'; the sentence should be streamlined.
  6. [Table 4] Several rows have inconsistent spacing and formatting (for example, 'ETOD [81] 2019 audio 6000 speeches 13'), which makes the table harder to read; please ensure uniform formatting across all rows.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation chain: this is a literature review with no fitted parameters or predictions, so none of its claims reduce to its own inputs; the only self-referential element is a minor, non-load-bearing citation to the authors' companion scoping review.

full rationale

This paper is a survey, not a derivation. It reports no new experiments, fits no parameters, and makes no quantitative predictions, so the classical circularity patterns (self-definitional equations, fitted inputs renamed as predictions, ansatz smuggled via citation) do not apply. The load-bearing content is the Fig. 2 taxonomy, Table 4 dataset summaries, and Tables 5-7 performance transcriptions, all of which are drawn from the cited primary literature and are checkable against external sources. The 'first systematic overview' novelty claim is a negative scope claim supported by the comparison in Section 4 and Table 2 against prior reviews such as [9] and [52]-[57]; it is not derived from the authors' own prior results. The companion scoping review [12] is cited only for background statements about emotion recognition and the broad impact of generative technology, and removing it would not change any taxonomic or trend conclusion. The paper's genuine weaknesses are selection-bias and reproducibility issues in Section 2: the search covers only IEEE Xplore, ScienceDirect, and Google Scholar; per-stage exclusion counts are not reported; the filter eliminating 'papers with incremental contributions' is subjective; and the stated 'peer-reviewed papers' criterion is contradicted by many arXiv preprints in the reference list. These issues could bias the comprehensiveness and the Section 11.1 qualitative finding that diffusion models are more promising, but bias is not circularity. No equation in the paper is equivalent by construction to its input, and no fitted value is relabeled as a prediction. The circularity score is therefore low.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

A review introduces no free parameters and no invented entities. The load-bearing assumptions are methodological: the completeness claim depends on a three-database search with un-itemized exclusions, and the cross-model trend claims depend on the comparability of transcription-level performance numbers across heterogeneous papers.

assumptions (3)
  • domain assumption The three-modality taxonomy with its sub-task splits (face reenactment, face manipulation, talking head; voice conversion, TTS, speech manipulation; text transfer, empathetic dialogue) is the correct organizing principle for the field.
    Section 1 and Fig. 2 impose this structure on all 230 reviewed papers; if the categories misalign with how the community actually organizes the work, the survey's map misrepresents the literature even if each entry is accurate.
  • domain assumption Discrete emotion models (Ekman basics, Plutchik wheel) and dimensional models (valence-arousal, valence-arousal-dominance) adequately cover the emotion labels used across synthesis tasks.
    Section 3 adopts these as the theoretical foundation, and the datasets in Table 4 label emotions within these schemes; the review inherits any limitations of the label taxonomies of the cited datasets.
  • domain assumption The performance numbers transcribed in Tables 5 to 7 are faithful to the cited papers and comparable across models and datasets.
    The Section 11.1 trend findings, such as diffusion models being more promising for facial emotion synthesis, depend on this comparability; several table entries report "figures" instead of values, and datasets differ across rows, so comparisons are only indicative.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Review of Human Emotion Synthesis Based on Generative Technology." pith.science (2026). https://pith.science/paper/L63UYAFN

@misc{pith2026241207116,
  author       = {Pith},
  title        = {Pith review of: A Review of Human Emotion Synthesis Based on Generative Technology},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/L63UYAFN}},
  note         = {Machine review of arXiv:2412.07116}
}
read the original abstract

Human emotion synthesis is a crucial aspect of affective computing. It involves using computational methods to mimic and convey human emotions through various modalities, with the goal of enabling more natural and effective human-computer interactions. Recent advancements in generative models, such as Autoencoders, Generative Adversarial Networks, Diffusion Models, Large Language Models, and Sequence-to-Sequence Models, have significantly contributed to the development of this field. However, there is a notable lack of comprehensive reviews in this field. To address this problem, this paper aims to address this gap by providing a thorough and systematic overview of recent advancements in human emotion synthesis based on generative models. Specifically, this review will first present the review methodology, the emotion models involved, the mathematical principles of generative models, and the datasets used. Then, the review covers the application of different generative models to emotion synthesis based on a variety of modalities, including facial images, speech, and text. It also examines mainstream evaluation metrics. Additionally, the review presents some major findings and suggests future research directions, providing a comprehensive understanding of the role of generative technology in the nuanced domain of emotion synthesis.

Figures

Figures reproduced from arXiv: 2412.07116 by the authors.

Figure 1
Figure 1. Schematic Diagram of Generation Technology for Human Emo [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Taxonomy of This Survey. approach was followed to ensure the comprehensiveness and relevance of the literature. The overall screening pro￾cess is shown in [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 4
Figure 4. Plutchik Wheel (left) and 2D Emotion Model (right). [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗
Figures from the paper (7 more)
Figure 3
Figure 3. Figure 3: A Comprehensive Review Methodology. up to November 2024, including both journal and confer￾ence papers; (2) Research focusing on the application of generative models to human emotion synthesis in various modalities; (3) Studies in which extensive experiments were condu…
Figure 5
Figure 5. Figure 5: A mask-based GAN [95] for face reenactment. The system [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: A DM-based talking head generation model from [23]. This archi [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: A two-stage model [145] for voice conversion. In the first stage, [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 8
Figure 8. Figure 8: A Tacotron2-based TTS model from [164]. This system integrated [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: A text emotion transfer model from [183]. During training, the [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: A empathetic dialogue generation system from [198]. The [PITH_FULL_IMAGE:figures/full_fig_p015_10.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Scalable Event Cloud Network for Event-based Classification

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A frequency-aware network operating on raw-like event clouds matches or beats prior event-based models on nine benchmarks while using roughly 0.1 G MACs, far below frame and voxel baselines.

Reference graph

Works this paper leans on

239 extracted references · 53 canonical work pages · cited by 1 Pith paper

  1. [23]

    EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model

    B. Zhang, X. Zhang, N. Cheng, J. Yu, J. Xiao, and J. Wang, “Emotalker: Emotionally editable talking face generation via diffusion model,” arXiv preprint arXiv:2401.08049, 2024

  2. [146]

    Durflex-evc: Duration-flexible emotional voice conversion with parallel gen- eration,

    H.-S. Oh, S.-H. Lee, D.-H. Cho, and S.-W. Lee, “Durflex-evc: Duration-flexible emotional voice conversion with parallel gen- eration,” arXiv preprint arXiv:2401.08095, 2024. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 23

  3. [166]

    Emomix: Emotion mixing via diffusion models for emotional speech syn- thesis,

    H. Tang, X. Zhang, J. Wang, N. Cheng, and J. Xiao, “Emomix: Emotion mixing via diffusion models for emotional speech syn- thesis,” arXiv preprint arXiv:2306.00648, 2023

  4. [210]

    DiffusEmp: A Diffusion Model-Based Framework with Multi-Grained Control for Empathetic Response Generation

    G. Bi, L. Shen, Y. Cao, M. Chen, Y. Xie, Z. Lin, and X. He, “Diffusemp: A diffusion model-based framework with multi- grained control for empathetic response generation,” arXiv preprint arXiv:2306.01657, 2023

  5. [225]

    Cause-aware empathetic response generation via chain-of-thought fine-tuning,

    X. Chen, C. Yang, M. Lan et al. , “Cause-aware empathetic response generation via chain-of-thought fine-tuning,” arXiv preprint arXiv:2408.11599, 2024

  6. [9]

    Generative adversarial networks in human emotion synthesis: A review,

    N. Hajarolasvadi, M. A. Ramirez, W. Beccaro, and H. Demirel, “Generative adversarial networks in human emotion synthesis: A review,” IEEE Access, vol. 8, pp. 218 499–218 529, 2020

  7. [36]

    Fine-grained quan- titative emotion editing for speech generation,

    S. Inoue, K. Zhou, S. Wang, and H. Li, “Fine-grained quan- titative emotion editing for speech generation,” arXiv preprint arXiv:2403.02002, 2024

  8. [133]

    Dreamtalk: When expressive talking head generation meets diffusion probabilistic models,

    Y. Ma, S. Zhang, J. Wang, X. Wang, Y. Zhang, and Z. Deng, “Dreamtalk: When expressive talking head generation meets diffusion probabilistic models,” arXiv preprint arXiv:2312.09767 , 2023

  9. [134]

    Dream-talk: Diffusion-based realis- tic emotional audio-driven method for single image talking face generation,

    C. Zhang, C. Wang, J. Zhang, H. Xu, G. Song, Y. Xie, L. Luo, Y. Tian, X. Guo, and J. Feng, “Dream-talk: Diffusion-based realis- tic emotional audio-driven method for single image talking face generation,” arXiv preprint arXiv:2312.13578, 2023

  10. [183]

    Textsettr: Few-shot text style extraction and tunable targeted restyling,

    P . Riley, N. Constant, M. Guo, G. Kumar, D. C. Uthus, and Z. Parekh, “Textsettr: Few-shot text style extraction and tunable targeted restyling,” in Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), 2021, pp. 3786–3800

  11. [193]

    Enhancing the emotional generation capability of large language models via emotional chain-of-thought,

    Z. Li, G. Chen, R. Shao, D. Jiang, and L. Nie, “Enhancing the emotional generation capability of large language models via emotional chain-of-thought,” arXiv preprint arXiv:2401.06836, 2024

  12. [200]

    An Adversarial Approach to High-Quality, Sentiment-Controlled Neural Dialogue Generation

    X. Kong, B. Li, G. Neubig, E. Hovy, and Y. Yang, “An adversarial approach to high-quality, sentiment-controlled neural dialogue generation,” arXiv preprint arXiv:1901.07129, 2019

  13. [224]

    Multi-level knowledge-enhanced prompting for empathetic dialogue generation,

    Z. Gu, Q. Zhu, H. He et al. , “Multi-level knowledge-enhanced prompting for empathetic dialogue generation,” in Proceedings of the 27th International Conference on Computer Supported Cooperative Work in Design (CSCWD). IEEE, 2024, pp. 3170–3175

  14. [125]

    Exprgan: Facial expres- sion editing with controllable expression intensity,

    H. Ding, K. Sricharan, and R. Chellappa, “Exprgan: Facial expres- sion editing with controllable expression intensity,” inProceedings of the AAAI Conference on Artificial Intelligence , vol. 32, 2018

  15. [93]

    Altering the conveyed facial emotion through automatic reenactment of video portraits,

    C. Groth, J.-P . Tauscher, S. Castillo, and M. Magnor, “Altering the conveyed facial emotion through automatic reenactment of video portraits,” in International Conference on Computer Animation and Social Agents. Springer, 2020, pp. 128–135

  16. [126]

    Generating realistic facial expressions through condi- tional cycle-consistent generative adversarial networks (ccycle- gan),

    G. Tesei, “Generating realistic facial expressions through condi- tional cycle-consistent generative adversarial networks (ccycle- gan),” 2019

  17. [127]

    Comp-gan: Compositional generative adversarial network in synthesizing and recognizing facial expression,

    W. Wang, Q. Sun, Y. Fu, T. Chen, C. Cao, Z. Zheng, G. Xu, H. Qiu, Y.-G. Jiang, and X. Xue, “Comp-gan: Compositional generative adversarial network in synthesizing and recognizing facial expression,” in Proceedings of the 27th ACM International Conference on Multimedia, 2019, pp. 211–219

  18. [142]

    Stargan for emotional speech conversion: Validated by data augmentation of end-to-end emotion recognition,

    G. Rizos, A. Baird, M. Elliott, and B. Schuller, “Stargan for emotional speech conversion: Validated by data augmentation of end-to-end emotion recognition,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 3502–3506

  19. [143]

    Multi-speaker and multi- domain emotional voice conversion using factorized hierarchical variational autoencoder,

    M. Elgaar, J. Park, and S. W. Lee, “Multi-speaker and multi- domain emotional voice conversion using factorized hierarchical variational autoencoder,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 7769–7773

  20. [148]

    English emotional voice conversion using stargan model,

    A. H. Meftah, A. A. Alashban, Y. A. Alotaibi, and S.-A. Selouani, “English emotional voice conversion using stargan model,” IEEE Access, 2023

  21. [150]

    Disentanglement of emo- tional style and speaker identity for expressive voice conversion,

    Z. Du, B. Sisman, K. Zhou, and H. Li, “Disentanglement of emo- tional style and speaker identity for expressive voice conversion,” arXiv preprint arXiv:2110.10326, 2021

  22. [153]

    Sequence-to-sequence emotional voice conversion with strength control,

    H. Choi and M. Hahn, “Sequence-to-sequence emotional voice conversion with strength control,” IEEE Access, vol. 9, pp. 42 674– 42 687, 2021

  23. [162]

    Controllable emotion transfer for end-to-end speech synthesis,

    T. Li, S. Yang, L. Xue, and L. Xie, “Controllable emotion transfer for end-to-end speech synthesis,” in 2021 12th International Sym- posium on Chinese Spoken Language Processing (ISCSLP) . IEEE, 2021, pp. 1–5

  24. [173]

    Cross-speaker emotion transfer through information perturbation in emotional speech synthesis,

    Y. Lei, S. Yang, X. Zhu, L. Xie, and D. Su, “Cross-speaker emotion transfer through information perturbation in emotional speech synthesis,” IEEE Signal Processing Letters , vol. 29, pp. 1948–1952, 2022

  25. [176]

    Speech synthesis with mixed emotions,

    K. Zhou, B. Sisman, R. Rana, B. W. Schuller, and H. Li, “Speech synthesis with mixed emotions,” IEEE Transactions on Affective Computing, 2022

  26. [179]

    Emotional dimension control in language model-based text-to-speech: Spanning a broad spec- trum of human emotions,

    K. Zhou, Y. Zhang, S. Zhao et al., “Emotional dimension control in language model-based text-to-speech: Spanning a broad spec- trum of human emotions,” arXiv preprint arXiv:2409.16681, 2024

  27. [184]

    Tet: Text emotion transfer,

    R. MohammadiBaghmolaei and A. Ahmadi, “Tet: Text emotion transfer,” Knowledge-Based Systems, vol. 262, p. 110236, 2023

  28. [196]

    Does gpt-3 generate em- pathetic dialogues? a novel in-context example selection method and automatic evaluation metric for empathetic dialogue gen- eration,

    Y.-J. Lee, C.-G. Lim, and H.-J. Choi, “Does gpt-3 generate em- pathetic dialogues? a novel in-context example selection method and automatic evaluation metric for empathetic dialogue gen- eration,” in Proceedings of the 29th International Conference on Computational Linguistics, 2022, pp. 669–683

  29. [217]

    Empathetic dialogue generation with pre-trained roberta-gpt2 and external knowl- edge,

    Y. Liu, W. Maier, W. Minker, and S. Ultes, “Empathetic dialogue generation with pre-trained roberta-gpt2 and external knowl- edge,” in Conversational AI for Natural Human-Centric Interaction: 12th International Workshop on Spoken Dialogue System Technology, IWSDS 2021, Singapore. Springer, 2022, pp. 67–81

  30. [222]

    Reinforcement learn- ing based emotional editing constraint conversation generation,

    J. Li, X. Sun, X. Wei, C. Li, and J. Tao, “Reinforcement learn- ing based emotional editing constraint conversation generation,” arXiv preprint arXiv:1904.08061, 2019

Show all 239 references
  1. [1]

    R. W. Picard, Affective computing. MIT press, 2000

  2. [2]

    Neuro or symbolic? fine-tuned transformer with unsupervised lda topic clustering for text senti- ment analysis,

    F. Ding, X. Kang, and F. Ren, “Neuro or symbolic? fine-tuned transformer with unsupervised lda topic clustering for text senti- ment analysis,” IEEE Transactions on Affective Computing , vol. 15, no. 2, pp. 493–507, 2024

  3. [3]

    Affective computing in education: A systematic review and future research,

    E. Yadegaridehkordi, N. F. B. M. Noor, M. N. B. Ayub, H. B. Affal, and N. B. Hussin, “Affective computing in education: A systematic review and future research,” Computers & education , vol. 142, p. 103649, 2019

  4. [4]

    An efficient framework for con- structing speech emotion corpus based on integrated active learn- ing strategies,

    F. Ren, Z. Liu, and X. Kang, “An efficient framework for con- structing speech emotion corpus based on integrated active learn- ing strategies,” IEEE Transactions on Affective Computing , vol. 13, no. 4, pp. 1929–1940, 2022

  5. [5]

    A survey of textual emotion recognition and its challenges,

    J. Deng and F. Ren, “A survey of textual emotion recognition and its challenges,” IEEE Transactions on Affective Computing , vol. 14, no. 1, pp. 49–67, 2023

  6. [6]

    Prompt consistency for multi- label textual emotion detection,

    Y. Zhou, X. Kang, and F. Ren, “Prompt consistency for multi- label textual emotion detection,” IEEE Transactions on Affective Computing, vol. 15, no. 1, pp. 121–129, 2024

  7. [7]

    Emotional speech synthesis: A review,

    M. Schröder, “Emotional speech synthesis: A review,” in Seventh European Conference on Speech Communication and Technology, 2001

  8. [8]

    Speech synthesis based on hidden markov models,

    K. Tokuda, Y. Nankaku, T. Toda, H. Zen, J. Yamagishi, and K. Oura, “Speech synthesis based on hidden markov models,” Proceedings of the IEEE, vol. 101, no. 5, pp. 1234–1252, 2013

  9. [10]

    A review of affective computing: From unimodal analysis to multimodal fusion,

    S. Poria, E. Cambria, R. Bajpai, and A. Hussain, “A review of affective computing: From unimodal analysis to multimodal fusion,” Information fusion, vol. 37, pp. 98–125, 2017

  10. [11]

    A compre- hensive survey of ai-generated content (aigc): A history of gen- erative ai from gan to chatgpt,

    Y. Cao, S. Li, Y. Liu, Z. Yan, Y. Dai, P . S. Yu, and L. Sun, “A compre- hensive survey of ai-generated content (aigc): A history of gen- erative ai from gan to chatgpt,” arXiv preprint arXiv:2303.04226 , 2023

  11. [12]

    Generative technology for human emotion recognition: A scoping review,

    F. Ma, Y. Yuan, Y. Xie, H. Ren, I. Liu, Y. He, F. Ren, F. R. Yu, and S. Ni, “Generative technology for human emotion recognition: A scoping review,” Information Fusion, p. 102753, 2024

  12. [13]

    An introduction to deep generative modeling,

    L. Ruthotto and E. Haber, “An introduction to deep generative modeling,” GAMM-Mitteilungen, vol. 44, no. 2, p. e202100008, 2021

  13. [14]

    Overview of the transformer-based models for nlp tasks,

    A. Gillioz, J. Casas, E. Mugellini, and O. A. Khaled, “Overview of the transformer-based models for nlp tasks,” in 2020 15th Conference on Computer Science and Information Systems (FedCSIS) . IEEE, 2020, pp. 179–183

  14. [15]

    Genera- tive ai,

    S. Feuerriegel, J. Hartmann, C. Janiesch, and P . Zschech, “Genera- tive ai,” Business & Information Systems Engineering, vol. 66, no. 1, pp. 111–126, 2024

  15. [16]

    Ccis-diff: A generative model with stable diffusion prior for controlled colonoscopy image synthesis,

    Y. Xie, J. Wang, T. Feng, F. Ma, and Y. Li, “Ccis-diff: A generative model with stable diffusion prior for controlled colonoscopy image synthesis,” arXiv preprint arXiv:2411.12198, 2024

  16. [17]

    Reducing the dimension- ality of data with neural networks,

    G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimension- ality of data with neural networks,” Science, vol. 313, no. 5786, pp. 504–507, 2006. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 20

  17. [18]

    iemotts: Toward robust cross-speaker emotion transfer and control for speech synthesis based on disentanglement be- tween prosody and timbre,

    G. Zhang, Y. Qin, W. Zhang, J. Wu, M. Li, Y. Gai, F. Jiang, and T. Lee, “iemotts: Toward robust cross-speaker emotion transfer and control for speech synthesis based on disentanglement be- tween prosody and timbre,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2023

  18. [19]

    Generative adver- sarial nets,

    I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde- Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adver- sarial nets,” in Advances in neural information processing systems , vol. 27, 2014

  19. [20]

    Ganimation: Anatomically-aware facial an- imation from a single image,

    A. Pumarola, A. Agudo, A. M. Martinez, A. Sanfeliu, and F. Moreno-Noguer, “Ganimation: Anatomically-aware facial an- imation from a single image,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 818–833

  20. [21]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P . Abbeel, “Denoising diffusion probabilistic models,” in Advances in neural information processing systems , vol. 33, 2020, pp. 6840–6851

  21. [22]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in neural information processing systems , vol. 30, 2017

  22. [24]

    Enhancing conversational agents with em- pathic abilities,

    J. Casas, T. Spring, K. Daher, E. Mugellini, O. A. Khaled, and P . Cudré-Mauroux, “Enhancing conversational agents with em- pathic abilities,” in Proceedings of the 21st ACM International Conference on Intelligent Virtual Agents, 2021, pp. 41–47

  23. [25]

    Sequence to sequence learning with neural networks,

    I. Sutskever, O. Vinyals, and Q. V . Le, “Sequence to sequence learning with neural networks,” in Advances in neural information processing systems, vol. 27, 2014

  24. [26]

    Delete, retrieve, generate: A simple approach to sentiment and style transfer,

    J. Li, R. Jia, H. He, and P . Liang, “Delete, retrieve, generate: A simple approach to sentiment and style transfer,” inProceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Lon...

  25. [27]

    Towards empathetic open-domain conversation models: A new bench- mark and dataset,

    H. Rashkin, E. M. Smith, M. Li, and Y.-L. Boureau, “Towards empathetic open-domain conversation models: A new bench- mark and dataset,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019, pp. 5370–5381

  26. [28]

    A review on face reenactment techniques,

    S. Dhere, S. B. Rathod, S. Aarankalle, Y. Lad, and M. Gandhi, “A review on face reenactment techniques,” in 2020 International Conference on Industry 4.0 Technology (I4Tech) . IEEE, 2020, pp. 191–194

  27. [29]

    Codeswap: Symmetrically face swapping based on prior codebook,

    X. Luo, X. Zhang, Y. Xie, X. Tong, W. Yu, H. Chang, F. Ma, and F. R. Yu, “Codeswap: Symmetrically face swapping based on prior codebook,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 6910–6919

  28. [30]

    Consistency preservation and feature entropy regularization for gan-based face editing,

    W. Xie, W. Lu, Z. Peng, and L. Shen, “Consistency preservation and feature entropy regularization for gan-based face editing,” IEEE Transactions on Multimedia, 2023

  29. [31]

    Human- computer interaction system: A survey of talking-head genera- tion,

    R. Zhen, W. Song, Q. He, J. Cao, L. Shi, and J. Luo, “Human- computer interaction system: A survey of talking-head genera- tion,” Electronics, vol. 12, no. 1, p. 218, 2023

  30. [32]

    Talking human face gen- eration: A survey,

    M. Toshpulatov, W. Lee, and S. Lee, “Talking human face gen- eration: A survey,” Expert Systems with Applications , vol. 219, p. 119678, 2023

  31. [33]

    An overview of voice conversion and its challenges: From statistical modeling to deep learning,

    B. Sisman, J. Yamagishi, S. King, and H. Li, “An overview of voice conversion and its challenges: From statistical modeling to deep learning,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 132–157, 2020

  32. [34]

    Emotional voice conver- sion: Theory, databases and esd,

    K. Zhou, B. Sisman, R. Liu, and H. Li, “Emotional voice conver- sion: Theory, databases and esd,” Speech Communication, vol. 137, pp. 1–18, 2022

  33. [35]

    Conventional and contemporary ap- proaches used in text to speech synthesis: a review,

    N. Kaur and P . Singh, “Conventional and contemporary ap- proaches used in text to speech synthesis: a review,” Artificial Intelligence Review, vol. 56, pp. 5837–5880, 2023

  34. [37]

    Investigating critical frequency bands and channels for eeg-based emotion recognition with deep neural networks,

    W.-L. Zheng and B.-L. Lu, “Investigating critical frequency bands and channels for eeg-based emotion recognition with deep neural networks,” IEEE Transactions on Autonomous Mental Development , vol. 7, no. 3, pp. 162–175, 2015

  35. [38]

    On the theory of the electrocardiogram,

    D. B. Geselowitz, “On the theory of the electrocardiogram,” Proceedings of the IEEE, vol. 77, no. 6, pp. 857–876, 1989

  36. [39]

    Emotion in human-computer interaction,

    S. Brave and C. Nass, “Emotion in human-computer interaction,” in The human-computer interaction handbook. CRC Press, 2007, pp. 103–118

  37. [40]

    Functional accounts of emotions,

    D. Keltner and J. J. Gross, “Functional accounts of emotions,” Cognition & Emotion, vol. 13, no. 5, pp. 467–480, 1999

  38. [41]

    Measures of emotion: A reviews,

    I. B. Mauss and M. D. Robinson, “Measures of emotion: A reviews,” Cognition and emotion, pp. 109–137, 2010

  39. [42]

    Emotion and decision making,

    J. S. Lerner, Y. Li, P . Valdesolo, and K. S. Kassam, “Emotion and decision making,” Annual review of psychology , vol. 66, pp. 799– 823, 2015

  40. [43]

    Parkinson, A

    B. Parkinson, A. H. Fischer, and A. S. Manstead, Emotion in social relations: Cultural, group, and interpersonal processes . Psychology press, 2005

  41. [44]

    Emotion recognition and artificial intelligence: A systematic review (2014–2023) and research recommendations,

    S. K. Khare, V . Blanes-Vidal, E. S. Nadimi, and U. R. Acharya, “Emotion recognition and artificial intelligence: A systematic review (2014–2023) and research recommendations,” Information Fusion, vol. 102019, 2023

  42. [45]

    Personality-aware person- alized emotion recognition from physiological signals,

    S. Zhao, G. Ding, J. Han, and Y. Gao, “Personality-aware person- alized emotion recognition from physiological signals,” in IJCAI, 2018, pp. 1660–1667

  43. [46]

    Personalized emotion recognition by personality- aware high-order learning of physiological signals,

    S. Zhao, A. Gholaminejad, G. Ding, Y. Gao, J. Han, and K. Keutzer, “Personalized emotion recognition by personality- aware high-order learning of physiological signals,” ACM Trans- actions on Multimedia Computing, Communications, and Applications (TOMM), vol. 15, no. 1s, pp. 1...

  44. [47]

    An argument for basic emotions,

    P . Ekman, “An argument for basic emotions,” Cognition & emo- tion, vol. 6, no. 3-4, pp. 169–200, 1992

  45. [48]

    Plutchik and H

    R. Plutchik and H. Kellerman, Theories of emotion . Academic press, 2013, vol. 1

  46. [49]

    Data aug- mentation for audio-visual emotion recognition with an efficient multimodal conditional gan,

    F. Ma, Y. Li, S. Ni, S.-L. Huang, and L. Zhang, “Data aug- mentation for audio-visual emotion recognition with an efficient multimodal conditional gan,” Applied Sciences , vol. 12, no. 1, p. 527, 2022

  47. [50]

    Real-time assessment of men- tal workload using psychophysiological measures and artificial neural networks,

    G. F. Wilson and C. A. Russell, “Real-time assessment of men- tal workload using psychophysiological measures and artificial neural networks,” Human factors, vol. 45, no. 4, pp. 635–644, 2003

  48. [51]

    Pleasure-arousal-dominance: A general frame- work for describing and measuring individual differences in temperament,

    A. Mehrabian, “Pleasure-arousal-dominance: A general frame- work for describing and measuring individual differences in temperament,” Current Psychology, vol. 14, pp. 261–292, 1996

  49. [52]

    Gener- ative adversarial networks for face generation: A survey,

    A. Kammoun, R. Slama, H. Tabia, T. Ouni, and M. Abid, “Gener- ative adversarial networks for face generation: A survey,” ACM Computing Surveys, vol. 55, no. 5, pp. 1–37, 2022

  50. [53]

    Gan-based facial attribute manipulation,

    Y. Liu, Q. Li, Q. Deng, Z. Sun, and M.-H. Yang, “Gan-based facial attribute manipulation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023

  51. [54]

    An overview of affective speech synthesis and conversion in the deep learning era,

    A. Triantafyllopoulos, B. W. Schuller, G. ˙Iymen, M. Sezgin, X. He, Z. Yang, P . Tzirakis, S. Liu, S. Mertes, E. André et al. , “An overview of affective speech synthesis and conversion in the deep learning era,” Proceedings of the IEEE, 2023

  52. [55]

    Generative adversarial networks for speech processing: A review,

    A. Wali, Z. Alamgir, S. Karim, A. Fawaz, M. B. Ali, M. Adan, and M. Mujtaba, “Generative adversarial networks for speech processing: A review,” Computer Speech & Language , vol. 72, p. 101308, 2022

  53. [56]

    A survey of controllable text generation using transformer-based pre-trained language models,

    H. Zhang, H. Song, S. Li, M. Zhou, and D. Song, “A survey of controllable text generation using transformer-based pre-trained language models,” ACM Computing Surveys , vol. 56, no. 3, pp. 1–37, 2023

  54. [57]

    A survey on text generation using generative adversarial networks,

    G. H. D. Rosa and J. P . Papa, “A survey on text generation using generative adversarial networks,” Pattern Recognition, vol. 119, p. 108098, 2021

  55. [58]

    Emotion recognition and detection methods: A comprehensive survey,

    A. Saxena, A. Khanna, and D. Gupta, “Emotion recognition and detection methods: A comprehensive survey,” Journal of Artificial Intelligence and Systems, vol. 2, no. 1, pp. 53–79, 2020

  56. [59]

    A review on sentiment analysis and emotion detection from text,

    P . Nandwani and R. Verma, “A review on sentiment analysis and emotion detection from text,” Social network analysis and mining , vol. 11, no. 1, p. 81, 2021

  57. [60]

    A survey on facial emotion recognition techniques: A state-of-the-art literature review,

    F. Z. Canal, T. R. Müller, J. C. Matias, G. G. Scotton, A. R. de Sa Junior, E. Pozzebon, and A. C. Sobieranski, “A survey on facial emotion recognition techniques: A state-of-the-art literature review,” Information Sciences, vol. 582, pp. 593–617, 2022

  58. [61]

    Deep generative models: Sur- vey,

    A. Oussidi and A. Elhassouny, “Deep generative models: Sur- vey,” in 2018 International conference on intelligent systems and computer vision (ISCV). IEEE, 2018, pp. 1–8

  59. [62]

    Auto-encoding variational bayes,

    D. P . Kingma and M. Welling, “Auto-encoding variational bayes,” arXiv preprint arXiv:1312.6114, 2013

  60. [63]

    Improving language understanding by generative pre-training,

    A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. , “Improving language understanding by generative pre-training,” 2018. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 21

  61. [64]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. D. M.-W. C. Kenton and L. K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of NAACL-HLT, 2019, pp. 4171–4186

  62. [65]

    Xlnet: Generalized autoregressive pretraining for lan- guage understanding,

    Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V . Le, “Xlnet: Generalized autoregressive pretraining for lan- guage understanding,” in Advances in neural information processing systems, vol. 32, 2019

  63. [66]

    Exploring the limits of transfer learning with a unified text-to-text transformer,

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P . J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of ma- chine learning research, vol. 21, no. 140, pp. 1–67, 2020

  64. [67]

    Language models are unsupervised multitask learners,

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog, vol. 1, no. 8, p. 9, 2019

  65. [68]

    Mllm-ta: Leveraging multimodal large language models for precise temporal video grounding,

    Y. Liu, H. Hou, F. Ma, S. Ni, and F. R. Yu, “Mllm-ta: Leveraging multimodal large language models for precise temporal video grounding,” IEEE Signal Processing Letters, 2024

  66. [69]

    Facial expression recognition from near-infrared videos,

    G. Zhao, X. Huang, M. Taini, S. Z. Li, and M. PietikäInen, “Facial expression recognition from near-infrared videos,” Image and vision computing, vol. 29, no. 9, pp. 607–619, 2011

  67. [70]

    Presentation and validation of the radboud faces database,

    O. Langner, R. Dotsch, G. Bijlstra, D. H. Wigboldus, S. T. Hawk, and A. V . Knippenberg, “Presentation and validation of the radboud faces database,” Cognition and emotion , vol. 24, no. 8, pp. 1377–1388, 2010

  68. [71]

    The extended cohn-kanade dataset (ck+): A com- plete dataset for action unit and emotion-specified expression,

    P . Lucey, J. F. Cohn, T. Kanade, J. Saragih, Z. Ambadar, and I. Matthews, “The extended cohn-kanade dataset (ck+): A com- plete dataset for action unit and emotion-specified expression,” in 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Works...

  69. [72]

    Compound facial expressions of emotion,

    S. Du, Y. Tao, and A. M. Martinez, “Compound facial expressions of emotion,” Proceedings of the National Academy of Sciences , vol. 111, no. 15, pp. E1454–E1462, 2014

  70. [73]

    Affectnet: A database for facial expression, valence, and arousal computing in the wild,

    A. Mollahosseini, B. Hasani, and M. H. Mahoor, “Affectnet: A database for facial expression, valence, and arousal computing in the wild,” IEEE Transactions on Affective Computing, vol. 10, no. 1, pp. 18–31, 2017

  71. [74]

    Disfa: A spontaneous facial action intensity database,

    S. M. Mavadati, M. H. Mahoor, K. Bartlett, P . Trinh, and J. F. Cohn, “Disfa: A spontaneous facial action intensity database,” IEEE Transactions on Affective Computing , vol. 4, no. 2, pp. 151– 160, 2013

  72. [75]

    Emo- tionet: An accurate, real-time algorithm for the automatic anno- tation of a million facial expressions in the wild,

    C. F. Benitez-Quiroz, R. Srinivasan, and A. M. Martinez, “Emo- tionet: An accurate, real-time algorithm for the automatic anno- tation of a million facial expressions in the wild,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 5562–5570

  73. [76]

    Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,

    K. Zhou, B. Sisman, R. Liu, and H. Li, “Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,” in ICASSP 2021 - 2021 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 920–924

  74. [77]

    The emotional voices database: Towards controlling the emotion dimension in voice generation systems,

    A. Adigwe, N. Tits, K. E. Haddad, S. Ostadabbas, and T. Du- toit, “The emotional voices database: Towards controlling the emotion dimension in voice generation systems,” arXiv preprint arXiv:1806.09514, 2018

  75. [78]

    A database of german emotional speech,

    F. Burkhardt, A. Paeschke, M. Rolfes, W. F. Sendlmeier, and B. Weiss, “A database of german emotional speech,” in Inter- speech, vol. 5, 2005, pp. 1517–1520

  76. [79]

    Et-gan: Cross-language emotion transfer based on cycle-consistent generative adversarial networks,

    X. Jia, J. Tai, H. Zhou, Y. Li, W. Zhang, H. Du, and Q. Huang, “Et-gan: Cross-language emotion transfer based on cycle-consistent generative adversarial networks,” arXiv preprint arXiv:1905.11173, 2019

  77. [80]

    A canadian french emo- tional speech dataset,

    P . Gournay, O. Lahaie, and R. Lefebvre, “A canadian french emo- tional speech dataset,” in Proceedings of the 9th ACM Multimedia Systems Conference, 2018, pp. 399–402

  78. [81]

    Emoq-tts: Emotion intensity quantization for fine-grained controllable emotional text-to-speech,

    C.-B. Im, S.-H. Lee, S.-B. Kim, and S.-W. Lee, “Emoq-tts: Emotion intensity quantization for fine-grained controllable emotional text-to-speech,” in ICASSP 2022 - 2022 IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 6317–6321

  79. [82]

    Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering,

    R. He and J. McAuley, “Ups and downs: Modeling the visual evolution of fashion trends with one-class collaborative filtering,” in Proceedings of the 25th International Conference on World Wide Web, 2016, pp. 507–517

  80. [83]

    Mojitalk: Generating emotional re- sponses at scale,

    X. Zhou and W. Y. Wang, “Mojitalk: Generating emotional re- sponses at scale,” in Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2018, pp. 1128–1137

  81. [84]

    Mead: A large-scale audio-visual dataset for emotional talking-face generation,

    K. Wang, Q. Wu, L. Song, Z. Yang, W. Wu, C. Qian, R. He, Y. Qiao, and C. C. Loy, “Mead: A large-scale audio-visual dataset for emotional talking-face generation,” in European Conference on Computer Vision. Springer, 2020, pp. 700–717

  82. [85]

    Crema-d: Crowd-sourced emotional multimodal actors dataset,

    H. Cao, D. G. Cooper, M. K. Keutmann, R. C. Gur, A. Nenkova, and R. Verma, “Crema-d: Crowd-sourced emotional multimodal actors dataset,” IEEE Transactions on Affective Computing , vol. 5, no. 4, pp. 377–390, 2014

  83. [86]

    Emotion recognition in speech using cross-modal transfer in the wild,

    S. Albanie, A. Nagrani, A. Vedaldi, and A. Zisserman, “Emotion recognition in speech using cross-modal transfer in the wild,” in Proceedings of the 26th ACM International Conference on Multimedia, 2018, pp. 292–301

  84. [87]

    Audio-visual feature se- lection and reduction for emotion classification,

    S. Haq, P . J. Jackson, and J. Edge, “Audio-visual feature se- lection and reduction for emotion classification,” in Proc. Int. Conf. on Auditory-Visual Speech Processing (AVSP’08) , 2008, pp. Tangalooma, Australia

  85. [88]

    The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,

    S. R. Livingstone and F. A. Russo, “The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,” PloS One, vol. 13, no. 5, p. e0196391, 2018

  86. [89]

    Iemocap: Interactive emotional dyadic motion capture database,

    C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, E. Mower, S. Kim, J. N. Chang, S. Lee, and S. S. Narayanan, “Iemocap: Interactive emotional dyadic motion capture database,” Language Resources and Evaluation, vol. 42, pp. 335–359, 2008

  87. [90]

    Icface: Interpretable and controllable face reenactment using gans,

    S. Tripathy, J. Kannala, and E. Rahtu, “Icface: Interpretable and controllable face reenactment using gans,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2020, pp. 3385–3394

  88. [91]

    Realistic face reenactment via self-supervised disentangling of identity and pose,

    X. Zeng, Y. Pan, M. Wang, J. Zhang, and Y. Liu, “Realistic face reenactment via self-supervised disentangling of identity and pose,” in Proceedings of the AAAI Conference on Artificial Intelli- gence, vol. 34, 2020, pp. 12 757–12 764

  89. [92]

    Emotion editing in head reenactment videos using latent space manipulation,

    V . Strizhkova, Y. Wang, D. Anghelone, D. Yang, A. Dantcheva, and F. Brémond, “Emotion editing in head reenactment videos using latent space manipulation,” in 2021 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021) . IEEE, 2021, pp. 1–8

  90. [94]

    All-in-one: Facial expression transfer, editing and recognition using a single network,

    K. Ali and C. E. Hughes, “All-in-one: Facial expression transfer, editing and recognition using a single network,” arXiv preprint arXiv:1911.07050, 2019

  91. [95]

    Semantic prior guided fine- grained facial expression manipulation,

    T. Xue, J. Yan, D. Zheng, and Y. Liu, “Semantic prior guided fine- grained facial expression manipulation,” Complex & Intelligent Systems, pp. 1–16, 2024

  92. [96]

    Wp2-gan: Wavelet-based multi-level gan for progressive facial expression translation with parallel gen- erators,

    J. Shao and T. Bui, “Wp2-gan: Wavelet-based multi-level gan for progressive facial expression translation with parallel gen- erators,” in Proc. British Mach. Vis. Conf., 2021, pp. 1388–1

  93. [97]

    Headgan: One- shot neural head synthesis and editing,

    M. C. Doukas, S. Zafeiriou, and V . Sharmanska, “Headgan: One- shot neural head synthesis and editing,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 14 398–14 407

  94. [98]

    Action unit driven facial expression synthesis from a single image with patch attentive gan,

    Y. Zhao, L. Yang, E. Pei, M. C. Oveneke, M. Alioscha-Perez, L. Li, D. Jiang, and H. Sahli, “Action unit driven facial expression synthesis from a single image with patch attentive gan,” in Computer Graphics Forum , vol. 40. Wiley Online Library, 2021, pp. 47–61

  95. [99]

    Local and global perception generative adversarial network for facial expression synthesis,

    Y. Xia, W. Zheng, Y. Wang, H. Yu, J. Dong, and F.-Y. Wang, “Local and global perception generative adversarial network for facial expression synthesis,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 3, pp. 1443–1452, 2021

  96. [100]

    Dft-net: Disentanglement of face deformation and texture synthesis for expression editing,

    J. Wang, J. Zhang, Z. Lu, and S. Shan, “Dft-net: Disentanglement of face deformation and texture synthesis for expression editing,” in 2019 IEEE International Conference on Image Processing (ICIP) . IEEE, 2019, pp. 3881–3885

  97. [101]

    Sargan: Spatial attention-based resid- uals for facial expression manipulation,

    A. Akram and N. Khan, “Sargan: Spatial attention-based resid- uals for facial expression manipulation,” IEEE Transactions on Circuits and Systems for Video Technology, 2023

  98. [102]

    Gated switch- gan for multi-domain facial image translation,

    X. Zhang, Y. Zhu, W. Chen, W. Liu, and L. Shen, “Gated switch- gan for multi-domain facial image translation,” IEEE Transactions on Multimedia, vol. 24, pp. 1990–2003, 2021

  99. [103]

    Us-gan: On the importance of ultimate skip connection for facial expression synthesis,

    A. Akram and N. Khan, “Us-gan: On the importance of ultimate skip connection for facial expression synthesis,” Multimedia Tools and Applications, vol. 83, no. 3, pp. 7231–7247, 2024. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 22

  100. [104]

    Emostyle: One-shot facial expression editing using continuous emotion parameters,

    B. Azari and A. Lim, “Emostyle: One-shot facial expression editing using continuous emotion parameters,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 6385–6394

  101. [105]

    Ganmut: Learning interpretable conditional space for gamut of emotions,

    S. d’Apolito, D. P . Paudel, Z. Huang, A. Romero, and L. V . Gool, “Ganmut: Learning interpretable conditional space for gamut of emotions,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 568–577

  102. [106]

    Apprgan: Appearance-based gan for facial expression synthesis,

    Y. Peng and H. Yin, “Apprgan: Appearance-based gan for facial expression synthesis,” IET Image Processing , vol. 13, no. 14, pp. 2706–2715, 2019

  103. [107]

    Geometry guided adversarial facial expression synthesis,

    L. Song, Z. Lu, R. He, Z. Sun, and T. Tan, “Geometry guided adversarial facial expression synthesis,” in Proceedings of the 26th ACM International Conference on Multimedia, 2018, pp. 627–635

  104. [108]

    Dualpathgan: Facial reenacted emotion synthesis,

    J. Kong, H. Shen, and K. Huang, “Dualpathgan: Facial reenacted emotion synthesis,” IET Computer Vision, vol. 15, no. 7, pp. 501– 513, 2021

  105. [109]

    Fine-grained expression manip- ulation via structured latent space,

    J. Tang, Z. Shao, and L. Ma, “Fine-grained expression manip- ulation via structured latent space,” in 2020 IEEE International Conference on Multimedia and Expo (ICME) . IEEE, 2020, pp. 1–6

  106. [110]

    Styleclip: Text-driven manipulation of stylegan imagery,

    O. Patashnik, Z. Wu, E. Shechtman, D. Cohen-Or, and D. Lischin- ski, “Styleclip: Text-driven manipulation of stylegan imagery,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 2085–2094

  107. [111]

    Evogan: An evolutionary computation assisted gan,

    F. Liu, H. Wang, J. Zhang, Z. Fu, A. Zhou, J. Qi, and Z. Li, “Evogan: An evolutionary computation assisted gan,” Neurocom- puting, vol. 469, pp. 81–90, 2022

  108. [112]

    Expression conditional gan for facial expression-to-expression translation,

    H. Tang, W. Wang, S. Wu, X. Chen, D. Xu, N. Sebe, and Y. Yan, “Expression conditional gan for facial expression-to-expression translation,” in 2019 IEEE International Conference on Image Pro- cessing (ICIP). IEEE, 2019, pp. 4449–4453

  109. [113]

    Unmasking your expression: Expression- conditioned gan for masked face inpainting,

    S. Sola and D. Gera, “Unmasking your expression: Expression- conditioned gan for masked face inpainting,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 5907–5915

  110. [114]

    Facial expres- sion editing with continuous emotion labels,

    A. Lindt, P . Barros, H. Siqueira, and S. Wermter, “Facial expres- sion editing with continuous emotion labels,” in 2019 14th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2019). IEEE, 2019, pp. 1–8

  111. [115]

    Ad- versarial autoencoders,

    A. Makhzani, J. Shlens, N. Jaitly, I. Goodfellow, and B. Frey, “Ad- versarial autoencoders,” arXiv preprint arXiv:1511.05644, 2015

  112. [116]

    Speech driven talking face generation from a single image and an emotion condition,

    S. E. Eskimez, Y. Zhang, and Z. Duan, “Speech driven talking face generation from a single image and an emotion condition,” IEEE Transactions on Multimedia, vol. 24, pp. 3480–3490, 2021

  113. [117]

    Realistic speech- driven facial animation with gans,

    K. Vougioukas, S. Petridis, and M. Pantic, “Realistic speech- driven facial animation with gans,” International Journal of Com- puter Vision, vol. 128, no. 5, pp. 1398–1413, 2020

  114. [118]

    Talking face generation with expression-tailored generative adversarial network,

    D. Zeng, H. Liu, H. Lin, and S. Ge, “Talking face generation with expression-tailored generative adversarial network,” in Proceed- ings of the 28th ACM International Conference on Multimedia , 2020, pp. 1716–1724

  115. [119]

    Efficient emotional adaptation for audio-driven talking-head generation,

    Y. Gan, Z. Yang, X. Yue, L. Sun, and Y. Yang, “Efficient emotional adaptation for audio-driven talking-head generation,” in Proceed- ings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 22 634–22 645

  116. [120]

    Flowvqtalker: High-quality emotional talking face generation through normalizing flow and quantiza- tion,

    S. Tan, B. Ji, and Y. Pan, “Flowvqtalker: High-quality emotional talking face generation through normalizing flow and quantiza- tion,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26 317–26 327

  117. [121]

    Edtalk: Efficient disentan- glement for emotional talking head synthesis,

    S. Tan, B. Ji, M. Bi, and et al., “Edtalk: Efficient disentan- glement for emotional talking head synthesis,” arXiv preprint arXiv:2404.01647, 2024

  118. [122]

    2cet-gan: Pixel-level gan model for human facial expression transfer,

    X. Hu, N. Aldausari, and G. Mohammadi, “2cet-gan: Pixel-level gan model for human facial expression transfer,” in Proceedings of the 1st International Workshop on Multimedia Content Generation and Evaluation: New Methods and Practice , 2023, pp. 49–56

  119. [123]

    Star- gan: Unified generative adversarial networks for multi-domain image-to-image translation,

    Y. Choi, M. Choi, M. Kim, J.-W. Ha, S. Kim, and J. Choo, “Star- gan: Unified generative adversarial networks for multi-domain image-to-image translation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 8789–8797

  120. [124]

    Ugan: Untraceable gan for multi-domain face translation,

    D. Zhu, S. Liu, W. Jiang, C. Gao, T. Wu, Q. Wang, and G. Guo, “Ugan: Untraceable gan for multi-domain face translation,”arXiv preprint arXiv:1907.11418, 2019

  121. [128]

    Self-supervised emotion represen- tation disentanglement for speech-preserving facial expression manipulation,

    Z. Xu, T. Chen, Z. Yang et al., “Self-supervised emotion represen- tation disentanglement for speech-preserving facial expression manipulation,” in Proceedings of ACM Multimedia 2024, 2024

  122. [129]

    Emmn: Emotional motion memory network for audio-driven emotional talking face generation,

    S. Tan, B. Ji, and Y. Pan, “Emmn: Emotional motion memory network for audio-driven emotional talking face generation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 22 146–22 156

  123. [130]

    Style2talker: High-resolution talking head generation with emotion style and art style,

    ——, “Style2talker: High-resolution talking head generation with emotion style and art style,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, 2024, pp. 5079–5087

  124. [131]

    Talking face generation with audio-deduced emotional landmarks,

    S. Zhai, M. Liu, Y. Li, Z. Gao, L. Zhu, and L. Nie, “Talking face generation with audio-deduced emotional landmarks,” IEEE Transactions on Neural Networks and Learning Systems, 2023

  125. [132]

    Stochastic latent talking face generation towards emotional expressions and head poses,

    Z. Sheng, L. Nie, M. Zhang, X. Chang, and Y. Yan, “Stochastic latent talking face generation towards emotional expressions and head poses,” IEEE Transactions on Circuits and Systems for Video Technology, 2023

  126. [135]

    Continuously controllable facial expression editing in talking face videos,

    Z. Sun, Y. H. Wen, T. Lv et al., “Continuously controllable facial expression editing in talking face videos,” IEEE Transactions on Affective Computing, 2023

  127. [136]

    Eat-face: Emotion-controllable audio-driven talking face generation via diffusion model,

    H. Wang, X. Jia, and X. Cao, “Eat-face: Emotion-controllable audio-driven talking face generation via diffusion model,” inPro- ceedings of the 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG). IEEE, 2024, pp. 1–10

  128. [137]

    Converting anyone’s emotion: Towards speaker-independent emotional voice conver- sion,

    K. Zhou, B. Sisman, M. Zhang, and H. Li, “Converting anyone’s emotion: Towards speaker-independent emotional voice conver- sion,” arXiv preprint arXiv:2005.07025, 2020

  129. [138]

    Vaw-gan for disentanglement and recomposition of emotional elements in speech,

    K. Zhou, B. Sisman, and H. Li, “Vaw-gan for disentanglement and recomposition of emotional elements in speech,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 415–422

  130. [139]

    Transforming spectrum and prosody for emotional voice conversion with non-parallel training data,

    ——, “Transforming spectrum and prosody for emotional voice conversion with non-parallel training data,” arXiv preprint arXiv:2002.00198, 2020

  131. [140]

    Unpaired image-to- image translation using cycle-consistent adversarial networks,

    J.-Y. Zhu, T. Park, P . Isola, and A. A. Efros, “Unpaired image-to- image translation using cycle-consistent adversarial networks,” in Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 2223–2232

  132. [141]

    An improved cyclegan- based emotional voice conversion model by augmenting tempo- ral dependency with a transformer,

    C. Fu, C. Liu, C. T. Ishi, and H. Ishiguro, “An improved cyclegan- based emotional voice conversion model by augmenting tempo- ral dependency with a transformer,” Speech Communication, vol. 144, pp. 110–121, 2022

  133. [144]

    Nonparallel emotional speech conversion,

    J. Gao, D. Chakraborty, H. Tembine, and O. Olaleye, “Nonparallel emotional speech conversion,” arXiv preprint arXiv:1811.01174 , 2018

  134. [145]

    Attention-based interactive disentangling network for instance-level emotional voice conversion,

    Y. Chen, L. Yang, Q. Chen, J.-H. Lai, and X. Xie, “Attention-based interactive disentangling network for instance-level emotional voice conversion,” arXiv preprint arXiv:2312.17508, 2023

  135. [147]

    Expressive voice con- version: A joint framework for speaker identity and emotional style transfer,

    Z. Du, B. Sisman, K. Zhou, and H. Li, “Expressive voice con- version: A joint framework for speaker identity and emotional style transfer,” in 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2021, pp. 594–601

  136. [149]

    Nonparallel emotional voice conversion for unseen speaker-emotion pairs us- ing dual domain adversarial network & virtual domain pairing,

    N. Shah, M. Singh, N. Takahashi, and N. Onoe, “Nonparallel emotional voice conversion for unseen speaker-emotion pairs us- ing dual domain adversarial network & virtual domain pairing,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processin...

  137. [151]

    Emotion intensity and its control for emotional voice conversion,

    K. Zhou, B. Sisman, R. Rana, B. W. Schuller, and H. Li, “Emotion intensity and its control for emotional voice conversion,” IEEE Transactions on Affective Computing, vol. 14, no. 1, pp. 31–48, 2022

  138. [152]

    Textless speech emotion conversion using discrete & decom- posed representations,

    F. Kreuk, A. Polyak, J. Copet, E. Kharitonov, T. A. Nguyen, M. Rivière, W.-N. Hsu, A. Mohamed, E. Dupoux, and Y. Adi, “Textless speech emotion conversion using discrete & decom- posed representations,” in Proceedings of the 2022 Conference on Empirical Methods in Natural Langu...

  139. [154]

    Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis,

    Y. Lei, S. Yang, and L. Xie, “Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis,” in 2021 IEEE Spoken Language Technology Workshop (SLT) . IEEE, 2021, pp. 423–430

  140. [155]

    Diclet-tts: Diffusion model based cross-lingual emotion transfer for text-to-speech—a study between english and mandarin,

    T. Li, C. Hu, J. Cong, X. Zhu, J. Li, Q. Tian, Y. Wang, and L. Xie, “Diclet-tts: Diffusion model based cross-lingual emotion transfer for text-to-speech—a study between english and mandarin,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023

  141. [156]

    Exploring transfer learn- ing for low resource emotional tts,

    N. Tits, K. E. Haddad, and T. Dutoit, “Exploring transfer learn- ing for low resource emotional tts,” in Intelligent Systems and Applications: Proceedings of the 2019 Intelligent Systems Conference (IntelliSys) Volume 1. Springer, 2020, pp. 52–60

  142. [157]

    Improving emotional tts with an emotion intensity input from unsupervised extraction,

    B. Schnell and P . N. Garner, “Improving emotional tts with an emotion intensity input from unsupervised extraction,” in Proc. 11th ISCA Speech Synth. Workshop, 2021, pp. 60–65

  143. [158]

    End-to- end emotional speech synthesis using style tokens and semi- supervised training,

    P . Wu, Z. Ling, L. Liu, Y. Jiang, H. Wu, and L. Dai, “End-to- end emotional speech synthesis using style tokens and semi- supervised training,” in 2019 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIP A ASC). IEEE, 2019, pp. 623–627

  144. [159]

    Emotional speech synthesis with rich and granularized control,

    S.-Y. Um, S. Oh, K. Byun, I. Jang, C. Ahn, and H.-G. Kang, “Emotional speech synthesis with rich and granularized control,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2020, pp. 7254–7258

  145. [160]

    Gantron: Emotional speech syn- thesis with generative adversarial networks,

    E. Hortal and R. B. Alarcia, “Gantron: Emotional speech syn- thesis with generative adversarial networks,” arXiv preprint arXiv:2110.03390, 2021

  146. [161]

    Emodiff: Intensity con- trollable emotional text-to-speech with soft-label guidance,

    Y. Guo, C. Du, X. Chen, and K. Yu, “Emodiff: Intensity con- trollable emotional text-to-speech with soft-label guidance,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5

  147. [163]

    Msemotts: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis,

    Y. Lei, S. Yang, X. Wang, and L. Xie, “Msemotts: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, pp. 853–864, 2022

  148. [164]

    Cross-speaker emo- tion disentangling and transfer for end-to-end speech synthesis,

    T. Li, X. Wang, Q. Xie, Z. Wang, and L. Xie, “Cross-speaker emo- tion disentangling and transfer for end-to-end speech synthesis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 30, pp. 1448–1460, 2022

  149. [165]

    Towards multi- scale style control for expressive speech synthesis,

    X. Li, C. Song, J. Li, Z. Wu, J. Jia, and H. Meng, “Towards multi- scale style control for expressive speech synthesis,” arXiv preprint arXiv:2104.03521, 2021

  150. [167]

    Controlling the strength of emotions in speech-like emotional sound generated by wavenet,

    K. Matsumoto, S. Hara, and M. Abe, “Controlling the strength of emotions in speech-like emotional sound generated by wavenet,” in INTERSPEECH, 2020, pp. 3421–3425

  151. [168]

    Emotion selectable end-to-end text-based speech editing,

    T. Wang, J. Yi, R. Fu, J. Tao, Z. Wen, and C. Y. Zhang, “Emotion selectable end-to-end text-based speech editing,” Artificial Intelli- gence, vol. 329, p. 104076, 2024

  152. [169]

    Fastspeech: Fast, robust and controllable text to speech,

    Y. Ren, Y. Ruan, X. Tan, T. Qin, S. Zhao, Z. Zhao, and T.-Y. Liu, “Fastspeech: Fast, robust and controllable text to speech,” in Advances in Neural Information Processing Systems, vol. 32, 2019

  153. [170]

    Towards realistic emotional voice conversion using controllable emotional intensity,

    T. Qi, S. Wang, C. Lu et al. , “Towards realistic emotional voice conversion using controllable emotional intensity,” arXiv preprint arXiv:2407.14800, 2024

  154. [171]

    Zet-speech: Zero-shot adaptive emotion-controllable text-to-speech synthe- sis with diffusion and style-based models,

    M. Kang, W. Han, S. J. Hwang, and E. Yang, “Zet-speech: Zero-shot adaptive emotion-controllable text-to-speech synthe- sis with diffusion and style-based models,” arXiv preprint arXiv:2305.13831, 2023

  155. [172]

    Controlling emotion strength with relative attribute for end-to-end speech synthesis,

    X. Zhu, S. Yang, G. Yang, and L. Xie, “Controlling emotion strength with relative attribute for end-to-end speech synthesis,” in 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2019, pp. 192–199

  156. [174]

    Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,

    X. Cai, D. Dai, Z. Wu, X. Li, J. Li, and H. Meng, “Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processin...

  157. [175]

    Fine-grained emotional control of text-to-speech: Learning to rank inter-and intra-class emotion intensities,

    S. Wang, J. Guðnason, and D. Borth, “Fine-grained emotional control of text-to-speech: Learning to rank inter-and intra-class emotion intensities,” in ICASSP 2023-2023 IEEE International Con- ference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5

  158. [177]

    Emospeech: Guiding fastspeech2 to- wards emotional text to speech,

    D. Diatlova and V . Shutov, “Emospeech: Guiding fastspeech2 to- wards emotional text to speech,” arXiv preprint arXiv:2307.00024, 2023

  159. [178]

    Mm-tts: Multi-modal prompt based style transfer for expressive text-to-speech synthesis,

    W. Guan, Y. Li, T. Li et al., “Mm-tts: Multi-modal prompt based style transfer for expressive text-to-speech synthesis,” in Proceed- ings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 16, 2024, pp. 18 117–18 125

  160. [180]

    Ed-tts: Multi-scale emotion modeling using cross-domain emotion diarization for emotional speech synthesis,

    H. Tang, X. Zhang, N. Cheng et al., “Ed-tts: Multi-scale emotion modeling using cross-domain emotion diarization for emotional speech synthesis,” in Proceedings of the IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 12 146–12 150

  161. [181]

    Enhancing emo- tional text-to-speech controllability with natural language guid- ance through contrastive learning and diffusion models,

    X. Jing, K. Zhou, A. Triantafyllopoulos et al. , “Enhancing emo- tional text-to-speech controllability with natural language guid- ance through contrastive learning and diffusion models,” arXiv preprint arXiv:2409.06451, 2024

  162. [182]

    Rset: Remapping-based sorting method for emotion transfer speech synthesis,

    H. Shi, J. Wang, X. Zhang et al., “Rset: Remapping-based sorting method for emotion transfer speech synthesis,” in Proceedings of Asia-Pacific Web (APWeb) and Web-Age Information Manage- ment (WAIM) Joint International Conference on Web and Big Data . Springer Nature Singapore...

  163. [185]

    Imat: Unsupervised text attribute transfer via iterative matching and translation,

    Z. Jin, D. Jin, J. Mueller, N. Matthews, and E. Santus, “Imat: Unsupervised text attribute transfer via iterative matching and translation,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on ...

  164. [186]

    Mask and infill: JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 24 Applying masked language model to sentiment transfer,

    X. Wu, T. Zhang, L. Zang, J. Han, and S. Hu, “Mask and infill: JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 24 Applying masked language model to sentiment transfer,” Aug 2019

  165. [187]

    Unsupervised text style transfer with padded masked language models,

    E. Malmi, A. Severyn, and S. Rothe, “Unsupervised text style transfer with padded masked language models,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 8671–8680

  166. [188]

    Unsupervised text style transfer using language models as discriminators,

    Z. Yang, Z. Hu, C. Dyer, E. P . Xing, and T. Berg-Kirkpatrick, “Unsupervised text style transfer using language models as discriminators,” Advances in Neural Information Processing Systems, vol. 31, 2018

  167. [189]

    Style transfer from non-parallel text by cross-alignment,

    T. Shen, T. Lei, R. Barzilay, and T. Jaakkola, “Style transfer from non-parallel text by cross-alignment,” Advances in Neural Information Processing Systems, vol. 30, 2017

  168. [190]

    Emotional text gen- eration based on cross-domain sentiment transfer,

    R. Zhang, Z. Wang, K. Yin, and Z. Huang, “Emotional text gen- eration based on cross-domain sentiment transfer,” IEEE Access, vol. 7, pp. 100 081–100 089, 2019

  169. [191]

    Cycle- consistent adversarial autoencoders for unsupervised text style transfer,

    Y. Huang, W. Zhu, D. Xiong, Y. Zhang, C. Hu, and F. Xu, “Cycle- consistent adversarial autoencoders for unsupervised text style transfer,” in Proceedings of the 28th International Conference on Computational Linguistics, 2020, pp. 2213–2223

  170. [192]

    Towards fine-grained text sentiment transfer,

    F. Luo, P . Li, P . Yang, J. Zhou, Y. Tan, B. Chang, Z. Sui, and X. Sun, “Towards fine-grained text sentiment transfer,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019, pp. 2013–2022

  171. [194]

    Enhancing empathetic response generation by aug- menting llms with small-scale empathetic models,

    Z. Yang, Z. Ren, W. Yufeng, S. Peng, H. Sun, X. Zhu, and X. Liao, “Enhancing empathetic response generation by aug- menting llms with small-scale empathetic models,” arXiv preprint arXiv:2402.11801, 2024

  172. [195]

    Rational sensi- bility: Llm enhanced empathetic response generation guided by self-presentation theory,

    L. Sun, N. Xu, J. Wei, B. Yu, L. Bu, and Y. Luo, “Rational sensi- bility: Llm enhanced empathetic response generation guided by self-presentation theory,” arXiv preprint arXiv:2312.08702, 2023

  173. [197]

    Soulchat: Improving llms’ empathy, listening, and comfort abili- ties through fine-tuning with multi-turn empathy conversations,

    Y. Chen, X. Xing, J. Lin, H. Zheng, Z. Wang, Q. Liu, and X. Xu, “Soulchat: Improving llms’ empathy, listening, and comfort abili- ties through fine-tuning with multi-turn empathy conversations,” in Findings of the Association for Computational Linguistics: EMNLP 2023, 2023, pp...

  174. [198]

    Generating responses with a specific emotion in dialog,

    Z. Song, X. Zheng, L. Liu, M. Xu, and X.-J. Huang, “Generating responses with a specific emotion in dialog,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019, pp. 3685–3695

  175. [199]

    Emotional chatting machine: Emotional conversation generation with inter- nal and external memory,

    H. Zhou, M. Huang, T. Zhang, X. Zhu, and B. Liu, “Emotional chatting machine: Emotional conversation generation with inter- nal and external memory,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, 2018

  176. [201]

    Dual-view conditional vari- ational auto-encoder for emotional dialogue generation,

    M. Li, J. Zhang, X. Lu, and C. Zong, “Dual-view conditional vari- ational auto-encoder for emotional dialogue generation,” Trans- actions on Asian and Low-Resource Language Information Processing , vol. 21, no. 3, pp. 1–18, 2021

  177. [202]

    Generating emotional controllable response based on multi-task and dual attention framework,

    W. Xu, X. Gu, and G. Chen, “Generating emotional controllable response based on multi-task and dual attention framework,” IEEE Access, vol. 7, pp. 93 734–93 741, 2019

  178. [203]

    Affective neural response generation,

    N. Asghar, P . Poupart, J. Hoey, X. Jiang, and L. Mou, “Affective neural response generation,” in Advances in Information Retrieval: 40th European Conference on IR Research, ECIR 2018 . Springer, 2018, pp. 154–166

  179. [204]

    Affect-driven dialog generation,

    P . Colombo, W. Witon, A. Modi, J. Kennedy, and M. Kapadia, “Affect-driven dialog generation,” arXiv preprint arXiv:1904.02793, 2019

  180. [205]

    Automatic dialogue generation with expressed emotions,

    C. Huang, O. R. Zaiane, A. Trabelsi, and N. Dziri, “Automatic dialogue generation with expressed emotions,” in Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, vol. 2, 2018, pp. 49–54

  181. [206]

    Emotional dialogue generation based on transformer and conditional variational autoencoder,

    H. Lin and Z. Deng, “Emotional dialogue generation based on transformer and conditional variational autoencoder,” in 2022 IEEE 21st International Conference on Ubiquitous Computing and Communications (IUCC/CIT/DSCI/SmartCNS) . IEEE, 2022, pp. 386–393

  182. [207]

    Sentigan: Generating sentimental texts via mixture adversarial networks,

    K. Wang and X. Wan, “Sentigan: Generating sentimental texts via mixture adversarial networks,” in IJCAI, 2018, pp. 4446–4452

  183. [208]

    Automatic generation of sentimental texts via mixture adversarial networks,

    ——, “Automatic generation of sentimental texts via mixture adversarial networks,” Artificial Intelligence, vol. 275, pp. 540–558, 2019

  184. [209]

    Emotional dialog generation via multiple classifiers based on a generative adversarial network,

    W. Chen, X. Chen, and X. Sun, “Emotional dialog generation via multiple classifiers based on a generative adversarial network,” Virtual Reality & Intelligent Hardware, vol. 3, no. 1, pp. 18–32, 2021

  185. [211]

    Customizable text generation via conditional text generative adversarial net- work,

    J. Chen, Y. Wu, C. Jia, H. Zheng, and G. Huang, “Customizable text generation via conditional text generative adversarial net- work,” Neurocomputing, vol. 416, pp. 125–135, 2020

  186. [212]

    A recipe for arbitrary text style transfer with large language models,

    E. Reif, D. Ippolito, A. Yuan, A. Coenen, C. Callison-Burch, and J. Wei, “A recipe for arbitrary text style transfer with large language models,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , 2022, pp. 837–848

  187. [213]

    Text style transfer via learning style instance supported latent space,

    X. Yi, Z. Liu, W. Li, and M. Sun, “Text style transfer via learning style instance supported latent space,” in Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence, 2021, pp. 3801–3807

  188. [214]

    Dgst: A dual-generator network for text style transfer,

    X. Li, G. Chen, C. Lin, and R. Li, “Dgst: A dual-generator network for text style transfer,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2020, pp. 7131–7136

  189. [215]

    Re- inforced rewards framework for text style transfer,

    A. Sancheti, K. Krishna, B. V . Srinivasan, and A. Natarajan, “Re- inforced rewards framework for text style transfer,” in Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portugal, April 14–17, 2020, Proceedings, Part I , vol. 42....

  190. [216]

    Semi-supervised formality style trans- fer using language model discriminator and mutual information maximization,

    K. Chawla and D. Yang, “Semi-supervised formality style trans- fer using language model discriminator and mutual information maximization,” in Findings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 2340–2354

  191. [218]

    Harnessing the power of large language models for empathetic response generation: Empirical investigations and improvements,

    Y. Qian, W. Zhang, and T. Liu, “Harnessing the power of large language models for empathetic response generation: Empirical investigations and improvements,” in The 2023 Conference on Empirical Methods in Natural Language Processing , 2023

  192. [219]

    Emosen: Generating sentiment and emotion controlled re- sponses in a multimodal dialogue system,

    M. Firdaus, H. Chauhan, A. Ekbal, and P . Bhattacharyya, “Emosen: Generating sentiment and emotion controlled re- sponses in a multimodal dialogue system,” IEEE Transactions on Affective Computing, vol. 13, no. 3, pp. 1555–1566, 2020

  193. [220]

    Mime: Mimicking emotions for empa- thetic response generation,

    N. Majumder, P . Hong, S. Peng, J. Lu, D. Ghosal, A. Gelbukh, R. Mihalcea, and S. Poria, “Mime: Mimicking emotions for empa- thetic response generation,” arXiv preprint arXiv:2010.01454, 2020

  194. [221]

    Cem: Commonsense-aware empathetic response generation,

    S. Sabour, C. Zheng, and M. Huang, “Cem: Commonsense-aware empathetic response generation,” in Proceedings of the AAAI Con- ference on Artificial Intelligence, vol. 36, 2022, pp. 11 229–11 237

  195. [223]

    Emotional dialogue generation with generative adversarial networks,

    Y. Li and B. Wu, “Emotional dialogue generation with generative adversarial networks,” in 2020 IEEE 4th Information Technology, Networking, Electronic and Automation Control Conference (ITNEC) , vol. 1. IEEE, 2020, pp. 868–873

  196. [226]

    Audio-driven emotional video portraits,

    X. Ji, H. Zhou, K. Wang, W. Wu, C. C. Loy, X. Cao, and F. Xu, “Audio-driven emotional video portraits,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 14 080–14 089. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 25

  197. [227]

    Gaussianpu: A hybrid 2d-3d upsampling framework for enhancing color point clouds via 3d gaussian splatting,

    Z. Guo, Y. Xie, W. Xie, P . Huang, F. Ma, and F. R. Yu, “Gaussianpu: A hybrid 2d-3d upsampling framework for enhancing color point clouds via 3d gaussian splatting,” arXiv preprint arXiv:2409.01581, 2024

  198. [228]

    Interactive multi-level prosody control for expressive speech synthesis,

    T. Cornille, F. Wang, and J. Bekker, “Interactive multi-level prosody control for expressive speech synthesis,” in ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2022, pp. 8312–8316

  199. [229]

    Tacotron: Towards end- to-end speech synthesis,

    Y. Wang, R. Skerry-Ryan, D. Stanton, Y. Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y. Xiao, Z. Chen, S. Bengioet al., “Tacotron: Towards end- to-end speech synthesis,” arXiv preprint arXiv:1703.10135, 2017

  200. [230]

    Large language models (llms) and empathy—a systematic review,

    V . Sorin, D. Brin, Y. Barash, E. Konen, A. Charney, G. Nadkarni, and E. Klang, “Large language models (llms) and empathy—a systematic review,” medRxiv, pp. 2023–08, 2023

  201. [231]

    Multimodal ma- chine learning: A survey and taxonomy,

    T. Baltrušaitis, C. Ahuja, and L.-P . Morency, “Multimodal ma- chine learning: A survey and taxonomy,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 41, no. 2, pp. 423– 443, 2018

  202. [232]

    Visualrwkv: Exploring recurrent neural networks for visual language models,

    H. Hou, P . Zeng, F. Ma, and F. R. Yu, “Visualrwkv: Exploring recurrent neural networks for visual language models,” arXiv preprint arXiv:2406.13362, 2024

  203. [233]

    A comprehensive review of data-driven co-speech gesture gen- eration,

    S. Nyatsanga, T. Kucherenko, C. Ahuja, G. E. Henter, and M. Neff, “A comprehensive review of data-driven co-speech gesture gen- eration,” Computer Graphics Forum, vol. 42, pp. 569–596, 2023

  204. [234]

    Virtual and augmented reality for developing emotional intelligence skills,

    C. Papoutsi, A. Drigas, and C. Skianis, “Virtual and augmented reality for developing emotional intelligence skills,” Int. J. Recent Contrib. Eng. Sci. IT (IJES), vol. 9, no. 3, pp. 35–53, 2021

  205. [235]

    An overview on edge computing research,

    K. Cao, Y. Liu, G. Meng, and Q. Sun, “An overview on edge computing research,” IEEE access, vol. 8, pp. 85 714–85 728, 2020

  206. [236]

    Wearables and the in- ternet of things (iot), applications, opportunities, and challenges: A survey,

    F. J. Dian, R. Vahidnia, and A. Rahmati, “Wearables and the in- ternet of things (iot), applications, opportunities, and challenges: A survey,” IEEE access, vol. 8, pp. 69 200–69 211, 2020

  207. [237]

    The impact of artificial intelligence on animation filmmaking: Tools, trends, and future implications,

    M. Izani, A. Razak, D. Rehad, and M. Rosli, “The impact of artificial intelligence on animation filmmaking: Tools, trends, and future implications,” in 2024 International Visualization, Informatics and Technology Conference (IVIT). IEEE, 2024, pp. 57–62

  208. [238]

    Original research article revolutionizing filmmaking: A comparative analysis of conventional and ai-generated film production in the era of virtual reality,

    A. Channa, A. Sharma, M. Singh, P . Malhotra, A. Bajpai, and P . Whig, “Original research article revolutionizing filmmaking: A comparative analysis of conventional and ai-generated film production in the era of virtual reality,” Journal of Autonomous Intelligence, vol. 7, no. 4, 2024

  209. [239]

    The application and user acceptance of aigc in network audiovisual field: Based on the perspective of social cognitive theory,

    H. Sun and Y. He, “The application and user acceptance of aigc in network audiovisual field: Based on the perspective of social cognitive theory,” in Proceedings of the 7th International Conference on Computer Science and Application Engineering , 2023, pp. 1–5

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.