Pith. sign in

REVIEW 4 major objections 6 minor 77 references

Looking at the user’s face lets a speech synthesizer produce more natural, emotionally fitting replies in conversation.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-31 15:04 UTC pith:22ZYC5VF

load-bearing objection Solid multimodal CSS systems paper with a real dataset and a compact AU tokenizer; the face-causality claim is under-supported, but the engineering package still deserves referee time. the 4 major comments →

arxiv 2607.24430 v1 pith:22ZYC5VF submitted 2026-07-27 cs.HC cs.CLeess.AS

Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis

classification cs.HC cs.CLeess.AS
keywords Conversational Speech SynthesisFacial Expression ModelingAction UnitsAUTokenizerDualDPOMultimodal DialogueEmpathetic SpeechVSDD-1K
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Conversational speech systems usually hear the user and read the transcript, but they ignore the face—where much of the emotion actually lives. This paper argues that those facial cues can be packed into a single discrete token per video frame and fed to a language-model backbone so the system can choose both an emotion label and speech that match how the user looks as well as what they said. The authors build a tokenizer supervised by facial Action Unit combinations, add a preference-training step that jointly ranks better face tokens and better speech tokens, and release a thousand-hour face-to-face dialogue corpus scraped and cleaned automatically from real interviews and podcasts. On several multimodal dialogue test sets the resulting system beats both non-visual conversational synthesizers and prior vision-aware models on speaker similarity, prosody match, emotion accuracy, and human ratings of naturalness and emotional fit. A sympathetic reader cares because empathetic spoken agents in homes, cars, and care settings need more than text and audio; this work claims a practical way to give them the missing visual channel without exploding sequence length.

Core claim

FacialTalker shows that encoding each user face frame as one Action-Unit-supervised discrete token, modeling it jointly with text and speech inside an LLM, and post-training with dual preference constraints on face and speech sequences, yields conversational speech that is more natural, expressive, and context-aligned than strong baselines that lack this facial pathway—or that use coarser visual encoders.

What carries the argument

AUTokenizer: a single-codebook Finite Scalar Quantization tokenizer that maps each facial frame to one discrete token under supervision from combinations of facial Action Units, plus DualDPO, which extends direct preference optimization to rank both visual and speech token sequences together.

Load-bearing premise

The load-bearing premise is that Action Unit labels from small micro-expression corpora, compressed into one token per frame, carry enough spontaneous conversational affect to improve target emotion and prosody—not merely that more data or speech-side training would have produced the same gains.

What would settle it

Retrain FacialTalker on the same large dialogue data but replace AUTokenizer with a scrambled or constant face token (or drop DualDPO’s visual preference term) and check whether emotion accuracy, prosody distance, and human MOS_E on MultiDialog, AvaMERG, and VSDD-1K collapse back to non-visual or CLIP-based baselines.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Empathetic CSS systems can treat user face video as a first-class context stream without multi-token image codecs.
  • A single AU-supervised face token per frame is a usable interface between facial analysis and autoregressive speech LLMs.
  • Joint preference optimization on visual and speech tokens is a transferable post-training recipe for multimodal dialogue agents.
  • The automated VSDD-1K pipeline supplies open, large-scale face-aligned dialogue data for other speech and affect tasks.
  • Agents in homes, cockpits, and eldercare can condition replies on how the user looks, not only on what they say.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If one AU token suffices here, the same compression may unlock face-conditioned spoken dialogue on smaller on-device LLMs where multi-token vision is too costly.
  • Failures will likely concentrate on AU combinations rare in CASME II/DISFA but common in spontaneous talk (e.g., polite smiles vs genuine joy), suggesting a need for in-the-wild AU labels.
  • DualDPO’s visual rejected samples are model-generated face tokens; the same idea could regularize other sparse nonverbal streams such as gaze or posture tokens.
  • Releasing 1K hours of real face-to-face talk may matter as much as the model for follow-on work on turn-taking and micro-expression-aware dialogue.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper presents FacialTalker, an LLM-based (Qwen2.5-0.5B, initialized from CosyVoice2) conversational speech synthesis system that conditions on user facial expressions in addition to text/speech history. Three main components are claimed: (i) AUTokenizer, a ConvNeXt-Tiny + FSQ visual tokenizer that compresses each face frame into a single discrete token supervised by Action Unit combination labels from CASME II/DISFA; (ii) DualDPO, a post-training stage applying DPO preference pairs to both speech and facial token sequences (rejected samples drawn from the Stage-3 model's own outputs); and (iii) VSDD-1K, a 1,033-hour automatically constructed video-speech dialogue dataset from interviews/podcasts with >85% valid-face frames. Evaluation on MultiDialog, AvaMERG, and VSDD-1K (Tables 4–5) shows consistent gains over non-visual CSS baselines and visual baselines (Empatheia, EmpathyEar, UniTalker) in SIM, PDTW, ACC_E, and MOS with 95% CIs from 50 raters, supported by ablations (FT-base, FT-CLIP, AUTokenizer-VQ) and AU F1 benchmarks (Table 3).

Significance. If the central claim holds, this is a useful contribution: it is, to my knowledge, the first LLM-based CSS system to incorporate frame-level facial affect at this scale, and VSDD-1K — a 1K-hour open-source synchronized face-speech dialogue corpus with a documented, reproducible construction pipeline (§5.1) — is independently valuable to the community. The experimental apparatus is substantial for the venue: three datasets, objective metrics external to the training losses (WavLM/CAM++ SIM, Emotion2vec ACC_E, PDTW), 50-rater MOS with confidence intervals, ablations of both proposed components, LOSO cross-validation on AU corpora, and a promised code/dataset release. The single-token-per-frame FSQ design is a clean, cheap interface to LLM sequence modeling and the DualDPO extension is a natural idea. The main limitation on significance is attribution: the experiments as designed cannot yet isolate the facial modality's causal contribution, and the head-to-head with the strongest baseline (UniTalker) is confounded.

major comments (4)
  1. [§7.4 Ablation Study] §7.4 / Tables 4–5: there is no ablation that removes, masks, or shuffles the facial input at inference or training, so the paper never demonstrates that the facial modality is causally responsible for the gains over text/speech-only context. The ablations vary the visual encoder (FT-CLIP) and remove DualDPO (FT-base), but every condition still receives aligned faces. A face-shuffle or <IGNORE>-mask condition is essentially free — the <IGNORE> token mechanism already exists (§5.1.4) — and is load-bearing for the abstract's claim that facial cues make the output 'more natural, expressive, and better aligned with the conversational context.' If performance is nearly unchanged under shuffled faces (a plausible outcome given how much context the speech/text history carries in CSS), the framing of the paper changes materially. This experiment must be added.
  2. [§7.3] The sentence 'FT-base and UniTalker mainly differ in their visual tokenizer design' is not supported. FacialTalker is Qwen2.5-0.5B initialized from CosyVoice2 (170k h pretraining, §4.3.1) trained on ~1,724 h with a four-stage curriculum; UniTalker is a different architecture, initialization, and training corpus. The SIM 0.90→0.92 and PDTW 42.01→39.29 gaps on MultiDialog can plausibly come from backbone, initialization, or data scale rather than AU tokens vs. 128 landmarks. Either provide a matched-backbone comparison (e.g., swap landmark tokens into the same LLM pipeline) or explicitly downgrade this to a systems comparison and remove the tokenizer-attribution language.
  3. [§4.1.3 / §6.1 / Table 3] Two linked issues. (1) AUTokenizer's AU-supervision comes from CASME II (247 samples, 26 subjects) and DISFA (27 subjects) — small, posed/micro-expression corpora with non-overlapping AU label sets (§6.1) — yet the tokenizer is deployed on spontaneous podcast/interview faces in VSDD-1K. No measurement of token quality exists on the target domain (no AU labels there, no reconstruction or probe metric on VSDD-1K faces). The transfer assumption is asserted, not tested. (2) Table 3 shows AUTokenizer is second-tier as an AU classifier: 0.65 avg F1 on DISFA vs. VL-FAU 0.66, and 0.75 on CASME II vs. AULLM 0.81 and SSSNet-LED 0.79. That is acceptable for a tokenizer whose real job is compression, but the paper should then provide direct evidence that the single FSQ token (levels [8,8,8,8,8], i.e., ≤32,768 codes) retains affectively relevant information — e.g., codebook utilization, emotion probe
  4. [§4.3.1 DualDPO] The DualDPO gains, while consistent, are modest (e.g., MultiDialog ACC_E 0.76→0.79, MOS_N 4.16→4.22 over FT-base), and the rejected samples are the Stage-3 model's own generations. This self-preference construction is standard practice, but it carries a known risk of reward hacking toward the model's own distribution, and the paper gives no analysis of what the visual-side DPO actually changes (e.g., do chosen-vs-rejected facial token sequences differ in ways correlated with downstream emotion accuracy?). A brief analysis of preference-pair statistics and a sentence on why self-generated rejected samples are appropriate here would strengthen §4.3.1; if the visual DPO contributes little independently of the speech DPO, that should be reported.
minor comments (6)
  1. [Tables 4–5] Objective metrics in Tables 4–5 report no variance or CIs (only MOS has them). Given that some margins are small (SIM 0.91 vs 0.92 on VSDD-1K; FT-base vs full PDTW 41.52 vs 40.10), please report seed variance or a paired significance test over the 200-sample evaluation sets.
  2. [§6.3] The MOS raters are described as 'participants who speak English as a second language.' For MOS_N (naturalness) judgments of English speech, please justify this choice or report a subset with native-speaker raters; naturalness ratings are known to be rater-population-sensitive.
  3. [§7.5 / Fig. 5] Fig. 5's attention-visualization methodology is unspecified (which layer/heads are visualized for AUTokenizer? how is CLIP's 'facial expression' prompt attention computed and made comparable to a single-token FSQ model?). As presented the comparison is suggestive but not controlled.
  4. [§4.1.3] The number N of learnable AU queries (§4.1.3) is never stated, nor are training hyperparameters, inference cost, or latency of the full pipeline. Given that efficiency ('significantly reducing training and inference costs') is a stated motivation, a parameter/FLOPs/latency table versus UniTalker and the CLIP variant would be appropriate.
  5. [§3 / Table 1 / Eq. (2)] Typos/notation: §3 'should be not only natural' is missing a 'be'; Table 1's 'Modal' column renders as '( ,a,t)' with a missing modality symbol; Eq. (2) would benefit from a sentence of explanation; the alignment between 25 FPS face tokens and the speech token rate (and the <IGNORE> padding ratio in practice) should be stated explicitly since it affects context length.
  6. [§6.1] VSDD-1K test evaluation is an 8:1:1 random split of the same distribution used for training (§6.1). A held-out-source or held-out-speaker evaluation (VSDD-1K has >1498 speakers) would substantially strengthen the generalization claims and is feasible with the data already collected.

Circularity Check

0 steps flagged

No derivation-chain circularity: claims are empirical system results evaluated on external metrics, not predictions forced by their inputs.

full rationale

FacialTalker is an engineering/ML systems paper. Its load-bearing claims are comparative empirical results (Tables 3–5: AU F1, SIM, PDTW, ACC_E, MOS) on held-out dialogue data against external baselines and human raters. AUTokenizer is supervised by independent AU combination labels from CASME II/DISFA and scored with F1; speech quality uses third-party metrics (WavLM, CAM++, Emotion2vec, DNSMOS, UTMOS) and human MOS—none of which equal the training losses by construction. DualDPO (§4.3.1) uses Stage-3 model outputs as rejected pairs and ground-truth tokens as chosen pairs; that is standard preference optimization, not a fitted parameter renamed as a prediction, and it does not make the reported test metrics tautological. Self-citations to the authors’ prior CSS stack (GPT-Talker, ECSS, UniTalker, etc.) appear as related work and baselines, not as uniqueness theorems or ansatzes that force the present result. There is no self-definitional equation, no uniqueness import, and no renaming of a known closed-form result. Experimental gaps (missing face-shuffle/removal ablations, backbone confounds) are correctness/attribution issues, not circularity of the derivation chain. Score 0; steps empty.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 4 invented entities

The central CSS gains rest on domain assumptions that facial AUs are the right affect interface for speech style, that discrete multimodal next-token modeling on a small LLM is adequate context fusion, and that automatically mined interview/podcast video is a valid training distribution. Engineering free parameters (FSQ grid, AU query count, stage schedule, preference construction) and newly named modules (AUTokenizer, DualDPO, VSDD-1K) carry the method; none are derived from first principles.

free parameters (5)
  • FSQ levels [8,8,8,8,8] (single facial token codebook geometry) = [8,8,8,8,8]
    Quantization grid chosen for AUTokenizer; determines token capacity and is not derived from a uniqueness theorem.
  • Number of learnable AU queries N and 128-d fused AU feature size = N AU queries; 128-d concat (paper)
    Architectural widths for the AU decoder path; set by design to match AU supervision.
  • DualDPO preference-pair construction (Stage-3 self-samples as rejected)
    Chosen/rejected definition is a training design choice that shapes the post-SFT objective; not fixed by external theory.
  • Pipeline thresholds (VAD pause >5s; SNR-related cutoff 'below 4'; 25 FPS / 16 kHz) = pause>5s; threshold 4; 25fps/16kHz mono
    Hand-set cleaning rules that define what enters VSDD-1K and thus the learned distribution.
  • LLM backbone scale and init (Qwen2.5-0.5B from CosyVoice2 170k-h pretrain) = Qwen2.5-0.5B; CosyVoice2 init
    Capacity and initialization strongly condition all reported CSS metrics; selected, not predicted.
axioms (6)
  • domain assumption Combinations of facial Action Units are a sufficient supervisory signal for conversational facial affect relevant to empathetic speech style.
    Invoked in §2.2 and §4.1.3 to justify AUTokenizer over CLIP/landmarks; AU labels come from micro-expression datasets, not CSS end metrics.
  • ad hoc to paper One discrete token per frame is enough visual bandwidth for LLM autoregressive multimodal context modeling in CSS.
    Core efficiency claim of AUTokenizer vs 64×64 general visual tokenizers; supported empirically in-paper but assumed adequate a priori.
  • domain assumption User-only face streams (agent face omitted) match real deployment and still supply the affect needed for agent speech.
    Stated in task definition §3; shapes all context sequences.
  • domain assumption Automatically collected interview/podcast dialogues with ASD-filtered faces are a valid large-scale proxy for natural multimodal conversation.
    §5 pipeline and Table 1/2 comparisons; underpins claims that data scarcity is solved.
  • domain assumption Standard next-token CE plus DPO-style preference optimization improves joint visual–speech contextual understanding.
    §4.3 training strategy; imports LLM alignment practice into CSS without new theory.
  • domain assumption BPE text tokens, SenseVoice+FSQ speech tokens, and emotion category tokens form a compatible discrete interface with face tokens for a single LLM.
    §4.1 tokenization stack following CosyVoice2-style design.
invented entities (4)
  • AUTokenizer independent evidence
    purpose: Map each face frame to one discrete token under AU-combination supervision for LLM CSS.
    New module (ConvNeXt + multiscale fusion + AU queries + FSQ + AUPredictor); independent_evidence partial via AU F1 on public DISFA/CASME II, but CSS utility is in-paper only.
  • DualDPO no independent evidence
    purpose: Joint preference constraints on visual and speech token sequences after SFT.
    Named extension of DPO to two modalities; no external benchmark isolates DualDPO outside this paper’s ablations.
  • FacialTalker no independent evidence
    purpose: End-to-end facial-expression-aware conversational speech synthesis system on an LLM backbone.
    System-level assembly of tokenizer, LLM, CFM synthesizer, and DualDPO; evaluated only in this work’s tables.
  • VSDD-1K independent evidence
    purpose: Large-scale synchronized video–speech dialogue resource (~1033 h) for multimodal CSS training.
    New corpus from automated internet pipeline; intended public release under CC BY-NC-SA 4.0 provides a falsifiable external handle if released as described.

pith-pipeline@v1.2.0-grok45-kimik3 · 23809 in / 4736 out tokens · 104919 ms · 2026-07-31T15:04:17.222486+00:00 · methodology

0 comments
read the original abstract

Conversational Speech Synthesis is a fundamental component of human-computer interaction, aiming to generate contextually appropriate, expressive, and empathetic speech. However, facial expressions encode subtle and rich affective cues that are crucial for empathetic speech interaction, whereas existing approaches often overlook this important modality. In addition, the lack of large-scale natural conversational datasets with both speech and visual modalities also limits the development of visual affect understanding in conversational settings.To address these limitations, we propose FacialTalker, a facial-expression-aware CSS framework built upon a large language model backbone. To efficiently encode facial expressions, we propose AUTokenizer, a single-codebook visual tokenizer that discretizes each frame-level facial expression into a compact token, trained with supervision from combinations of facial Action Units. We further introduce a dual direct preference optimization (DualDPO) strategy, which extends the DPO by jointly imposing preference constraints on both visual and speech token sequences, to enhance the model's understanding of facial expressions and speech semantics in multimodal conversational contexts. Moreover, we construct VSDD-1K, a large-scale multimodal dialogue dataset collected through a fully automated pipeline from real-world Internet conversations, comprising over 1,033 hours of synchronized speaker videos and speech, with more than 85\% of frames containing valid faces. Extensive objective and subjective experiments demonstrate that FacialTalker consistently outperforms strong baselines in facial-expression perception and speech synthesis quality, generating speech that is more natural, expressive, and better aligned with the conversational context. The results also validate the effectiveness of our training strategy and dataset construction pipeline.

Figures

Figures reproduced from arXiv: 2607.24430 by Haizhou Li, Rui Liu, Shuwei He, Yifan Hu.

Figure 1
Figure 1. Figure 1: (a) Previous CSS methods lack visual perception [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: The overview of FacialTalker. The framework consists of four main parts: (1) [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: The overall architecture of AUTokenizer. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: The construction pipeline of the VSDD-1K visual [PITH_FULL_IMAGE:figures/full_fig_p005_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Visualization of Facial Expression Attention Be [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

77 extracted references · 17 linked inside Pith

  1. [1]

    Inclusion AI, Bowen Ma, Cheng Zou, Canxiang Yan, Chunxiang Jin, Chunjie Shen, Chenyu Lian, Dandan Zheng, Fudong Wang, Furong Xu, et al. 2025. Ming-flash- omni: A sparse, unified architecture for multimodal perception and generation. arXiv preprint arXiv:2510.24821(2025)

  2. [2]

    Keyu An, Qian Chen, Chong Deng, Zhihao Du, Changfeng Gao, Zhifu Gao, Yue Gu, Ting He, Hangrui Hu, Kai Hu, et al. 2024. Funaudiollm: Voice understanding and generation foundation models for natural interaction between humans and llms.arXiv preprint arXiv:2407.04051(2024)

  3. [3]

    2022.Face-to-face dialogue: Theory, research, and applica- tions

    Janet Beavin Bavelas. 2022.Face-to-face dialogue: Theory, research, and applica- tions. Oxford University Press

  4. [4]

    Fadi Boutros, Meiling Fang, Marcel Klemt, Biying Fu, and Naser Damer. 2023. CR- FIQA: face image quality assessment by learning sample relative classifiability. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 5836–5845

  5. [5]

    Carlos Busso, Murtaza Bulut, Chi-Chun Lee, Abe Kazemzadeh, Emily Mower, Samuel Kim, Jeannette N Chang, Sungbok Lee, and Shrikanth S Narayanan. 2008. IEMOCAP: Interactive emotional dyadic motion capture database.Language resources and evaluation42, 4 (2008), 335–359

  6. [6]

    Rajdeep Chatterjee, Saptarshi Mazumdar, R Simon Sherratt, Rohit Halder, Tanmoy Maitra, and Debasis Giri. 2021. Real-time speech emotion analysis for smart home assistants.IEEE Transactions on Consumer Electronics67, 1 (2021), 68–76

  7. [7]

    Sanyuan Chen, Shujie Liu, Long Zhou, Yanqing Liu, Xu Tan, Jinyu Li, Sheng Zhao, Yao Qian, and Furu Wei. 2024. Vall-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers.arXiv preprint arXiv:2406.05370 (2024)

  8. [8]

    Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, et al . 2022. Wavlm: Large-scale self-supervised pre-training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing16, 6 (2022), 1505–1518

  9. [9]

    Yirong Chen, Weiquan Fan, Xiaofen Xing, Jianxin Pang, Minlie Huang, Wenjing Han, Qianfeng Tie, and Xiangmin Xu. 2022. CPED: A large-scale Chinese per- sonalized and emotional dialogue dataset for conversational AI.arXiv preprint arXiv:2205.14727(2022)

  10. [10]

    Alexandre Défossez, Laurent Mazaré, Manu Orsini, Amélie Royer, Patrick Pérez, Hervé Jégou, Edouard Grave, and Neil Zeghidour. 2024. Moshi: a speech-text foundation model for real-time dialogue.arXiv preprint arXiv:2410.00037(2024)

  11. [11]

    Zhihao Du, Yuxuan Wang, Qian Chen, Xian Shi, Xiang Lv, Tianyu Zhao, Zhifu Gao, Yexin Yang, Changfeng Gao, Hui Wang, et al . 2024. Cosyvoice 2: Scal- able streaming speech synthesis with large language models.arXiv preprint arXiv:2412.10117(2024)

  12. [12]

    Patrick Esser, Robin Rombach, and Bjorn Ommer. 2021. Taming transformers for high-resolution image synthesis. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 12873–12883

  13. [13]

    Hao Fei, Han Zhang, Bin Wang, Lizi Liao, Qian Liu, and Erik Cambria. 2024. Empathyear: An open-source avatar multimodal empathetic chatbot. InProceed- ings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations). 61–71

  14. [14]

    Xuri Ge, Junchen Fu, Fuhai Chen, Shan An, Nicu Sebe, and Joemon M Jose

  15. [15]

    Haohan Guo, Shaofei Zhang, Frank K Soong, Lei He, and Lei Xie. 2021. Conversa- tional end-to-end tts for voice agents. In2021 IEEE Spoken Language Technology Workshop (SLT). IEEE, 403–409

  16. [16]

    Guohong Hu, Xing Lan, Hanyu Jiang, Jiayi Lyu, and Jian Xue. 2024. Towards unified facial action unit recognition framework by large language models.arXiv preprint arXiv:2409.08444(2024)

  17. [17]

    Hangrui Hu, Xinfa Zhu, Ting He, Dake Guo, Bin Zhang, Xiong Wang, Zhifang Guo, Ziyue Jiang, Hongkun Hao, Zishan Guo, et al. 2026. Qwen3-TTS Technical Report.arXiv preprint arXiv:2601.15621(2026)

  18. [18]

    Yifan Hu, Rui Liu, Guanglai Gao, and Haizhou Li. 2024. FCTalker: Fine and Coarse Grained Context Modeling for Expressive Conversational Speech Synthesis. In 14th IEEE International Symposium on Chinese Spoken Language Processing, ISCSLP 2024, Beijing, China, November 7-10, 2024, Yanmin Qian, Qin Jin, Zhijian Ou, Zhenhua Ling, Zhiyong Wu, Ya Li, Lei Xie, a...

  19. [19]

    Yifan Hu, Rui Liu, Yi Ren, Xiang Yin, and Haizhou Li. 2025. UniTalker: Conver- sational Speech-Visual Synthesis. InProceedings of the 33rd ACM International Conference on Multimedia. 10248–10257

  20. [20]

    Byungho Jo, Donghyeon Cho, In Kyu Park, and Sungeun Hong. 2023. IFQA: Interpretable face quality assessment. InProceedings of the IEEE/CVF winter conference on applications of computer vision. 3444–3453

  21. [21]

    Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.Advances in neural information processing systems33 (2020), 17022–17033

  22. [22]

    Keon Lee, Kyumin Park, and Daeyoung Kim. 2023. Dailytalk: Spoken dialogue dataset for conversational text-to-speech. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5

  23. [23]

    Jidong Leng and Qiang Yan. 2025. Eijl: Popularity prediction of social media advertisements based on multimodal emotional interaction and joint learning. Data Intelligence7, 4 (2025), 1129–1146

  24. [24]

    Jingbei Li, Yi Meng, Chenyi Li, Zhiyong Wu, Helen Meng, Chao Weng, and Dan Su. 2022. Enhancing speaking styles in conversational text-to-speech synthesis with graph-based multi-modal context modeling. InICASSP 2022-2022 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 7917–7921

  25. [25]

    Jingbei Li, Yi Meng, Xixin Wu, Zhiyong Wu, Jia Jia, Helen Meng, Qiao Tian, Yup- ing Wang, and Yuxuan Wang. 2022. Inferring speaking styles from multi-modal conversational context by multi-scale relational graph convolutional networks. In Proceedings of the 30th ACM International Conference on Multimedia. 5811–5820

  26. [26]

    Yifan Li, Anh Dao, Wentao Bao, Zhen Tan, Tianlong Chen, Huan Liu, and Yu Kong

  27. [27]

    Yante Li, Xiaohua Huang, and Guoying Zhao. 2021. Micro-expression action unit detection with spatial and channel attention.Neurocomputing436 (2021), 221–231

  28. [28]

    InEuropean Conference on Computer Vision

    Facial affective behavior analysis with instruction tuning. InEuropean Conference on Computer Vision. Springer, 165–186

  29. [29]

    Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017. Dailydialog: A manually labelled multi-turn dialogue dataset. InProceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 986–995

  30. [30]

    Yante Li, Wei Peng, and Guoying Zhao. 2021. Micro-expression action unit de- tection with dual-view attentive similarity-preserving knowledge distillation. In 2021 16th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2021). IEEE, 01–08

  31. [31]

    Guan-Ting Lin, Prashanth Gurunath Shivakumar, Ankur Gandhe, Chao- Han Huck Yang, Yile Gu, Shalini Ghosh, Andreas Stolcke, Hung-yi Lee, and Ivan Bulyko. 2024. Paralinguistics-enhanced large language modeling of spoken dialogue. InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 10316–10320

  32. [32]

    Junhua Liao, Haihan Duan, Kanghui Feng, Wanbing Zhao, Yanbing Yang, Liangyin Chen, and Yanru Chen. 2025. Lr-asd: Lightweight and robust network for active speaker detection.International Journal of Computer Vision133, 7 (2025), 4749– 4769

  33. [33]

    Rui Liu, Yifan Hu, Yi Ren, Xiang Yin, and Haizhou Li. 2024. Emotion rendering for conversational speech synthesis with heterogeneous graph-based context modeling. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 18698–18706

  34. [34]

    Yaron Lipman, Ricky TQ Chen, Heli Ben-Hamu, Maximilian Nickel, and Matt Le

  35. [35]

    Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. 2022. A convnet for the 2020s. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 11976–11986

  36. [36]

    Zhishu Liu, Kaishen Yuan, Bo Zhao, Yong Xu, and Zitong Yu. 2025. Au-llm: Micro-expression action unit detection via enhanced llm-based feature fusion. In Chinese Conference on Biometric Recognition. Springer, 355–365

  37. [37]

    Rui Liu, Yifan Hu, Yi Ren, Xiang Yin, and Haizhou Li. 2024. Generative expressive conversational speech synthesis. InProceedings of the 32nd ACM International Conference on Multimedia. 4187–4196

  38. [38]

    Ziyang Ma, Zhisheng Zheng, Jiaxin Ye, Jinchao Li, Zhifu Gao, Shiliang Zhang, and Xie Chen. 2024. emotion2vec: Self-supervised pre-training for speech emotion MM ’26, November 10–14, 2026, Rio de Janeiro, Brazil Yifan Hu, Shuwei He, Rui Liu, and Haizhou Li. representation. InFindings of the Association for Computational Linguistics: ACL

  39. [39]

    Brais Martinez, Michel F Valstar, Bihan Jiang, and Maja Pantic. 2017. Automatic analysis of facial actions: A survey.IEEE transactions on affective computing10, 3 (2017), 325–347

  40. [40]

    Cheng Luo, Siyang Song, Weicheng Xie, Linlin Shen, and Hatice Gunes. 2022. Learning multi-dimensional edge feature-based au relation graph for facial action unit recognition.arXiv preprint arXiv:2205.01782(2022)

  41. [41]

    Fabian Mentzer, David Minnen, Eirikur Agustsson, and Michael Tschannen. 2023. Finite scalar quantization: Vq-vae made simple.arXiv preprint arXiv:2309.15505 (2023)

  42. [42]

    Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. 2016. V-net: Fully convolutional neural networks for volumetric medical image segmentation. In 2016 fourth international conference on 3D vision (3DV). Ieee, 565–571

  43. [43]

    S Mohammad Mavadati, Mohammad H Mahoor, Kevin Bartlett, Philip Trinh, and Jeffrey F Cohn. 2013. Disfa: A spontaneous facial action intensity database.IEEE Transactions on Affective Computing4, 2 (2013), 151–160

  44. [44]

    Se Park, Chae Kim, Hyeongseop Rha, Minsu Kim, Joanna Hong, Jeonghun Yeo, and Yong Ro. 2024. Let’s go real talk: Spoken dialogue model for face-to-face conversation. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 16334–16348

  45. [45]

    Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2019. Meld: A multimodal multi-party dataset for emotion recognition in conversations. InProceedings of the 57th annual meeting of the association for computational linguistics. 527–536

  46. [46]

    Jinhui Pang, Xinyun Yang, Xiaoyao Qiu, Zixuan Wang, and Taisheng Huang

  47. [47]

    MMAF: Masked Multi-modal Attention Fusion to Reduce Bias of Visual Features for Named Entity Recognition.Data Intelligence6, 4 (2024), 1114–1133

  48. [48]

    Yi Ren, Chenxu Hu, Xu Tan, Tao Qin, Sheng Zhao, Zhou Zhao, and Tie-Yan Liu. 2020. Fastspeech 2: Fast and high-quality end-to-end text to speech.arXiv preprint arXiv:2006.04558(2020)

  49. [49]

    Tal Ridnik, Emanuel Ben-Baruch, Nadav Zamir, Asaf Noy, Itamar Friedman, Matan Protter, and Lihi Zelnik-Manor. 2021. Asymmetric loss for multi-label classification. InProceedings of the IEEE/CVF international conference on computer vision. 82–91

  50. [50]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learning. PmLR, 8748–8763

  51. [51]

    Chandan KA Reddy, Vishak Gopal, and Ross Cutler. 2021. DNSMOS: A non- intrusive perceptual objective speech quality metric to evaluate noise suppressors. InICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 6493–6497

  52. [52]

    Yuki Saito, Yuto Nishimura, Shinnosuke Takamichi, Kentaro Tachibana, and Hiroshi Saruwatari. 2022. STUDIES: Corpus of Japanese empathetic dialogue speech towards friendly voice agent.arXiv preprint arXiv:2203.14757(2022)

  53. [53]

    Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016. Neural machine translation of rare words with subword units. InProceedings of the 54th annual meeting of the association for computational linguistics (volume 1: long papers). 1715–1725

  54. [54]

    Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari. 2022. Utmos: Utokyo-sarulab system for voicemos challenge 2022.arXiv preprint arXiv:2204.02152(2022)

  55. [55]

    Yuki Saito, Eiji Iimori, Shinnosuke Takamichi, Kentaro Tachibana, and Hiroshi Saruwatari. 2023. CALLS: Japanese empathetic dialogue speech corpus of complaint handling and attentive listening in customer center.arXiv preprint arXiv:2305.13713(2023)

  56. [56]

    Junke Wang, Yi Jiang, Zehuan Yuan, Binyue Peng, Zuxuan Wu, and Yu-Gang Jiang. 2024. Omnitokenizer: A joint image-video tokenizer for visual generation. Advances in Neural Information Processing Systems37 (2024), 28281–28295

  57. [57]

    Ning Wang, Frank Broz, Alessandro Di Nuovo, Tony Belpaeme, and Angelo Cangelosi. 2016. A user-centric design of service robots speech interface for the elderly. InRecent Advances in Nonlinear Speech Processing. Springer, 275–283

  58. [58]

    Meituan LongCat Team, Bairui Wang, Bin Xiao, Bo Zhang, Bolin Rong, Borun Chen, Chang Wan, Chao Zhang, Chen Huang, Chen Chen, et al. 2025. Longcat- flash-omni technical report.arXiv preprint arXiv:2511.00279(2025)

  59. [59]

    Tuomas Varanka, Wei Peng, and Guoying Zhao. 2023. Learnable eulerian dy- namics for micro-expression action unit detection. InScandinavian Conference on Image Analysis. Springer, 385–400

  60. [60]

    Zhifei Xie and Changqiao Wu. 2024. Mini-omni2: Towards open-source gpt- 4o with vision, speech and duplex capabilities.arXiv preprint arXiv:2410.11190 (2024)

  61. [61]

    Jin Xu, Zhifang Guo, Hangrui Hu, Yunfei Chu, Xiong Wang, Jinzheng He, Yuxuan Wang, Xian Shi, Ting He, Xinfa Zhu, et al. 2025. Qwen3-omni technical report. arXiv preprint arXiv:2509.17765(2025)

  62. [62]

    Boyong Wu, Chao Yan, Chen Hu, Cheng Yi, Chengli Feng, Fei Tian, Feiyu Shen, Gang Yu, Haoyang Zhang, Jingbei Li, et al. 2025. Step-audio 2 technical report. arXiv preprint arXiv:2507.16632(2025)

  63. [63]

    Lei Wu, Jiyong Xue, Wenbo Li, Kan Wang, Xiang Zhang, and Gang Guo. 2022. Toward decreasing the driving risk: speech-based driver’s anger regulation in smart cockpit.IEEE Journal of Radio Frequency Identification6 (2022), 764–768

  64. [64]

    Wen-Jing Yan, Xiaobai Li, Su-Jing Wang, Guoying Zhao, Yong-Jin Liu, Yu-Hsin Chen, and Xiaolan Fu. 2014. CASME II: An improved spontaneous micro- expression database and the baseline evaluation.PloS one9, 1 (2014), e86041

  65. [65]

    Zehui Yang, Yifan Chen, Lei Luo, Runyan Yang, Lingxuan Ye, Gaofeng Cheng, Ji Xu, Yaohui Jin, Qingqing Zhang, Pengyuan Zhang, et al. 2022. Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational (RAMC) Speech Dataset. InProc. Interspeech 2022. 1736–1740

  66. [66]

    Hongfei Xue, Yuhao Liang, Bingshen Mu, Shiliang Zhang, Mengzhe Chen, Qian Chen, and Lei Xie. 2024. E-chat: Emotion-sensitive spoken dialogue system with large language models. In2024 IEEE 14th International Symposium on Chinese Spoken Language Processing (ISCSLP). IEEE, 586–590

  67. [67]

    Jinlong Xue, Yayue Deng, Fengping Wang, Ya Li, Yingming Gao, Jianhua Tao, Jianqing Sun, and Jiaen Liang. 2023. M 2-ctts: End-to-end multi-scale multi-modal conversational text-to-speech synthesis. InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1–5

  68. [68]

    Fan Zhang and Lin Chai. 2024. A review of research on micro-expression recog- nition algorithms based on deep learning.Neural Computing and Applications36, 29 (2024), 17787–17828

  69. [69]

    Han Zhang, Zixiang Meng, Meng Luo, Hong Han, Lizi Liao, Erik Cambria, and Hao Fei. 2025. Towards multimodal empathetic response generation: A rich text-speech-vision avatar-based benchmark. InProceedings of the ACM on Web Conference 2025. 2872–2881

  70. [70]

    Rohola Zandie, Mohammad H Mahoor, Julia Madsen, and Eshrat S Emamian

  71. [71]

    Siyi Zhou, Yiquan Zhou, Yi He, Xun Zhou, Jinchao Wang, Wei Deng, and Jingchen Shu. 2026. Indextts2: A breakthrough in emotionally expressive and duration- controlled auto-regressive zero-shot text-to-speech. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 40. 35139–35148

  72. [72]

    Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, and Xipeng Qiu. 2023. Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities. InFindings of the Association for Computational Linguistics: EMNLP 2023. 15757–15773

  73. [75]

    Jinming Zhao, Tenggan Zhang, Jingwen Hu, Yuchen Liu, Qin Jin, Xinchao Wang, and Haizhou Li. 2022. M3ED: Multi-modal multi-scene multi-label emotional dialogue database. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 5699–5710

  74. [77]

    Yixuan Zhou, Guoyang Zeng, Xin Liu, Xiang Li, Renjie Yu, Ziyang Wang, Runchuan Ye, Weiyue Sun, Jiancheng Gui, Kehan Li, et al . 2025. Voxcpm: Tokenizer-free TTS for context-aware speech generation and true-to-life voice cloning.arXiv preprint arXiv:2509.24650(2025)

  75. [2021]

    RyanSpeech: A Corpus for Conversational Text-to-Speech Synthesis. In Proc. Interspeech 2021. 2751–2755

  76. [2022]

    Flow matching for generative modeling.arXiv preprint arXiv:2210.02747 (2022)

  77. [2024]

    InProceedings of the 32nd ACM International Conference on Multimedia

    Towards end-to-end explainable facial action unit recognition via vision- language joint learning. InProceedings of the 32nd ACM International Conference on Multimedia. 8189–8198