REVIEW 5 major objections 8 minor 47 references
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos
T0 review · 5 major / 8 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read V-CASS claims that synthesizing speech whose emotional tone matches a video's visual cues improves viewers' understanding and engagement, and reports that 74.68% of participants preferred it over neutral speech.
desk verdict A plausible system paper with a genuinely useful formatting study, but the headline preference result is confounded by expressiveness and the statistics are sloppy. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the visual-to-vocal para-linguistic mapping: a structured body of knowledge, elicited from professional voice actors and visual-emotion perception research, that connects visual properties such as color, lighting, line, and shadow to recommended vocal expressions such as tone, pitch, pace, and volume. This mapping bridges pixels to prosody. In V-CASS it is embedded in the middle stage, where a large language model first classifies the visual attributes reported by the vision-language model, then maps those attributes to emotional states, and finally fuses the resulting emotional information with the transcript into a structured instruction that the expressive speech model consumes. The same mapping is what gives the knowledge-infused model its measured edge: its emotional descriptions matched human descriptions with an average embedding similarity of 0.70, versus 0.62 for the same model without the expert knowledge.
What would settle it
Recruit a fresh group of participants with no prior exposure to the study, blind them to the hypothesis, and compare V-CASS's expressive speech against neutral text-to-speech on both preference ratings and objective comprehension questions about the video's emotional intent; if naive participants show no preference or no comprehension advantage, the central claim is not supported.
Extended reading notes
Core claim
The central claim is that emotional alignment between speech and visual content is load-bearing for comprehension: viewers treat the speaker's tone as evidence about what the video means, so neutral narration leaves interpretation ambiguous and mismatched narration pushes it in the wrong direction. The formatting study quantified this with professional voice recordings: intent-to-perception consistency was 64.00% for neutral speech, 79.95% for vision-context-aligned speech, and 58.00% for emotionally contradictory speech. The paper then claims V-CASS reproduces the aligned condition automatically: a vision-language model extracts visual attributes such as lighting, color, and scene mood; a large language model, infused with expert visual-to-vocal mapping rules and prompted to reason step by step, converts those attributes together with the transcript into structured expressive instructions; and an expressive instruction-to-speech model produces the final speech. In a forced-preference user study, 74.68% of participants preferred V-CASS's output over neutral text-to-speech, and in a five-person case study blind and low-vision users found the expressive version more useful and immersive as audio description.
Load-bearing premise
The study assumes that the 30 participants, who had already completed the formatting study and were aware of its purpose, were not biased by that prior exposure when they later preferred V-CASS's expressive speech in the user study.
Editorial extensions
If this is right
- If the central claim holds, automatic video commentary systems should condition on visual para-linguistic cues, not just transcribed facts, to avoid ambiguous or distorted interpretations.
- The formatting study shows misaligned emotion is actively harmful, so systems that default to neutral or guessed emotion risk misleading viewers; deriving emotional delivery from visual context is the safer design.
- Expressive instruction-based speech synthesis can be driven automatically by a vision-language model and a knowledge-infused language model, removing the need for hand-authored style prompts per clip.
- V-CASS-style outputs can serve as richer audio descriptions for blind and low-vision users, conveying mood and atmosphere in addition to factual content.
- Infusing expert mapping knowledge measurably improves agreement with human emotional judgments, so the mapping itself is a reusable asset independent of the underlying language model.
Reading between the lines
- A stronger test than self-reported preference would measure objective comprehension: ask naive viewers, who have never seen the study or its aims, to answer questions about a video's narrative and emotional intent after hearing each speech version. The 74.68% preference could partly reflect the contrast with deliberately flat synthetic speech rather than a genuine understanding gain.
- Because V-CASS separates visual emotion reading, visual-to-vocal translation, and expressive synthesis into distinct stages, each can be swapped or improved independently; a small specialized model distilled from the expert mapping could replace the general-purpose language model, making the pipeline cheaper and easier to audit.
- The finding that contradictory speech actively distorts interpretation suggests a design rule for uncertain AI narration: when the visual emotion is unclear, a neutral delivery may be less harmful than a confident guess, since a wrong emotion is worse than none.
- The five-person blind and low-vision case study motivates a larger accessibility trial that measures task performance, such as answering questions about scene tone, with V-CASS audio descriptions versus standard factual audio descriptions.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes V-CASS, a three-stage pipeline for vision-context-aware expressive speech synthesis in automatic video commentary: Stage I uses Gemini to extract visual para-linguistic cues from video keyframes; Stage II uses GPT-4o with Chain-of-Thought and expert knowledge elicited from professional voice actors to translate those cues into vocal expressiveness instructions; Stage III uses VoxInstruct to synthesize expressive speech. The authors first report a formatting study with 30 participants showing that vision-context-aligned speech yields higher consistency (79.95%) than neutral speech (64.00%) and that emotionally contradictory speech increases misinterpretation. They then report user studies in which 74.68% of participants prefer V-CASS speech over neutral TTS, and a case study with five blind and low-vision participants indicating preference for context-aligned audio descriptions. The paper claims V-CASS enhances emotional and attitudinal resonance, user audio-visual understanding, and engagement, and has potential for accessibility applications.
Significance. If fully supported, this work addresses a real gap in automatic video commentary: most systems convey only semantic content and ignore visual para-linguistic cues. The proposed pipeline is practical, combining off-the-shelf foundation models, and the formatting study usefully demonstrates that speech emotion can dominate or even distort visual interpretation. Strengths include the expert-knowledge elicitation from three voice actors, the effort to integrate that knowledge via Chain-of-Thought prompting, the provision of a demo/appendix, and the accessibility motivation for blind and low-vision users. However, the evaluation has a load-bearing confound: the preference experiment compares V-CASS only against neutral TTS, not against an expressive but vision-unaware baseline, so the 74.68% preference cannot be attributed specifically to vision-context alignment. Several arithmetic and reporting inconsistencies further weaken the quantitative claims.
major comments (5)
- [Section V, second experiment] The central claim—that vision-context alignment drives the benefit—is not supported by the experiment as designed. The preference study compares V-CASS expressive speech only against 'neutral speech generated by a basic TTS.' Because Section III, Table I shows that any emotionally charged speech, including emotionally contradictory speech, dominates interpretation relative to neutral speech, a preference for expressive speech over neutral speech is expected and does not isolate the contribution of vision-context awareness. The experiment lacks an expressive non-context-aware baseline, such as VoxInstruct driven by the same instruction template but with emotion labels derived from the transcript or randomly assigned, or a prompt that omits Stage I visual cues. Consequently, the 74.68% preference is confounded by general expressiveness and cannot support the specific claim that V-CASS's vision-context awareness enhances understanding.
- [Section V, '74.68%' statistic] The headline statistic is arithmetically inconsistent with the stated sample size. 74.68% of 30 participants is 22.404, and 74.68% of 300 video-level judgments (30 participants × 10 videos) is 224.04, both of which are non-integers. Additionally, the text says 'participants preferred' although the experiment collects per-video preference, making the unit of analysis ambiguous. Please report the raw counts and clarify whether the percentage is over participants, video-level judgments, or some aggregation.
- [Section III, Table I vs text] The text states that neutral speech yields a 'correct understanding rate of 65%' and that emotionally contradictory speech yields 'a significantly lower correct understanding rate of 60%', but Table I reports 64.00% and 58.00%, respectively. These discrepancies between the narrative and the table make it difficult to trust the reported effects; please reconcile all numbers.
- [Section III, participant demographics] The participant description says 30 participants, with 14 males and 15 females, which sums to 29. The reported percentages (89% aged 18–25, 76% students, 40% arts and design, 36% science, 23% human-computer interaction) yield non-integer counts for N=30, and the discipline percentages sum to 99%. This suggests a counting or rounding error that should be corrected, as it affects the credibility of the study's reporting.
- [Section V, first experiment and Table II] No statistical significance testing is reported for the comparison between the knowledge-infused LLM and the non-knowledge-infused LLM. The average Sentence-BERT similarities (0.70 vs 0.62 across 10 videos) could be within noise, and details on the exact prompts, temperature, and number of runs are absent. Additionally, reusing the same 30 participants from the formatting study risks priming: their prior exposure to aligned and contradictory speech examples and their knowledge of the study aims could bias preferences toward expressive speech. The assertion that this reuse 'reduces the learning process' does not address the potential direction of the bias.
minor comments (8)
- [Throughout] There are several typos, e.g., 'stduy' in Section VIII and 'V edio' in the Table II header.
- [Equation (2)] Equation (2) uses the set-intersection symbol (Cv ∩ K_expert) where Cv and K_expert are textual descriptions; this notation is not meaningful and should be replaced with, for instance, a concatenation or an explicit prompt-construction step.
- [Section III, formatting study] The paper does not describe how the 'original video intent' labels were assigned in the formatting study; please specify the labeling procedure and the number of labelers involved.
- [Section III, formatting study] The statement that 'each video sample was rated by at least ten participants' should be clarified: with 30 participants and three speech conditions, it is unclear how ratings were distributed across the experimental cells.
- [Section VI, BLV case study] The BLV case study is qualitative, with only five participants and no quantitative measures; the current wording ('participants were more accurate') overstates what is actually reported.
- [Section IV, prompts] The appendix/demo link is appreciated; the authors should consider including the exact prompts used for Gemini and GPT-4o in the appendix to improve reproducibility.
- [Table I] The column grouping in Table I is easy to misread; consider adding explicit subheaders for 'Positive Intent' and 'Negative Intent' for the PPT/PNT and NPT/NNT pairs.
- [Throughout] The phrase 'formatting study' appears throughout; if 'formative study' is intended, it should be corrected for consistency.
Circularity Check
No circular derivation chain: V-CASS is an empirical system integration; the preference result is confounded by expressiveness and participant priming, but no prediction is forced by construction.
full rationale
The paper does not contain a derivation chain whose predictions reduce to its inputs. V-CASS is an integration of external foundation models (Gemini, GPT-4o, VoxInstruct) with a prompt engineered from interview-derived visual-to-vocal mapping knowledge; no parameters are fitted and no predicted quantity is defined in terms of its own outcome. The formatting study independently establishes that emotionally charged speech, whether aligned or contradictory, dominates neutral speech in shaping interpretation, and this empirical result is an input to the system design rather than a tautology. The main user study compares V-CASS expressive speech only against neutral TTS, and it re-invites the same 30 participants from the formatting study, whose prior exposure to the study aims is acknowledged in the paper; these are experimental confounds (generic expressiveness and priming) that weaken the attribution of the 74.68% preference to vision-context alignment, but they are not circularity in the derivation. The only self-citation, VoxInstruct [21], is used as the Stage III synthesizer and shares authors with the present paper; however, VoxInstruct is a separately published, independently evaluated system rather than a uniqueness theorem invoked to forbid alternatives, so the citation is not load-bearing. Score 2 reflects this minor self-citation and the methodological proximity between the formatting study and the user study, not a circular derivation.
Assumptions & free parameters
assumptions (3)
- ad hoc to paper Visual-to-vocal para-linguistic mapping knowledge K_expert, elicited from three voice actors, is a valid and generalizable mapping that an LLM can apply.
- domain assumption The selected videos and transcripts are neutral and balanced enough to isolate the effect of speech emotion on user understanding.
- domain assumption Pretrained foundation models (Gemini, GPT-4o, VoxInstruct) operate as described, with no failure modes that would break the pipeline.
Cite this review
Pith. "Pith review of V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos." pith.science (2026). https://pith.science/paper/RYVG7OPS
@misc{pith2026250616716,
author = {Pith},
title = {Pith review of: V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos},
year = {2026},
howpublished = {\url{https://pith.science/paper/RYVG7OPS}},
note = {Machine review of arXiv:2506.16716}
}
read the original abstract
Automatic video commentary systems are widely used on multimedia social media platforms to extract factual information about video content. However, current systems may overlook essential para-linguistic cues, including emotion and attitude, which are critical for fully conveying the meaning of visual content. The absence of these cues can limit user understanding or, in some cases, distort the video's original intent. Expressive speech effectively conveys these cues and enhances the user's comprehension of videos. Building on these insights, this paper explores the usage of vision-context-aware expressive speech in enhancing users' understanding of videos in video commentary systems. Firstly, our formatting study indicates that semantic-only speech can lead to ambiguity, and misaligned emotions between speech and visuals may distort content interpretation. To address this, we propose a method called vision-context-aware speech synthesis (V-CASS). It analyzes para-linguistic cues from visuals using a vision-language model and leverages a knowledge-infused language model to guide the expressive speech model in generating context-aligned speech. User studies show that V-CASS enhances emotional and attitudinal resonance, as well as user audio-visual understanding and engagement, with 74.68% of participants preferring the system. Finally, we explore the potential of our method in helping blind and low-vision users navigate web videos, improving universal accessibility.
Figures
Reference graph
Works this paper leans on
-
[1]
Video Summarization Using Deep Neural Networks: A Survey,
E. Apostolidis, E. Adamantidou, A. I. Metsai, V . Mezaris, and I. Pa- tras, “Video Summarization Using Deep Neural Networks: A Survey,” Proceedings of the IEEE, vol. 109, no. 11, pp. 1838–1863, 2021
work page 2021
-
[2]
Exploring Video Captioning Techniques: A Comprehensive Survey on Deep Learning Methods,
S. Islam, A. Dash, A. Seum, A. H. Raj, T. Hossain, and F. M. Shah, “Exploring Video Captioning Techniques: A Comprehensive Survey on Deep Learning Methods,”SN Computer Science, vol. 2, no. 2, p. 120, Apr. 2021
work page 2021
-
[3]
S. Amirian, K. Rasheed, T. R. Taha, and H. R. Arabnia, “Automatic Image and Video Caption Generation With Deep Learning: A Concise Review and Algorithmic Overlap,”IEEE Access, vol. 8, pp. 218 386– 218 400, 2020
work page 2020
-
[4]
T. Liu and X. Yuan, “Paralinguistic and spectral feature extraction for speech emotion classification using machine learning techniques,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2023, no. 1, p. 23, May 2023
work page 2023
-
[5]
V ocal communication of emotion: A review of research paradigms,
K. Scherer, “V ocal communication of emotion: A review of research paradigms,”Speech Communication, vol. 40, no. 1-2, pp. 227–256, Apr. 2003
work page 2003
-
[6]
Does speech rate influence intertemporal decisions? an experimental investigation,
J. I. Chen, T.-S. He, and H.-Y . Liao, “Does speech rate influence intertemporal decisions? an experimental investigation,”PLOS ONE, vol. 17, no. 2, p. e0264356, Feb. 2022
work page 2022
-
[7]
Rhythmic and speech rate effects in the perception of durational cues,
J. Steffman, “Rhythmic and speech rate effects in the perception of durational cues,”Attention, Perception, & Psychophysics, vol. 83, no. 8, pp. 3162–3182, Nov. 2021
work page 2021
-
[8]
M. M. Bradley and P. J. Lang, “Emotion and Motivation,” inHandbook of Psychophysiology, 3rd ed., J. T. Cacioppo, L. G. Tassinary, and G. Berntson, Eds. Cambridge: Cambridge University Press, 2007, pp. 581–607
work page 2007
Show all 47 references
-
[9]
Language and Emotion: Introduction to the Special Issue,
K. A. Lindquist, “Language and Emotion: Introduction to the Special Issue,”Affective Science, vol. 2, no. 2, pp. 91–98, Jun. 2021
2021
-
[10]
Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies,
Y . Gao, L. Fischer, A. Lintner, and S. Ebling, “Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies,” 2024
2024
-
[11]
Audio Description in the UK: What works, what doesn’t, and understanding the need for personalising access,
M. Lopez, G. Kearney, and K. Hofst ¨adter, “Audio Description in the UK: What works, what doesn’t, and understanding the need for personalising access,”British journal of visual impairment, vol. 36, no. 3, pp. 274– 291, 2018, publisher: SAGE Publications Sage UK: London, England
2018
-
[12]
Audio description: The visual made verbal,
J. Snyder, “Audio description: The visual made verbal,”International Congress Series, vol. 1282, pp. 935–939, Sep. 2005
2005
-
[13]
Ambient Lights Influence Perception and Decision-Making,
S. Song and S. Yamada, “Ambient Lights Influence Perception and Decision-Making,”Frontiers in Psychology, vol. 9, p. 2685, Jan. 2019
2019
-
[14]
Kobayasi,Colorist: a practical handbook for personal and profes- sional use
S. Kobayasi,Colorist: a practical handbook for personal and profes- sional use. Kodansha International: Tokyo u.a, 1998
1998
-
[15]
Bordwell and K
D. Bordwell and K. Thompson,Film art: an introduction, 10th ed. New York, NY: McGraw-Hill, 2013
2013
-
[16]
Tacotron: Towards end-to-end speech synthesis,
Y . Wang, R. Skerry-Ryan, D. Stanton, Y . Wu, R. J. Weiss, N. Jaitly, Z. Yang, Y . Xiao, Z. Chen, S. Bengio, and others, “Tacotron: Towards end-to-end speech synthesis,”arXiv preprint arXiv:1703.10135, 2017
2017 arXiv
-
[17]
WaveNet: A Gener- ative Model for Raw Audio,
A. v. d. Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “WaveNet: A Gener- ative Model for Raw Audio,” 2016
2016
-
[18]
Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,
Y . A. Li, C. Han, V . Raghavan, G. Mischler, and N. Mesgarani, “Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models,”Advances in Neural Information Processing Systems, vol. 36, pp. 19 594–19 621, 2023
2023
-
[19]
Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. G ¨olge, and M. A. Ponti, “Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,” inInternational conference on machine learning. PMLR, 2022, pp. 2709–2720
2022
-
[20]
InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt,
D. Yang, S. Liu, R. Huang, C. Weng, and H. Meng, “InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt,” 2023
2023
-
[21]
V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,
Y . Zhou, X. Qin, Z. Jin, S. Zhou, S. Lei, S. Zhou, Z. Wu, and J. Jia, “V oxinstruct: Expressive human instruction-to-speech generation with unified multilingual codec language modelling,” inProceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 554–563
2024
-
[22]
TextrolSpeech: A Text Style Control Speech Corpus with Codec Language Text-to-Speech Models,
S. Ji, J. Zuo, M. Fang, Z. Jiang, F. Chen, X. Duan, B. Huai, and Z. Zhao, “TextrolSpeech: A Text Style Control Speech Corpus with Codec Language Text-to-Speech Models,” inICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024...
2024
-
[23]
PromptTTS 2: Describing and Generating V oices with Text Prompt,
Y . Leng, Z. Guo, K. Shen, X. Tan, Z. Ju, Y . Liu, Y . Liu, D. Yang, L. Zhang, K. Song, L. He, X.-Y . Li, S. Zhao, T. Qin, and J. Bian, “PromptTTS 2: Describing and Generating V oices with Text Prompt,” 2023
2023
-
[24]
What Does Your Face Sound Like? 3D Face Shape towards V oice,
Z. Yang, Z. Wu, Y . Shan, and J. Jia, “What Does Your Face Sound Like? 3D Face Shape towards V oice,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 11, pp. 13 905–13 913, Jun. 2023
2023
-
[25]
Multimodal Machine Learning: A Survey and Taxonomy,
T. Baltru ˇsaitis, C. Ahuja, and L.-P. Morency, “Multimodal Machine Learning: A Survey and Taxonomy,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 41, no. 2, pp. 423–443, 2019
2019
-
[26]
Multimodal Transformer for Unaligned Multimodal Language Sequences,
Y .-H. H. Tsai, S. Bai, P. P. Liang, J. Z. Kolter, L.-P. Morency, and R. Salakhutdinov, “Multimodal Transformer for Unaligned Multimodal Language Sequences,” inProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, A. Korhonen, D. Traum, and L...
2019
-
[27]
MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech Synthesis,
W. Guan, Y . Li, T. Li, H. Huang, F. Wang, J. Lin, L. Huang, L. Li, and Q. Hong, “MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech Synthesis,”Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 16, pp. 18 117–18 125, Mar. 2024
2024
-
[28]
Face2Speech: Towards Multi-Speaker Text-to-Speech Synthesis Using an Embedding Vector Predicted from a Face Image
S. Goto, K. Onishi, Y . Saito, K. Tachibana, and K. Mori, “Face2Speech: Towards Multi-Speaker Text-to-Speech Synthesis Using an Embedding Vector Predicted from a Face Image.” inINTERSPEECH, 2020, pp. 1321–1325
2020
-
[29]
Imaginary V oice: Face-Styled Diffusion Model for Text-to-Speech,
J. Lee, J. Son Chung, and S.-W. Chung, “Imaginary V oice: Face-Styled Diffusion Model for Text-to-Speech,” inICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5
2023
-
[30]
Face-based V oice Conversion: Learning the V oice behind a Face,
H.-H. Lu, S.-E. Weng, Y .-F. Yen, H.-H. Shuai, and W.-H. Cheng, “Face-based V oice Conversion: Learning the V oice behind a Face,” in Proceedings of the 29th ACM International Conference on Multimedia, ser. MM ’21. New York, NY , USA: Association for Computing Machinery, 2021,...
2021
-
[31]
EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model,
D. Li, X. Liu, B. Xing, B. Xia, Y . Zong, B. Wen, and H. K ¨alvi¨ainen, “EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model,” 2024
2024
-
[32]
Prompt-to-Prompt Image Editing with Cross Attention Control,
A. Hertz, R. Mokady, J. Tenenbaum, K. Aberman, Y . Pritch, and D. Cohen-Or, “Prompt-to-Prompt Image Editing with Cross Attention Control,” 2022
2022
-
[33]
On the Opportunities and Risks of Foundation Models,
B. et al., “On the Opportunities and Risks of Foundation Models,” 2021
2021
-
[34]
Learning Transferable Visual Models From Natural Language Super- vision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning Transferable Visual Models From Natural Language Super- vision,” 2021
2021
-
[35]
ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision,
W. Kim, B. Son, and I. Kim, “ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision,” 2021
2021
-
[36]
Video (language) modeling: a baseline for generative models of natural videos,
M. Ranzato, A. Szlam, J. Bruna, M. Mathieu, R. Collobert, and S. Chopra, “Video (language) modeling: a baseline for generative models of natural videos,” 2014
2014
-
[37]
ESCoT: Towards Interpretable Emotional Support Dialogue Systems,
T. Zhang, X. Zhang, J. Zhao, L. Zhou, and Q. Jin, “ESCoT: Towards Interpretable Emotional Support Dialogue Systems,” 2024
2024
-
[38]
Chain-of-thought prompting elicits reasoning in large language models,
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhou, and others, “Chain-of-thought prompting elicits reasoning in large language models,”Advances in neural information processing systems, vol. 35, pp. 24 824–24 837, 2022
2022
-
[39]
AutoFoley: Artificial Synthesis of Syn- chronized Sound Tracks for Silent Videos With Deep Learning,
S. Ghose and J. J. Prevost, “AutoFoley: Artificial Synthesis of Syn- chronized Sound Tracks for Silent Videos With Deep Learning,”IEEE Transactions on Multimedia, vol. 23, pp. 1895–1907, 2021
1907
-
[40]
MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generation,
L. Ruan, Y . Ma, H. Yang, H. He, B. Liu, J. Fu, N. J. Yuan, Q. Jin, and B. Guo, “MM-Diffusion: Learning Multi-Modal Diffusion Models for Joint Audio and Video Generation,” 2022
2022
-
[41]
V ocoder-Based Speech Synthesis from Silent Videos,
D. Michelsanti, O. Slizovskaia, G. Haro, E. G ´omez, Z.-H. Tan, and J. Jensen, “V ocoder-Based Speech Synthesis from Silent Videos,” 2020
2020
-
[42]
Sonicvisionlm: Playing sound with vision language models,
Z. Xie, S. Yu, Q. He, and M. Li, “Sonicvisionlm: Playing sound with vision language models,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26 866–26 875
2024
-
[43]
‘The problem-centred expert interview’. Combining qual- itative interviewing approaches for investigating implicit expert knowl- edge,
S. D ¨oringer, “‘The problem-centred expert interview’. Combining qual- itative interviewing approaches for investigating implicit expert knowl- edge,”International Journal of Social Research Methodology, vol. 24, no. 3, pp. 265–278, May 2021
2021
-
[44]
Pleasure-arousal-dominance: A general framework for describing and measuring individual differences in Temperament,
A. Mehrabian, “Pleasure-arousal-dominance: A general framework for describing and measuring individual differences in Temperament,”Cur- rent Psychology, vol. 14, no. 4, pp. 261–292, Dec. 1996
1996
-
[45]
Pleasure, Arousal, Dominance: Mehrabian and Russell revisited,
I. Bakker, T. Van Der V oordt, P. Vink, and J. De Boon, “Pleasure, Arousal, Dominance: Mehrabian and Russell revisited,”Current Psy- chology, vol. 33, no. 3, pp. 405–421, Sep. 2014
2014
-
[46]
Gemini: A Family of Highly Capable Multimodal Models,
Gemini Team, “Gemini: A Family of Highly Capable Multimodal Models,” 2023
2023
-
[47]
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,
N. Reimers and I. Gurevych, “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,” 2019
2019
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.