Pith. sign in

REVIEW 2 cited by

AlignCap: Aligning Speech Emotion Captioning to Human Preferences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.19134 v1 pith:NPAXUZKL submitted 2024-10-24 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords speechhumanaligncapcaptioningemotionaligningalignmentemotional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speech Emotion Captioning (SEC) has gradually become an active research task. The emotional content conveyed through human speech are often complex, and classifying them into fixed categories may not be enough to fully capture speech emotions. Describing speech emotions through natural language may be a more effective approach. However, existing SEC methods often produce hallucinations and lose generalization on unseen speech. To overcome these problems, we propose AlignCap, which Aligning Speech Emotion Captioning to Human Preferences based on large language model (LLM) with two properties: 1) Speech-Text Alignment, which minimizing the divergence between the LLM's response prediction distributions for speech and text inputs using knowledge distillation (KD) Regularization. 2) Human Preference Alignment, where we design Preference Optimization (PO) Regularization to eliminate factuality and faithfulness hallucinations. We also extract emotional clues as a prompt for enriching fine-grained information under KD-Regularization. Experiments demonstrate that AlignCap presents stronger performance to other state-of-the-art methods on Zero-shot SEC task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations

    cs.MM 2025-05 conditional novelty 6.0 of 10

    EmotionTalk provides 19,250 utterances from 744 Chinese dyadic dialogues with emotion, sentiment, and speaking-style caption annotations.

  2. EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Using Rank-N-Contrast loss on valence-arousal rankings instead of symmetric cross-entropy improves ordinal consistency and cross-modal alignment for emotional speech and text.

Pith tools