Pith. sign in

REVIEW 3 cited by

CTAL: Pre-training Cross-modal Transformer for Audio-and-Language Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.00181 v1 pith:PBSCI72O submitted 2021-09-01 cs.SD cs.AI

classification cs.SDcs.AI
keywords audio-and-languagecross-modalmodeltasksaudio-languagectalfine-tuningfusion
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Existing audio-language task-specific predictive approaches focus on building complicated late-fusion mechanisms. However, these models are facing challenges of overfitting with limited labels and low model generalization abilities. In this paper, we present a Cross-modal Transformer for Audio-and-Language, i.e., CTAL, which aims to learn the intra-modality and inter-modality connections between audio and language through two proxy tasks on a large amount of audio-and-language pairs: masked language modeling and masked cross-modal acoustic modeling. After fine-tuning our pre-trained model on multiple downstream audio-and-language tasks, we observe significant improvements across various tasks, such as, emotion classification, sentiment analysis, and speaker verification. On this basis, we further propose a specially-designed fusion mechanism that can be used in fine-tuning phase, which allows our pre-trained model to achieve better performance. Lastly, we demonstrate detailed ablation studies to prove that both our novel cross-modality fusion component and audio-language pre-training methods significantly contribute to the promising results.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion

    cs.SD 2025-06 conditional novelty 6.0 of 10

    The paper contributes a 1.2M-caption and 6M-QA multimodal audio dataset generated by an LLM that fuses speech, music, sound, and visual cues, and reports downstream gains on retrieval and understanding.

  2. DiVR: incorporating context from diverse VR scenes for human trajectory prediction

    cs.AI 2024-11 conditional novelty 6.0 of 10

    A cross-modal transformer using heterogeneous scene graphs predicts VR user trajectories more accurately than gaze- and point-cloud-only baselines on the CREATTIVE3D dataset.

  3. Audio-Language Models for Audio-Centric Tasks: A Systematic Survey

    cs.SD 2025-01 conditional novelty 5.0 of 10

    A systematic survey that categorizes audio-language models by architecture, training objective, and application, covering speech, music, and general audio.

Pith tools