LTBM merges similar audio tokens under a temporal locality constraint and shows task-dependent advantages over global merging for captioning versus multiple-choice understanding tasks in experiments with Qwen2-Audio and Audio Flamingo 3.
InProceed- ings of the IEEE/CVF International Conference on Computer Vision
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Locality Matters for Training-Free Audio Token Compression in Audio-Language Models
LTBM merges similar audio tokens under a temporal locality constraint and shows task-dependent advantages over global merging for captioning versus multiple-choice understanding tasks in experiments with Qwen2-Audio and Audio Flamingo 3.