Pith. sign in

REVIEW 20 cited by

GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11768 v1 pith:4UZILHES submitted 2024-06-17 cs.SD cs.AIcs.CLeess.AS

classification cs.SDcs.AIcs.CLeess.AS
keywords audiocomplexgamareasoningunderstandingabilitiesaudio-languagecapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Perceiving and understanding non-speech sounds and non-verbal speech is essential to making decisions that help us interact with our surroundings. In this paper, we propose GAMA, a novel General-purpose Large Audio-Language Model (LALM) with Advanced Audio Understanding and Complex Reasoning Abilities. We build GAMA by integrating an LLM with multiple types of audio representations, including features from a custom Audio Q-Former, a multi-layer aggregator that aggregates features from multiple layers of an audio encoder. We fine-tune GAMA on a large-scale audio-language dataset, which augments it with audio understanding capabilities. Next, we propose CompA-R (Instruction-Tuning for Complex Audio Reasoning), a synthetically generated instruction-tuning (IT) dataset with instructions that require the model to perform complex reasoning on the input audio. We instruction-tune GAMA with CompA-R to endow it with complex reasoning abilities, where we further add a soft prompt as input with high-level semantic evidence by leveraging event tags of the input audio. Finally, we also propose CompA-R-test, a human-labeled evaluation dataset for evaluating the capabilities of LALMs on open-ended audio question-answering that requires complex reasoning. Through automated and expert human evaluations, we show that GAMA outperforms all other LALMs in literature on diverse audio understanding tasks by margins of 1%-84%. Further, GAMA IT-ed on CompA-R proves to be superior in its complex reasoning and instruction following capabilities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering

    cs.IR 2026-04 unverdicted novelty 7.0 of 10

    Jamendo-MT-QA is a new dataset and benchmark for multi-track comparative music question answering, constructed via an LLM-assisted pipeline from Creative Commons Jamendo tracks and used to evaluate audio-language models.

  2. See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models

    cs.CV 2025-12 unverdicted novelty 7.0 of 10

    AV-SpeakerBench is a new speaker-centered benchmark showing that top multimodal models still struggle with fine-grained audiovisual speech understanding, with Gemini 2.5 Pro leading but open models lagging on fusion.

  3. MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models

    cs.CL 2026-04 conditional novelty 6.5 of 10

    Current Omni LLMs extract modality-specific cues yet fail to integrate them for accurate multicontext safety judgments, performing better on physical than social/illegal risks.

  4. Empowering Long-form Omni-modal Understanding with Robust Audio Perception

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Decoupled audio-visual caption and CoT-QA datasets plus two-stage fine-tuning measurably strengthen auditory perception and cross-modal reasoning in a 7B omni-modal LLM.

  5. Unlocking In-Context Learning in Audio-Language Models from Decentralized Medical Audio

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    FSC uses unsupervised clustering for pseudo-label episodes and a three-stage federated pipeline to achieve 71.6% accuracy in 2-way 2-shot in-context diagnosis of respiratory and cardiac audio conditions.

  6. Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding

    cs.SD 2026-04 unverdicted novelty 6.0 of 10

    HyPeR is a hybrid perception-reasoning framework that uses a new hierarchical PAQA dataset and PAUSE tokens to improve large audio language models' handling of multi-speaker and ambiguous audio.

  7. Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models

    cs.SD 2026-04 unverdicted novelty 6.0 of 10

    Temporal Contrastive Decoding mitigates temporal smoothing bias in unified large audio-language models by contrasting logits from original and blurred audio inputs during decoding, yielding consistent gains on MMAU an...

  8. EvA: An Evidence-First Audio Understanding Paradigm for LALMs

    cs.SD 2026-03 conditional novelty 6.0 of 10

    Preserving multi-scale non-speech evidence via hierarchical aggregation and non-compressive time-aligned fusion measurably lifts LALM perception more than reasoning.

  9. The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents

    cs.CR 2025-07 unverdicted novelty 6.0 of 10

    A multi-agent audio-language model framework can automatically profile private attributes, such as age, health, and income, directly from general audio recordings.

  10. Attacker's Noise Can Manipulate Your Audio-based LLM in the Real World

    cs.CR 2025-07 conditional novelty 6.0 of 10

    Adversarial audio noise, optimized with audio augmentations, can trigger and distort the behavior of audio-based LLMs both digitally and when played through the air.

  11. LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    LLaDA-V is a diffusion-based multimodal large language model that reaches competitive or state-of-the-art results on visual instruction tasks while using a non-autoregressive architecture.

  12. FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation

    eess.AS 2026-01 conditional novelty 5.0 of 10

    A hierarchical Q-Former compresses speech to about 1.67 tokens/sec, enabling hour-long audio processing with near-linear memory scaling and competitive benchmark scores.

  13. Improving Audio Event Recognition with Consistency Regularization

    cs.SD 2025-09 conditional novelty 5.0 of 10

    Consistency regularization improves audio event recognition on AudioSet by about 2 mAP, both in supervised and semi-supervised settings.

  14. From Sound to Sight: Towards AI-authored Music Videos

    cs.SD 2025-08 conditional novelty 5.0 of 10

    This paper presents two off-the-shelf model pipelines (CLAP or LALM, an LLM, and a text-to-video model) for generating music videos from arbitrary songs, validated by a preliminary five-participant user study with mod...

  15. MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses

    eess.AS 2025-07 reject novelty 5.0 of 10

    MMW combines a Mamba-based Mix Block, a Frame Diarization Mamba layer, and multi-scale GRPO to reduce side-talk interference in Whisper ASR, reporting WER as low as 3.71%.

  16. Kimi-Audio Technical Report

    eess.AS 2025-04 unverdicted novelty 5.0 of 10

    Kimi-Audio is an open-source audio foundation model that achieves state-of-the-art results on speech recognition, audio understanding, question answering, and conversation after pre-training on more than 13 million ho...

  17. TinyGiantALM: A Compact Audio-Language Model for Intent-Aware Reasoning under Resource Constraints

    cs.SD 2026-06 unverdicted novelty 4.0 of 10

    TinyGiantALM, a compact 1.5B audio-language model with instruction-aware refinement, achieves 46.4% zero-shot accuracy on MMAR and outperforms models up to 8x larger in mixed-modality tasks.

  18. Frame-Aligned Fusion of Canary and WavLM for Non-Intrusive Intelligibility Prediction of Hearing-Aid-Processed Speech

    eess.AS 2026-05 unverdicted novelty 4.0 of 10

    Frame-aligned fusion of Canary and WavLM encoders, with WavLM temporally prepared via learnable strided convolution, outperforms other fusion strategies and reaches Eval RMSE 24.96 and Corr 0.796 on non-intrusive inte...

  19. Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents

    cs.CL 2026-05 unverdicted novelty 4.0 of 10

    Audio language models are benchmarked on five semantic and paralinguistic reasoning tasks to reveal limitations in handling spoken audio evidence, accent variation, and domain shifts.

  20. BoSS: Beyond-Semantic Speech

    cs.SD 2025-07 conditional novelty 4.0 of 10

    Current spoken-language models perform poorly on a new five-task evaluation of beyond-semantic speech signals, including dialect, emotion, age, and non-verbal cues.

Pith tools