Pith. sign in

REVIEW 4 cited by

PAM: Prompting Audio-Language Models for Audio Quality Assessment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.00282 v1 pith:SOCIOHNI submitted 2024-02-01 eess.AS cs.SD

classification eess.AScs.SD
keywords audioqualityhumanlisteningmetricmetricsscorestasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While audio quality is a key performance metric for various audio processing tasks, including generative modeling, its objective measurement remains a challenge. Audio-Language Models (ALMs) are pre-trained on audio-text pairs that may contain information about audio quality, the presence of artifacts, or noise. Given an audio input and a text prompt related to quality, an ALM can be used to calculate a similarity score between the two. Here, we exploit this capability and introduce PAM, a no-reference metric for assessing audio quality for different audio processing tasks. Contrary to other "reference-free" metrics, PAM does not require computing embeddings on a reference dataset nor training a task-specific model on a costly set of human listening scores. We extensively evaluate the reliability of PAM against established metrics and human listening scores on four tasks: text-to-audio (TTA), text-to-music generation (TTM), text-to-speech (TTS), and deep noise suppression (DNS). We perform multiple ablation studies with controlled distortions, in-the-wild setups, and prompt choices. Our evaluation shows that PAM correlates well with existing metrics and human listening scores. These results demonstrate the potential of ALMs for computing a general-purpose audio quality metric.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation

    cs.CL 2025-07 conditional novelty 7.0 of 10

    With prompt engineering (audio concatenation plus in-context examples), large audio models rank speech synthesis systems in line with human preferences, reaching up to 0.91 Spearman correlation.

  2. AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences

    cs.SD 2025-07 conditional novelty 6.0 of 10

    AudioBERTScore, a training-free metric combining max-norm and p-norm similarity of audio embeddings, correlates more strongly with human subjective scores for text-to-audio synthesis than conventional metrics.

  3. A Survey on Evaluation Metrics for Music Generation

    cs.SD 2025-08 conditional novelty 5.0 of 10

    A taxonomy and critical review of evaluation metrics for music generation, identifying gaps such as weak correlation with human perception and lack of standardization.

  4. A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A unified taxonomy and comparative meta-evaluation of automatic evaluation methods across text, vision, and speech generation, concluding that LLM-based evaluators dominate current practice.

Pith tools