Open-source audio-language models perform far below humans on a new 600-item benchmark of temporal reasoning in sound, and their accuracy does not track a proposed perturbation-based uncertainty measure.
Test-time Uncertainty Measure We propose to measure the uncertainty in the decision making for a test sample using data perturbations
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
Open-source audio-language models perform far below humans on a new 600-item benchmark of temporal reasoning in sound, and their accuracy does not track a proposed perturbation-based uncertainty measure.