VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track

Bohan Li; Jian Gao; Jing Peng; Kai Yu; Shuai Fan; Tao Liu; Wenming Tu; Xie Chen; Yanru Huo; Yixuan Wang

arxiv: 2606.07264 · v1 · pith:BIYHPIZJnew · submitted 2026-06-05 · 📡 eess.AS

VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track

Wenming Tu , Jian Gao , Yanru Huo , Yixuan Wang , Jing Peng , Bohan Li , Ziyang Ma , Tao Liu

show 4 more authors

Shuai Fan Kai Yu Xie Chen Zilong Zheng

This is my paper

classification 📡 eess.AS

keywords agentaudioreasoningvisatrackinferenceinterspeechmulti-modal

0 comments

read the original abstract

Audio reasoning requires multi-step, evidence-grounded inference over temporally dynamic and acoustically mixed signals, exceeding conventional perception tasks such as ASR or captioning. We present VISA, our submission to the Interspeech 2026 Audio Reasoning Challenge (Agent Track), evaluated via the MMAR Rubrics for correctness and reasoning quality. Under a "LALM as a Tool" paradigm, VISA strengthens large audio language models with auxiliary multi-modal evidence while avoiding heavy orchestration. The system integrates three components: multi-modal feature extraction for complementary audio and acoustic-visual clues, model-voting inference with consistency checking for stable predictions, and fine-grained category-aware routing to resolve disagreements and select rubric-aligned reasoning chains. On the official Agent Track leaderboard, VISA ranks 2nd overall with a 66.23% Rubrics score. It also achieves 77.40% Accuracy, the highest among all systems listed across both the Single Model and Agent tracks.

This paper has not been read by Pith yet.

VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track

discussion (0)