REVIEW 5 major objections 4 minor 86 references
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
T0 review · 5 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper introduces OmniAVS, a benchmark of 61,095 omnimodal referring expressions for audio-visual segmentation, and OISA, a multimodal LLM that reaches 41.1% J&F on it, beating prior best by 5.0 points.
desk verdict New omnimodal referring dataset is a real contribution, but the missing audio-ablation control leaves its central audio-understanding claim unproven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanisms are Audio-Visual Interleaving and Query Propagation. Audio-Visual Interleaving divides the audio token sequence into clips and places each clip immediately after its corresponding frame's vision tokens, forming a synchronized sequence of the form $\{v_1, a_1, v_2, a_2, \ldots, v_N, a_N\}$ without adding parameters. Query Propagation updates the [SEG] token frame-by-frame in the mask decoder, rather than using one fixed token for all frames, so the query tracks object motion and avoids identity switches. The [SEG] token, produced by the MLLM from the interleaved multimodal context, is fed into the mask head for segmentation.
What would settle it
Re-annotate a random 5% of OmniAVS test expressions with an independent annotation team and measure inter-annotator agreement on the referred-object masks; if agreement falls near the 5.0-point J&F gap between OISA-1B and LISA-13B, the ranking would not be trustworthy. Alternatively, rerun OISA-1B with the audio and video tokens interleaved in random order instead of the frame-aligned order; if J&F does not drop substantially, the claimed benefit of Audio-Visual Interleaving is not load-bearing.
Extended reading notes
Core claim
The central claim is that a new benchmark, OmniAVS, can push referring audio-visual segmentation from surface-level acoustic attributes to semantic content understanding and reasoning, and that a multimodal LLM-based model can be adapted to this harder task. OISA-1B accomplishes this by interleaving audio and visual tokens for temporal synchronization and by propagating a single segmentation query across frames, achieving state-of-the-art results on OmniAVS and strong transfer to related referring and reasoning segmentation tasks.
Load-bearing premise
The benchmark's reliability rests on the assumption that its 61,095 expressions and the associated mask annotations are clean and unambiguous; the paper reports no inter-annotator agreement or quality checks, so if those labels are noisy, the reported rankings and difficulty comparisons could change.
Editorial extensions
If this is right
- If OmniAVS becomes a standard benchmark, referring segmentation evaluation will include audio-content reasoning and explanation quality, not just acoustic event detection.
- The eight-expression-type interface could push development of models that accept arbitrary combinations of text, speech, sound, and image as referring input, closer to human interaction.
- Audio-Visual Interleaving and Query Propagation are architecture-agnostic enough to be incorporated into other MLLM-based segmentation systems, potentially improving video-level referring segmentation broadly.
- The explanations provided for reasoning expressions enable quantifying a model's interpretability, a dimension absent from prior referring audio-visual benchmarks.
Reading between the lines
- Because speech expressions are generated by converting text via TTS, there may be a gap between these utterances and natural spontaneous speech; a follow-up could test whether OISA's gains persist with human-spoken expressions.
- The full-temporal mask convention, inherited from MeViS, may be ill-matched to expressions that are only valid in part of a video; a temporally-gated evaluation variant could reveal whether models locate objects during the relevant segment.
- The 5.0-point gain over LISA-13B may partly stem from the audio-text alignment pretraining stage rather than the interleaving mechanism itself; ablating that stage would clarify which contribution is decisive.
- The benchmark's emphasis on sound-content reasoning connects naturally to audio-visual question answering, so a joint model fine-tuned on both OmniAVS and A-VQA may improve both tasks through shared reasoning supervision.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes OmniAVS, a new referring audio-visual segmentation dataset with 2,104 videos and 61,095 referring expressions spanning eight modality combinations (text/speech with sound and/or image). The authors argue that existing RAVS datasets such as Ref-AVS rely on surface-level acoustic properties, whereas OmniAVS expressions require understanding audio content and performing reasoning, such as inferring illness from coughing sounds. The paper also introduces OISA-1B, an MLLM-based baseline with two technical components: audio-visual interleaving for temporal alignment and query propagation for mask decoding. Experiments report state-of-the-art results on OmniAVS (41.1% J&F, outperforming LISA-13B by 5.0 points) and competitive results on Ref-AVS, referring image/video segmentation, and ReVOS.
Significance. If the dataset annotations are reliable, OmniAVS is a useful resource that moves referring audio-visual segmentation from sound-presence cues toward audio-content understanding and reasoning. The OISA design choices, audio-visual interleaving and query propagation, are simple, parameter-free, and show consistent gains in the in-model ablations. The evaluation is performed on a held-out test split and does not rely on fitted constants, so the main benchmark claim is not circular. However, the absence of an audio-ablated control leaves the core claim that OmniAVS expressions demand audio-content understanding empirically unsupported, and the benchmark's reliability is not yet established because annotation quality is not measured. For these reasons the contribution is promising but not fully established.
major comments (5)
- [§5.2–5.3 (Tables 3 and 5)] No experiment removes or masks the video audio stream. All fusion ablations in Table 3 vary only how real audio tokens are combined, and all benchmark conditions in Table 5 feed audio into the model. Because many OmniAVS expressions are semantically redundant with text and vision (e.g., 'Who is most likely to be sick?' is inferable from visible coughing, and 'The dog warning' names the sound in the text), a vision-language model without audio understanding could plausibly achieve much of the reported 41.1% J&F. An audio-ablated control (e.g., replacing audio tokens with silence or zeros, evaluated for both OISA and adapted LISA) is needed to support the paper's central claim that OmniAVS demands audio-content understanding beyond sound-presence detection.
- [§3.2] The reliability of OmniAVS as a benchmark depends on annotation quality, but Section 3.2 reports no inter-annotator agreement, no double-annotation rate, and no quality-control metric for the 61,095 expressions or the 206k mask labels. The expression rules (e.g., 'emphasize the sound's content') leave room for subjective judgment, and annotator use of SAM2 assistance does not by itself validate mask correctness. Please report agreement statistics (e.g., mask IoU between annotators and expression-validity agreement) on a sample, and state how ambiguous or failing annotations were resolved.
- [§5.3 (Table 5)] The LISA baseline is enhanced only with the same audio-text alignment as OISA, while OISA additionally uses audio-visual interleaving and query propagation. The 5.0-point advantage over LISA-13B therefore conflates the proposed architectural components with the ability to reason about audio content. To support the claim that OISA outperforms existing methods on omnimodal reasoning, the comparison should give LISA the same interleaving and query-propagation benefits, or isolate each component's contribution on the OmniAVS test set.
- [§5.3–5.4 (Tables 5–8)] All reported metrics are single-run point estimates with no variance or significance testing. This matters for claims such as the 5.0-point gain over LISA-13B and the 0.2-point improvement over VISA on ReVOS; without error bars or multiple seeds, small differences may not be reproducible. Please report standard deviations over multiple runs, or at least provide evidence that the main conclusions are stable under different random seeds.
- [§5.4 (Table 6)] The discussion of the Ref-AVS Null split is speculative: the claim that EEMC's 0.7% S score 'likely' reflects overfitting to audio patterns rather than genuine null understanding is not tested. Because OISA-1B's S=9.8% is substantially worse on this split, the paper should either provide an error analysis supporting the overfitting explanation or qualify the claim of 'greatly surpassing' EEMC on Ref-AVS.
minor comments (4)
- [§5.3] The sentence 'splits VII (text+speech+image) and VIII (text+sound+image)' mislabels the expression types: according to Section 3.2, type VII is text+sound+image and type VIII is speech+sound+image.
- [§5.1] The terms 'dense frames' and 'sparse frames' are used without definition; please clarify how the dense and sparse frame subsets are selected and how they differ during training.
- [§5 (Evaluation Metrics)] For no-target expressions, J&F is set to 1 when the prediction is empty and 0 otherwise; this gives full credit for predicting an empty mask, which could inflate scores. Please report the proportion of no-target expressions in OmniAVS and the sensitivity of the overall J&F to this convention.
- [§4.2] Audio tokens are divided uniformly across the N sampled frames, but OmniAVS videos have annotation frame rates of 3–15 FPS; please clarify whether the audio segmentation uses actual timestamps or a simplified uniform division, and discuss the effect on audio-visual alignment for variable-FPS videos.
Circularity Check
No significant circularity: benchmark and method are evaluated on a held-out test split and against external benchmarks.
full rationale
The paper has two contributions: the OmniAVS dataset and the OISA method. The central numerical claims (OISA-1B reaching 41.1% J&F on OmniAVS, outperforming LISA-13B by 5.0 points, and 58.0% J&F on Ref-AVS) are produced by training on training splits and evaluating on held-out test splits; no parameter is fitted to the test set and no reported metric is a renamed training objective. The audio-content emphasis of OmniAVS is built through annotation rules in Section 3.2, not derived from OISA's outputs, and OISA's results do not define the benchmark's content. Citations to prior work by the same authors, such as MeViS for the full-temporal mask protocol or MOVE, MOSEv2, and MMT-Bench as related baselines, are methodological or bibliographic and do not carry the derivation of any reported result. The absence of an audio-ablated control in Section 5 is a valid experimental concern about whether OISA's scores prove audio-content understanding, but it is not circularity: a missing control does not make the result equivalent to its input. The paper is also self-contained against external benchmarks (Ref-AVS, MeViS, RefCOCO, ReVOS), further supporting the independence of the method's assessment. No load-bearing step reduces, by construction or by self-citation, to the paper's own inputs.
Assumptions & free parameters
free parameters (1)
- Frame sampling counts =
10 training frames, 32 inference frames, 4 dense frames
assumptions (4)
- domain assumption Uniform frame sampling and interleaved audio clips preserve audio-visual synchronization.
- domain assumption Full-temporal masks are valid when an object only partially matches an expression.
- domain assumption The [SEG] token output by the LLM can be used directly as a mask query.
- ad hoc to paper Videos selected for informative audio and complex scenes represent the target task.
Cite this review
Pith. "Pith review of Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation." pith.science (2026). https://pith.science/paper/ODETY2BB
@misc{pith2026250722886,
author = {Pith},
title = {Pith review of: Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation},
year = {2026},
howpublished = {\url{https://pith.science/paper/ODETY2BB}},
note = {Machine review of arXiv:2507.22886}
}
read the original abstract
Referring audio-visual segmentation (RAVS) has recently seen significant advancements, yet challenges remain in integrating multimodal information and deeply understanding and reasoning about audiovisual content. To extend the boundaries of RAVS and facilitate future research in this field, we propose Omnimodal Referring Audio-Visual Segmentation (OmniAVS), a new dataset containing 2,104 videos and 61,095 multimodal referring expressions. OmniAVS stands out with three key innovations: (1) 8 types of multimodal expressions that flexibly combine text, speech, sound, and visual cues; (2) an emphasis on understanding audio content beyond just detecting their presence; and (3) the inclusion of complex reasoning and world knowledge in expressions. Furthermore, we introduce Omnimodal Instructed Segmentation Assistant (OISA), to address the challenges of multimodal reasoning and fine-grained understanding of audiovisual content in OmniAVS. OISA uses MLLM to comprehend complex cues and perform reasoning-based segmentation. Extensive experiments show that OISA outperforms existing methods on OmniAVS and achieves competitive results on other related tasks.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Qwen Technical Report
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, et al. Qwen Technical Report. arXiv, 2023. 6
2023
-
[2]
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities. arXiv preprint arXiv:2308.12966 ,
-
[3]
One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
Zechen Bai, Tong He, Haiyang Mei, Pichao W ANG, Ziteng Gao, Joya Chen, liulei, Zheng Zhang, and Mike Zheng Shou. One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos. In Adv. Neural Inform. Process. Syst., 2024. 2, 3, 6, 7, 8
2024
-
[4]
METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments
Satanjeev Banerjee and Alon Lavie. METEOR: An Au- tomatic Metric for MT Evaluation with Improved Correla- tion with Human Judgments. In Assoc. Comput. Linguist. Worksh., 2005. 6
2005
-
[5]
End-to-End Referring Video Object Segmentation with Mul- timodal Transformers
Adam Botach, Evgenii Zheltonozhskii, and Chaim Baskin. End-to-End Referring Video Object Segmentation with Mul- timodal Transformers. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 8
2022
-
[6]
Auditory Scene Analysis: The Perceptual Organization of Sound
Albert S Bregman. Auditory Scene Analysis: The Perceptual Organization of Sound. MIT press, 1994. 2
1994
-
[7]
COCO- Stuff: Thing and Stuff Classes in Context
Holger Caesar, Jasper Uijlings, and Vittorio Ferrari. COCO- Stuff: Thing and Stuff Classes in Context. In IEEE Conf. Comput. Vis. Pattern Recog., 2018. 6
2018
-
[8]
TIM: A Time Interval Ma- chine for Audio-Visual Action Recognition
Jacob Chalk, Jaesung Huh, Evangelos Kazakos, Andrew Zisserman, and Dima Damen. TIM: A Time Interval Ma- chine for Audio-Visual Action Recognition. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 3
work page 2024
Show all 86 references
-
[9]
GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Guoguo Chen, Shuzhou Chai, et al. GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio. In Proc. Interspeech 2021, 2021. 6
2021
-
[10]
VGGSound: A Large-scale Audio-Visual Dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zis- serman. VGGSound: A Large-scale Audio-Visual Dataset. In IEEE Int. Conf. Acoust. Speech Signal Process., 2020. 3
2020
-
[11]
Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts
Xianjie Chen, Roozbeh Mottaghi, Xiaobai Liu, Sanja Fidler, Raquel Urtasun, and Alan Yuille. Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts. In IEEE Conf. Comput. Vis. Pattern Recog.,
-
[12]
Vision Transformer Adapter for Dense Predictions
Zhe Chen, Yuchen Duan, Wenhai Wang, Junjun He, Tong Lu, Jifeng Dai, and Yu Qiao. Vision Transformer Adapter for Dense Predictions. In Int. Conf. Learn. Represent., 2023. 5, 6
2023
-
[13]
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open- Source Suites
Zhe Chen, Weiyun Wang, Hao Tian, Shenglong Ye, Zhang- wei Gao, Erfei Cui, Wenwen Tong, Kongzhi Hu, Jiapeng Luo, Zheng Ma, et al. How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open- Source Suites. arXiv preprint arXiv:2404.16821 , 2024. 2, 3, 6, 7
2024 arXiv
-
[14]
Masked-attention Mask Transformer for Universal Image Segmentation
Bowen Cheng, Ishan Misra, Alexander G Schwing, Alexan- der Kirillov, and Rohit Girdhar. Masked-attention Mask Transformer for Universal Image Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 5, 6
2022
-
[15]
VideoLLaMA 2: Advancing Spatial- Temporal Modeling and Audio Understanding in Video- LLMs
Zesen Cheng, Sicong Leng, Hang Zhang, Yifei Xin, Xin Li, Guanzheng Chen, Yongxin Zhu, Wenqi Zhang, Ziyang Luo, Deli Zhao, et al. VideoLLaMA 2: Advancing Spatial- Temporal Modeling and Audio Understanding in Video- LLMs. arXiv preprint arXiv:2406.07476, 2024. 5, 7
2024 arXiv
-
[16]
Qwen2-Audio Technical Report
Yunfei Chu, Jin Xu, Qian Yang, Haojie Wei, Xipin Wei, Zhifang Guo, Yichong Leng, Yuanjun Lv, Jinzheng He, Junyang Lin, et al. Qwen2-Audio Technical Report. arXiv preprint arXiv:2407.10759, 2024. 3
2024 arXiv
-
[17]
MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions
Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, and Chen Change Loy. MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions. In Int. Conf. Comput. Vis., 2023. 2, 3, 4, 6, 7, 8
2023
-
[18]
MOSE: A New Dataset for Video Object Segmentation in Complex Scenes
Henghui Ding, Chang Liu, Shuting He, Xudong Jiang, Philip HS Torr, and Song Bai. MOSE: A New Dataset for Video Object Segmentation in Complex Scenes. InInt. Conf. Comput. Vis., 2023. 6
2023
-
[19]
Multimodal referring segmentation: A survey
Henghui Ding, Song Tang, Shuting He, Chang Liu, Zuxuan Wu, and Yu-Gang Jiang. Multimodal referring segmentation: A survey. arXiv, 2025. 2
2025
-
[20]
MOSEv2: A more challenging dataset for video object segmentation in complex scenes
Henghui Ding, Kaining Ying, Chang Liu, Shuting He, Yu- Gang Jiang, Philip HS Torr, and Song Bai. MOSEv2: A more challenging dataset for video object segmentation in complex scenes. arXiv, 2025. 6
2025
-
[21]
VITA: Towards Open-Source Interactive Omni Multimodal LLM
Chaoyou Fu, Haojia Lin, Zuwei Long, Yunhang Shen, Meng Zhao, Yifan Zhang, Xiong Wang, Di Yin, Long Ma, Xiawu Zheng, et al. VITA: Towards Open-Source Interactive Omni Multimodal LLM. arXiv, 2024. 3, 4, 6
2024
-
[22]
A VSegFormer: Audio-Visual Segmentation with Transformer
Shengyi Gao, Zhe Chen, Guo Chen, Wenhai Wang, and Tong Lu. A VSegFormer: Audio-Visual Segmentation with Transformer. In AAAI, 2024. 8
2024
-
[23]
https://github.com/RVC- Boss/ GPT-SoVITS, 2024
GPT-SoVITS. https://github.com/RVC- Boss/ GPT-SoVITS, 2024. 4
2024
-
[24]
Open- V ocabulary Audio-Visual Semantic Segmentation
Ruohao Guo, Liao Qu, Dantong Niu, Yanyu Qi, Wenzhen Yue, Ji Shi, Bowei Xing, and Xianghua Ying. Open- V ocabulary Audio-Visual Semantic Segmentation. In ACM Int. Conf. Multimedia, 2024. 2
2024
-
[25]
Multi- modal Instruction Tuned LLMs with Fine-grained Visual Perception
Junwen He, Yifan Wang, Lijun Wang, Huchuan Lu, Jun- Yan He, Jin-Peng Lan, Bin Luo, and Xuansong Xie. Multi- modal Instruction Tuned LLMs with Fine-grained Visual Perception. In IEEE Conf. Comput. Vis. Pattern Recog. ,
-
[26]
Decoupling Static and Hierarchical Motion Perception for Referring Video Seg- mentation
Shuting He and Henghui Ding. Decoupling Static and Hierarchical Motion Perception for Referring Video Seg- mentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 2
2024
-
[27]
A Generalized Framework for Video Instance Segmentation
Miran Heo, Sukjun Hwang, Jeongseok Hyun, Hanjung Kim, Seoung Wug Oh, Joon-Young Lee, and Seon Joo Kim. A Generalized Framework for Video Instance Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 6
2023
-
[28]
Deep clustering: Discriminative embeddings for segmentation and separation
John R Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe. Deep clustering: Discriminative embeddings for segmentation and separation. In IEEE Int. Conf. Acoust. Speech Signal Process., 2016. 7
2016
-
[29]
LoRA: Low-Rank Adaptation of Large Language Models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-Rank Adaptation of Large Language Models. In Int. Conf. Learn. Represent., 2022. 6 9
2022
-
[30]
Egocentric Audio-Visual Object Localization
Chao Huang, Yapeng Tian, Anurag Kumar, and Chenliang Xu. Egocentric Audio-Visual Object Localization. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 3
2023
-
[31]
Video Object Segmentation with Language Referring Expressions
Anna Khoreva, Anna Rohrbach, and Bernt Schiele. Video Object Segmentation with Language Referring Expressions. In ACCV, 2019. 2, 6, 8
2019
-
[32]
Segment Anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer White- head, Alexander C Berg, Wan-Yen Lo, et al. Segment Anything. In Int. Conf. Comput. Vis., 2023. 3, 5
2023
-
[33]
LISA: Reasoning Segmen- tation via Large Language Model
Xin Lai, Zhuotao Tian, Yukang Chen, Yanwei Li, Yuhui Yuan, Shu Liu, and Jiaya Jia. LISA: Reasoning Segmen- tation via Large Language Model. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 2, 3, 5, 6, 7, 8
2024
-
[34]
TVQA: Localized, Compositional Video Question Answer- ing
Jie Lei, Licheng Yu, Mohit Bansal, and Tamara L Berg. TVQA: Localized, Compositional Video Question Answer- ing. In Proc. of the Conf. on Empirical Methods in Nat. Lang. Process., 2018. 3, 7
2018
-
[35]
Learning to Answer Questions in Dynamic Audio-Visual Scenarios
Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu, Ji- Rong Wen, and Di Hu. Learning to Answer Questions in Dynamic Audio-Visual Scenarios. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 3
2022
-
[36]
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
Guangyao Li, Henghui Du, and Di Hu. Boosting Audio Visual Question Answering via Key Semantic-Aware Cues. In ACM Int. Conf. Multimedia, pages 5997–6005, 2024
2024
-
[37]
Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? arXiv, 2025
Jia Li, Wenjie Zhao, Ziru Huang, Yunhui Guo, and Yapeng Tian. Do Audio-Visual Segmentation Models Truly Segment Sounding Objects? arXiv, 2025. 3
2025
-
[38]
Robust Referring Video Object Segmentation with Cyclic Structural Consensus
Xiang Li, Jinglu Wang, Xiaohao Xu, Xiao Li, Bhiksha Raj, and Yan Lu. Robust Referring Video Object Segmentation with Cyclic Structural Consensus. In Int. Conf. Comput. Vis.,
-
[39]
Baichuan-Omni Technical Report
Yadong Li, Haoze Sun, Mingan Lin, Tianpeng Li, Guosheng Dong, Tao Zhang, Bowen Ding, Wei Song, Zhenglin Cheng, Yuqi Huo, et al. Baichuan-Omni Technical Report. arXiv preprint arXiv:2410.08565, 2024. 3
2024
-
[40]
Losh: Long-short text joint prediction network for referring video object segmentation
Linfeng Yuan and Miaojing Shi and Zijie Yue and Qijun Chen. Losh: Long-short text joint prediction network for referring video object segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2024. 2
2024
-
[41]
GRES: Generalized Referring Expression Segmentation
Chang Liu, Henghui Ding, and Xudong Jiang. GRES: Generalized Referring Expression Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 5, 6, 8
2023
-
[42]
Primitivenet: decomposing the global constraints for referring segmenta- tion
Chang Liu, Xudong Jiang, and Henghui Ding. Primitivenet: decomposing the global constraints for referring segmenta- tion. Visual Intelligence, 2024. 2
2024
-
[43]
Visual Instruction Tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual Instruction Tuning. In Adv. Neural Inform. Process. Syst., 2023. 3
2023
-
[44]
Improved Baselines with Visual Instruction Tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved Baselines with Visual Instruction Tuning. InIEEE Conf. Comput. Vis. Pattern Recog., 2024. 7
2024
-
[45]
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models
Shuo Liu, Kaining Ying, Hao Zhang, Yue Yang, Yuqi Lin, Tianle Zhang, Chuanhao Li, Yu Qiao, Ping Luo, Wenqi Shao, et al. ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Ablation Capability for Large Vision-Language Models. In Adv. Neural Inform. Proc...
2024
-
[46]
Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Sep- aration
Yi Luo and Nima Mesgarani. Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Sep- aration. IEEE/ACM Trans. Audio Speech Lang. Process., 27 (8), 2019. 7
2019
-
[47]
Stepping Stones: A Progressive Training Strategy for Audio- Visual Semantic Segmentation
Juncheng Ma, Peiwen Sun, Yaoting Wang, and Di Hu. Stepping Stones: A Progressive Training Strategy for Audio- Visual Semantic Segmentation. In Eur. Conf. Comput. Vis.,
-
[48]
Generation and Comprehension of Unambiguous Object Descriptions
Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. Generation and Comprehension of Unambiguous Object Descriptions. In IEEE Conf. Comput. Vis. Pattern Recog., 2016. 6, 8
2016
-
[49]
V-Net: Fully Convolutional Neural Networks for V olumetric Medical Image Segmentation
Fausto Milletari, Nassir Navab, and Seyed-Ahmad Ahmadi. V-Net: Fully Convolutional Neural Networks for V olumetric Medical Image Segmentation. In IEEE Int. Conf. 3D Vis. ,
-
[50]
https://platform.openai.com/docs/ guides/text-to-speech, 2023
OpenAI. https://platform.openai.com/docs/ guides/text-to-speech, 2023. 4
2023
-
[51]
https://openai.com/index/hello- gpt-4o, 2024
OpenAI. https://openai.com/index/hello- gpt-4o, 2024. 2, 3
2024
-
[52]
Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks
Wenwen Pan, Haonan Shi, Zhou Zhao, Jieming Zhu, Xi- uqiang He, Zhigeng Pan, Lianli Gao, Jun Yu, Fei Wu, and Qi Tian. Wnet: Audio-Guided Video Object Segmentation via Wavelet-Based Cross-Modal Denoising Networks. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 4
2022
-
[53]
DetGPT: Detect What You Need via Reasoning
Renjie Pi, Jiahui Gao, Shizhe Diao, Rui Pan, Hanze Dong, Jipeng Zhang, Lewei Yao, Jianhua Han, Hang Xu, Lingpeng Kong, et al. DetGPT: Detect What You Need via Reasoning. In Proc. of the Conf. on Empirical Methods in Nat. Lang. Process., 2023. 2, 3
2023
-
[54]
Robust Speech Recognition via Large-Scale Weak Supervision
Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust Speech Recognition via Large-Scale Weak Supervision. InInt. Conf. Mach. Learn., 2023. 6
2023
-
[55]
PACO: Parts and Attributes of Common Objects
Vignesh Ramanathan, Anmol Kalia, Vladan Petrovic, Yi Wen, Baixue Zheng, Baishan Guo, Rui Wang, Aaron Mar- quez, Rama Kovvuri, Abhishek Kadian, et al. PACO: Parts and Attributes of Common Objects. In IEEE Conf. Comput. Vis. Pattern Recog., 2023. 6
2023
-
[56]
SAM 2: Segment Anything in Images and Videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, et al. SAM 2: Segment Anything in Images and Videos. arXiv preprint arXiv:2408.00714, 2024. 4
2024 arXiv
-
[57]
URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark
Seonguk Seo, Joon-Young Lee, and Bohyung Han. URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark. In Eur. Conf. Comput. Vis., 2020. 2, 4, 6, 8
2020
-
[58]
video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
Guangzhi Sun, Wenyi Yu, Changli Tang, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, Yuxuan Wang, and Chao Zhang. video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models. In Int. Conf. Mach. Learn., 2024. 3, 5
2024
-
[59]
Auto- ACD: A Large-scale Dataset for Audio-Language Represen- tation Learning
Luoyi Sun, Xuenan Xu, Mengyue Wu, and Weidi Xie. Auto- ACD: A Large-scale Dataset for Audio-Language Represen- tation Learning. In ACM Int. Conf. Multimedia, 2024. 6 10
2024
-
[60]
Unveiling and Mitigating Bias in Audio Visual Segmentation
Peiwen Sun, Honggang Zhang, and Di Hu. Unveiling and Mitigating Bias in Audio Visual Segmentation. In ACM Int. Conf. Multimedia, 2024. 3
2024
-
[61]
SALMONN: Towards Generic Hearing Abilities for Large Language Models
Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun Ma, and Chao Zhang. SALMONN: Towards Generic Hearing Abilities for Large Language Models. In Int. Conf. Learn. Represent., 2023. 3
2023
-
[62]
Audio-Visual Event Localization in Unconstrained Videos
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chen- liang Xu. Audio-Visual Event Localization in Unconstrained Videos. In Eur. Conf. Comput. Vis., 2018. 3
2018
-
[63]
Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-VL: Enhancing Vision-Language Model’s Perception of the World at Any Resolution. arXiv preprint arXiv:2409.12191, 2024. 3
2024 arXiv
-
[64]
Audio-Visual Grounding Referring Ex- pression for Robotic Manipulation
Yefei Wang, Kaili Wang, Yi Wang, Di Guo, Huaping Liu, and Fuchun Sun. Audio-Visual Grounding Referring Ex- pression for Robotic Manipulation. InIEEE Int. Conf. Robot. Autom., 2022. 3
2022
-
[66]
Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer
Yaoting Wang, Weisong Liu, Guangyao Li, Jian Ding, Di Hu, and Xi Li. Prompting Segmentation with Sound Is Generalizable Audio-Visual Source Localizer. In AAAI,
-
[67]
Can Textual Semantics Mitigate Sounding Object Segmentation Preference? In Eur
Yaoting Wang, Peiwen Sun, Yuanchao Li, Honggang Zhang, and Di Hu. Can Textual Semantics Mitigate Sounding Object Segmentation Preference? In Eur. Conf. Comput. Vis., 2024
2024
-
[68]
Ref-A VS: Refer and Segment Objects in Audio-Visual Scenes
Yaoting Wang, Peiwen Sun, Dongzhan Zhou, Guangyao Li, Honggang Zhang, and Di Hu. Ref-A VS: Refer and Segment Objects in Audio-Visual Scenes. In Eur. Conf. Comput. Vis.,
-
[69]
Language as Queries for Referring Video Object Segmentation
Jiannan Wu, Yi Jiang, Peize Sun, Zehuan Yuan, and Ping Luo. Language as Queries for Referring Video Object Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog. ,
-
[70]
VISA: Reasoning Video Object Segmentation via Large Language Models
Cilin Yan, Haochen Wang, Shilin Yan, Xiaolong Jiang, Yao Hu, Guoliang Kang, Weidi Xie, and Efstratios Gavves. VISA: Reasoning Video Object Segmentation via Large Language Models. In Eur. Conf. Comput. Vis., 2024. 2, 3, 4, 6, 8
2024
-
[71]
Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation
Shilin Yan, Renrui Zhang, Ziyu Guo, Wenchao Chen, Wei Zhang, Hongyang Li, Yu Qiao, Hao Dong, Zhongjiang He, and Peng Gao. Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation. In AAAI, 2024. 3, 7
2024
-
[72]
A VQA: A Dataset for Audio-Visual Question Answering on Videos
Pinci Yang, Xin Wang, Xuguang Duan, Hong Chen, Runze Hou, Cong Jin, and Wenwu Zhu. A VQA: A Dataset for Audio-Visual Question Answering on Videos. In ACM Int. Conf. Multimedia, 2022. 3
2022
-
[73]
LA VT: Language-Aware Vision Transformer for Referring Image Segmentation
Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Heng- shuang Zhao, and Philip HS Torr. LA VT: Language-Aware Vision Transformer for Referring Image Segmentation. In IEEE Conf. Comput. Vis. Pattern Recog., 2022. 8
2022
-
[74]
Isda: Position-aware instance segmentation with deformable attention
Kaining Ying, Zhenhua Wang, Cong Bai, and Pengfei Zhou. Isda: Position-aware instance segmentation with deformable attention. In IEEE Int. Conf. Acoust. Speech Signal Process.,
-
[75]
CTVIS: Consistent Training for Online Video Instance Segmentation
Kaining Ying, Qing Zhong, Weian Mao, Zhenhua Wang, Hao Chen, Lin Yuanbo Wu, Yifan Liu, Chengxiang Fan, Yunzhi Zhuge, and Chunhua Shen. CTVIS: Consistent Training for Online Video Instance Segmentation. In Int. Conf. Comput. Vis., 2023. 6
2023
-
[76]
MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Kaining Ying, Fanqing Meng, Jin Wang, Zhiqian Li, Han Lin, Yue Yang, Hao Zhang, Wenbo Zhang, Yuqi Lin, Shuo Liu, Jiayi Lei, Quanfeng Lu, Runjian Chen, Peng Xu, Renrui Zhang, Haozhe Zhang, Peng Gao, Yali Wang, Yu Qiao, Ping Luo, Kaipeng Zhang, and Wenqi Shao. MMT-Bench: A Compr...
2024
-
[77]
MOVE: Motion-guided few-shot video object segmentation
Kaining Ying, Hengrui Hu, and Henghui Ding. MOVE: Motion-guided few-shot video object segmentation. In Int. Conf. Comput. Vis., 2025. 6
2025
-
[78]
Modeling Context in Referring Expressions
Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg. Modeling Context in Referring Expressions. In Eur. Conf. Comput. Vis., 2016. 6, 8
2016
-
[79]
MOTR: End-to-End Multiple-Object Tracking with Transformer
Fangao Zeng, Bin Dong, Yuang Zhang, Tiancai Wang, Xiangyu Zhang, and Yichen Wei. MOTR: End-to-End Multiple-Object Tracking with Transformer. In Eur. Conf. Comput. Vis., 2022. 6
2022
-
[80]
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Hang Zhang, Xin Li, and Lidong Bing. Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding. In Proc. of the Conf. on Empirical Methods in Nat. Lang. Process., 2023. 5
2023
-
[81]
LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero- initialized Attention
Renrui Zhang, Jiaming Han, Chris Liu, Aojun Zhou, Pan Lu, Yu Qiao, Hongsheng Li, and Peng Gao. LLaMA-Adapter: Efficient Fine-tuning of Large Language Models with Zero- initialized Attention. In Int. Conf. Learn. Represent., 2024. 3
2024
-
[82]
GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
Shilong Zhang, Peize Sun, Shoufa Chen, Min Xiao, Wenqi Shao, Wenwei Zhang, Yu Liu, Kai Chen, and Ping Luo. GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest. arXiv preprint arXiv:2307.03601, 2023. 3
2023 arXiv
-
[83]
DVIS: Decoupled Video Instance Segmentation Framework
Tao Zhang, Xingye Tian, Yu Wu, Shunping Ji, Xuebo Wang, Yuan Zhang, and Pengfei Wan. DVIS: Decoupled Video Instance Segmentation Framework. In Int. Conf. Comput. Vis., 2023. 6
2023
-
[84]
ViLLa: Video Reasoning Segmentation with Large Language Model
Rongkun Zheng, Lu Qi, Xi Chen, Yi Wang, Kun Wang, Yu Qiao, and Hengshuang Zhao. ViLLa: Video Reasoning Segmentation with Large Language Model. arXiv preprint arXiv:2407.14500, 2024. 2, 6
2024 arXiv
-
[85]
Scene Parsing through ADE20K Dataset
Bolei Zhou, Hang Zhao, Xavier Puig, Sanja Fidler, Adela Barriuso, and Antonio Torralba. Scene Parsing through ADE20K Dataset. In IEEE Conf. Comput. Vis. Pattern Recog., 2017. 6
2017
-
[86]
Audio-Visual Segmentation
Jinxing Zhou, Jianyuan Wang, Jiayi Zhang, Weixuan Sun, Jing Zhang, Stan Birchfield, Dan Guo, Lingpeng Kong, Meng Wang, and Yiran Zhong. Audio-Visual Segmentation. In Eur. Conf. Comput. Vis., 2022. 2, 3, 8
2022
-
[87]
Tracking with Human-Intent Reasoning
Jiawen Zhu, Zhi-Qi Cheng, Jun-Yan He, Chenyang Li, Bin Luo, Huchuan Lu, Yifeng Geng, and Xuansong Xie. Tracking with Human-Intent Reasoning. arXiv, 2023. 8 11
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.