REVIEW 3 major objections 5 minor 45 references
DGFNet: End-to-End Audio-Visual Source Separation Based on Dynamic Gating Fusion
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read A learnable gating weight at the encoder bottleneck lets audio-visual source separation choose per sample how much vision to trust, gaining 0.62 dB SDR over the iQuery baseline on MUSIC.
desk verdict Plausible incremental architecture, but the ablation table contradicts the main baseline, making the headline gain uninterpretable; send to review with a request for variance and code. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Dynamic Gating Fusion Module (DGFM): a data-dependent convex combination of a fused audio-visual feature and the unmodified audio feature at the encoder bottleneck. A sigmoid gating weight $\sigma$ is computed from $1\times1$ convolutional projections of both candidates, and the final fused feature is $\sigma F_{av} + (1-\sigma)F_{mid}^a$. This replaces the fixed channel-wise concatenation or pixel-wise multiplication used by earlier systems, letting the network modulate the modality contribution per sample. The paper also adds an audio attention module adapted from the Efficient Multi-scale Attention module to the decoder upsampling path, and feeds the fused features into a query-based Audio-Visual Transformer decoder.
What would settle it
Reproduce iQuery with the identical training recipe (same STFT, object detector, optimizer, and test split) and compare SDR. If the faithful baseline reaches 10.95 dB—the value listed in the ablation table—rather than the 10.63 dB in the main table, DGFNet's advantage shrinks from 0.62 to about 0.30 dB, and the claimed margin depends on an unexplained setup difference rather than the gating module alone.
Extended reading notes
Core claim
The central claim is that a Dynamic Gating Fusion Module (DGFM) placed at the bottleneck of a U-Net audio encoder improves audio-visual source separation by adaptively reweighting audio and visual features. The fused audio-visual feature $F_{av}$ is computed by pixel-wise multiplication of object features and intermediate audio features; both $F_{av}$ and the audio feature $F_{mid}^a$ pass through $1\times1$ convolutions, their outputs are summed, and a sigmoid produces the gating coefficient $\sigma$. The module replaces the audio feature with $\sigma F_{av} + (1-\sigma)F_{mid}^a$, so the model can lean on vision when it is informative and fall back on audio when it is not. The paper reports that DGFNet outperforms the iQuery baseline on both datasets, and its ablation shows that removing only the gating module reduces SDR, supporting the module as the source of the gain.
Load-bearing premise
The central claim assumes the iQuery baseline was reproduced faithfully and that the main-table and ablation-table numbers are comparable; if those conditions differ, the reported gains cannot be credited to the gating module alone.
Editorial extensions
If this is right
- Other encoder-decoder audio-visual separation models can adopt the same bottleneck gating without altering their decoder, since DGFM only replaces the intermediate audio feature.
- The gain transfers across datasets with different instrument vocabularies (11 vs 21 classes), so the mechanism is not overfit to a single distribution.
- Because the learned $\sigma$ stays near 0.5 on average but shifts at the extremes, the dynamic part mainly matters for hard or ambiguous visual inputs; a fixed equal-weight fusion would lose those cases.
- Injecting vision at the bottleneck rather than only in the decoder means the generated separation mask is informed by visual evidence earlier, which should improve separation when the visual object is detectable.
Reading between the lines
- The gating coefficient could be conditioned on an explicit reliability estimate, such as object-detection confidence or estimated signal-to-noise ratio, making the adaptation interpretable and testable; the paper learns it implicitly.
- One could compare the distribution of $\sigma$ across instrument categories to check whether visually salient instruments receive systematically higher visual weights than occluded or small ones.
- The same convex gating formula is a natural fit for other audio-visual tasks with time-varying modality reliability, such as active source separation or audio-visual navigation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes DGFNet, an end-to-end audio-visual source separation model built on the iQuery architecture. The main additions are a Dynamic Gating Fusion Module (DGFM) that combines intermediate audio features with object-level visual features at the U-Net bottleneck using a learned sigmoid gate, and an audio attention module inserted after each upsampling layer of the U-Net decoder. The authors report SDR gains of 0.62 dB over iQuery on MUSIC and 0.36 dB on MUSIC-21, with corresponding SIR and SAR improvements, and present an ablation study of fusion variants plus an analysis of the learned gating weight distribution. The paper claims to be the first to combine bottleneck feature fusion with decoder-side interaction for this task.
Significance. If the reported gains are reproducible, DGFM is a simple and plausible improvement over iQuery: the gating mechanism is a natural way to let the model down-weight the visual modality when it is uninformative, and the audio attention module is a reasonable enhancement to the U-Net decoder. The paper also provides qualitative spectrogram comparisons and a distributional analysis of the gating weights, which help explain the mechanism. However, the empirical support is currently not strong enough to establish the central claim. The baseline SDR for iQuery is 10.63 dB in Table 1 but 10.95 dB in Table 3, no variance or significance testing is reported, no code is released, and the ablation does not isolate the audio attention module from DGFM. These issues make the headline 0.62 dB gain uninterpretable as evidence for the proposed components. The contribution is therefore significant only conditionally on the experiments being made consistent and statistically reliable.
major comments (3)
- [Tables 1 and 3, Sections 4.2.1 and 4.2.3] The baseline is inconsistent across the two tables. Table 1 reports iQuery at 10.63 dB SDR on MUSIC, while Table 3 lists the "Baseline" at 10.95 dB SDR on the same dataset. Recomputing the gains from Table 3, DGFNet improves over the no-fusion baseline by only 0.30 dB, and the DGFM-only variant improves by 0.22 dB, not the 0.62 dB claimed in Section 4.2.1. The authors must explain why the two baselines differ, specify the exact training/evaluation protocol for each row, and report all comparisons under identical conditions. Without a consistent baseline, the improvement cannot be attributed to the fusion module.
- [Section 4.2, Tables 1 and 2] No statistical reliability is established. The paper reports no number of training runs, no standard deviation or error bars, and no significance tests for any metric. The claimed improvements are on the order of 0.22 to 0.68 dB, which can fall within run-to-run variation for this type of model. The abstract's statement that the method achieves "significant performance improvements" is therefore not supported by the evidence presented. The authors should report results across multiple seeds with means and variances, and ideally a paired significance test against the reproduced iQuery baseline.
- [Table 3, Section 4.2.3] The ablation does not isolate the contribution of the audio attention module. The row "DGFNet(Ours)" includes both DGFM and the audio attention module, while "+DGFM" includes only the fusion module. Consequently the difference between "+DGFM" and "DGFNet(Ours)" reflects the audio attention module, but this component is never ablated separately against the baseline. Moreover, no fusion ablation is reported on MUSIC-21 at all. The authors should provide separate ablations for DGFM and audio attention on both datasets before claiming that dynamic gating is the key component.
minor comments (5)
- [Section 3.3, Figure 4] The dimensions in the DGFM description are underspecified: object features F_O are introduced as R^{C_O}, while intermediate audio features F_mid are R^{C_A x F_S x T_S}. The paper says the two are fused by pixel-wise multiplication along the channel dimension, but it does not state how the object features are projected or broadcast to the audio feature map. This should be spelled out for reproducibility.
- [Table 3] The row labels are ambiguous. "Baseline" is described as "the model without fusion at the bottleneck layer," which should coincide with the iQuery baseline from Table 1, yet the numbers differ. The caption should state explicitly which model components are present in each row, including the audio attention module.
- [Section 4.2.4, Figures 6 and 7] The weight distribution analysis is descriptive only. To support the claim that gating adaptively balances modalities, the authors should correlate the learned sigma values with separation quality or with visual-audio correspondence for individual test samples.
- [Section 4.1.3] The text says that some baseline numbers are taken from [2] and that iQuery was run in the authors' environment with "the same training settings," but it is not stated which other methods were reproduced locally and which were copied. A clear statement of reproduced versus cited results is needed for fair comparison.
- [Section 1, Contributions] The claim "we are the first to propose the combination of bottleneck feature fusion with decoder interaction decoding strategy" is not substantiated by a literature search or a positioning discussion against methods that use cross-modal attention or gating in related tasks. The statement should be softened or supported with evidence.
Circularity Check
No significant circularity: the fusion method is an empirical architecture evaluated against external baselines, with no fitted parameter renamed as a prediction and no load-bearing self-citation chain.
full rationale
DGFNet is an empirical architecture paper. The claimed contribution, the Dynamic Gating Fusion Module, is defined by its own equations in Section 3.3 and evaluated against independently reported baselines plus an iQuery reimplementation under the same training settings. There is no step in which a fitted parameter is renamed as a prediction, no uniqueness theorem imported from the authors' prior work, and no ansatz whose justification reduces to a self-citation. The paper's self-citations in the introduction ([36]-[39]) concern audio-visual navigation, not the separation method, and are not load-bearing for the central claim. The discrepancy between Table 1 (iQuery 10.63 dB) and Table 3 (Baseline 10.95 dB) is an empirical consistency concern that bears on whether the reported 0.62 dB gain is attributable to the module, but it is not circular reasoning; it does not make the derivation equivalent to its own inputs. The method is self-contained against external benchmarks, so the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (5)
- standard math STFT and iSTFT provide invertible spectrogram representations.
- domain assumption Pre-trained Faster R-CNN or Detic object detectors reliably localize sounding instruments.
- domain assumption I3D motion features encode instrument motion relevant to sound.
- domain assumption The Mix-and-Separate self-supervised task creates valid training targets.
- domain assumption U-Net with skip connections is a suitable audio backbone.
Cite this review
Pith. "Pith review of DGFNet: End-to-End Audio-Visual Source Separation Based on Dynamic Gating Fusion." pith.science (2026). https://pith.science/paper/L5TV2LJL
@misc{pith2026250421366,
author = {Pith},
title = {Pith review of: DGFNet: End-to-End Audio-Visual Source Separation Based on Dynamic Gating Fusion},
year = {2026},
howpublished = {\url{https://pith.science/paper/L5TV2LJL}},
note = {Machine review of arXiv:2504.21366}
}
read the original abstract
Current Audio-Visual Source Separation methods primarily adopt two design strategies. The first strategy involves fusing audio and visual features at the bottleneck layer of the encoder, followed by processing the fused features through the decoder. However, when there is a significant disparity between the two modalities, this approach may lead to the loss of critical information. The second strategy avoids direct fusion and instead relies on the decoder to handle the interaction between audio and visual features. Nonetheless, if the encoder fails to integrate information across modalities adequately, the decoder may be unable to effectively capture the complex relationships between them. To address these issues, this paper proposes a dynamic fusion method based on a gating mechanism that dynamically adjusts the modality fusion degree. This approach mitigates the limitations of solely relying on the decoder and facilitates efficient collaboration between audio and visual features. Additionally, an audio attention module is introduced to enhance the expressive capacity of audio features, thereby further improving model performance. Experimental results demonstrate that our method achieves significant performance improvements on two benchmark datasets, validating its effectiveness and advantages in Audio-Visual Source Separation tasks.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Triantafyllos Afouras, Andrew Owens, Joon Son Chung, and Andrew Zisserman
-
[2]
Jiaben Chen, Renrui Zhang, Dongze Lian, Jiaqi Yang, Ziyao Zeng, and Jianbo Shi. 2023. iQuery: Instruments as Queries for Audio-Visual Sound Separation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023. IEEE, 14675–14686. https://doi.org/10. 1109/CVPR52729.2023.01410
arXiv 2023
-
[3]
Haoyue Cheng, Zhaoyang Liu, Wayne Wu, and Limin Wang. 2023. Filter-Recovery Network for Multi-Speaker Audio-Visual Speech Separation. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net. https://openreview.net/forum?id=fiB2RjmgwQ6
work page 2023
-
[4]
Ying Cheng, Ruize Wang, Zhihao Pan, Rui Feng, and Yuejie Zhang. 2020. Look, Listen, and Attend: Co-Attention Network for Self-Supervised Audio-Visual Representation Learning. In MM ’20: The 28th ACM International Conference on Multimedia, Virtual Event / Seattle, W A, USA, October 12-16, 2020. ACM, 3884–3892. https://doi.org/10.1145/3394171.3413869 ICMR ’...
arXiv 2020
-
[5]
Shuangrui Ding, Maomao Li, Tianyu Yang, Rui Qian, Haohang Xu, Qingyi Chen, Jue Wang, and Hongkai Xiong. 2022. Motion-aware Contrastive Video Represen- tation Learning via Foreground-background Merging. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022. IEEE, 9706–9716. https://doi.org/10.1...
arXiv 2022
-
[6]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. [n. d.]. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In9th Interna- tional Conference on Learning Representations,...
work page 2021
-
[7]
Freeman, and Michael Rubinstein
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Has- sidim, William T. Freeman, and Michael Rubinstein. 2018. Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation. ACM Trans. Graph. 37, 4 (2018), 112. https://doi.org/10.1145/3197517.3201357
arXiv 2018
-
[8]
Tenenbaum, and Antonio Torralba
Chuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum, and Antonio Torralba. 2020. Music Gesture for Visual Sound Separation. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, W A, USA, June 13-19, 2020 . Computer Vision Foundation / IEEE, 10475–10484. https: //doi.org/10.1109/CVPR42600.2020.01049
arXiv 2020
Show all 45 references
-
[9]
Ruohan Gao, Rogério Schmidt Feris, and Kristen Grauman. 2018. Learning to Separate Object Sounds by Watching Unlabeled Video. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part III (Lecture Notes in Computer Scie...
2018 doi
-
[10]
Ruohan Gao and Kristen Grauman. 2019. Co-Separating Sounds of Visual Objects. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019 . IEEE, 3878–3887. https://doi.org/ 10.1109/ICCV.2019.00398
2019
-
[12]
Griffin and Jae S
Daniel W. Griffin and Jae S. Lim. 1983. Signal estimation from modified short- time Fourier transform. In IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP ’83, Boston, Massachusetts, USA, April 14-16, 1983 . IEEE, 804–807. https://doi.org/10.11...
1983
-
[13]
Simon Haykin and Zhe Chen. 2005. The Cocktail Party Problem. Neural Comput. 17, 9 (2005), 1875–1902. https://doi.org/10.1162/0899766054322964
2005 doi
-
[14]
Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe
John R. Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe. 2016. Deep clustering: Discriminative embeddings for segmentation and separation. In 2016 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2016, Shanghai, China, March 20-25, 201...
2016
-
[15]
Abrar Hussain. 2016. Evaluation of multichannel speech signal separation using Independent Component Analysis. In 2016 IEEE Students’ Conference on Electrical, Electronics and Computer Science (SCEECS) . 1–7. https://doi.org/10.1109/SCEECS. 2016.7509339
2016
-
[17]
Yanli Ji, Shuo Ma, Xing Xu, Xuelong Li, and Heng Tao Shen. 2023. Self-Supervised Fine-Grained Cycle-Separation Network (FSCN) for Visual-Audio Separation. IEEE Trans. Multim. 25 (2023), 5864–5876. https://doi.org/10.1109/TMM.2022. 3200282
2023 doi
-
[18]
Vahid Ahmadi Kalkhorani, Anurag Kumar, Ke Tan, Buye Xu, and DeLiang Wang. [n. d.]. Audiovisual Speaker Separation with Full- and Sub-Band Modeling in the Time-Frequency Domain. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2024, Seoul, Rep...
2024
-
[20]
Sagnik Majumder and Kristen Grauman. 2022. Active Audio-Visual Separation of Dynamic Sound Sources. In Computer Vision - ECCV 2022 - 17th European Conference, Tel A viv, Israel, October 23-27, 2022, Proceedings, Part XXXIX (Lecture Notes in Computer Science, Vol. 13699). Sprin...
2022
-
[21]
Mingyuan Mao, Peng Gao, Renrui Zhang, Honghui Zheng, Teli Ma, Peng Yan, Errui Ding, Baochang Zhang, and Shumin Han. 2021. Dual-stream Network for Visual Recognition. Neural Information Processing Systems,Neural Informa- tion Processing Systems (2021). https://proceedings.neuri...
2021
-
[22]
Daliang Ouyang, Su He, Guozhong Zhang, Mingzhu Luo, Huaiyong Guo, Jian Zhan, and Zhijie Huang. 2023. Efficient Multi-Scale Attention Module with Cross-Spatial Learning. In IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP 2023, Rhodes Island, Gree...
2023
-
[23]
Germain, Sameer Khurana, Chiori Hori, and Jonathan Le Roux
Zexu Pan, Gordon Wichern, Yoshiki Masuyama, François G. Germain, Sameer Khurana, Chiori Hori, and Jonathan Le Roux. 2023. Scenario-Aware Audio-Visual TF-Gridnet for Target Speech Extraction. In IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2023, Taipei, Ta...
2023
-
[24]
Samuel Pegg, Kai Li, and Xiaolin Hu. 2024. RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation. In The Twelfth Interna- tional Conference on Learning Representations . https://openreview.net/forum?id= PEuDO2EiDr
2024
-
[25]
Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, and DanielP.W
Colin Raffel, Brian McFee, EricJ. Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, and DanielP.W. Ellis. 2014. MIR_EVAL: A Transparent Implementation of Common MIR Metrics. International Symposium/Conference on Music Information Retrieval (2014)
2014
-
[26]
Tanzila Rahman, Mengyu Yang, and Leonid Sigal. 2021. TriBERT: Full-body Human-centric Audio-visual Representation Learning for Visual Sound Separa- tion. abs/2110.13412 (2021). https://arxiv.org/abs/2110.13412
2021 arXiv
-
[27]
Girshick, and Jian Sun
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2017. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.IEEE Trans. Pattern Anal. Mach. Intell. 39, 6 (2017), 1137–1149. https://doi.org/10.1109/TPAMI. 2016.2577031
2017
-
[28]
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015 - 18th International Conference Munich, Germany, October 5 - 9, 2015, Proceedi...
2015 doi
-
[29]
Zengjie Song and Zhaoxiang Zhang. 2024. Visually Guided Sound Source Separa- tion With Audio-Visual Predictive Coding. IEEE Transactions on Neural Networks and Learning Systems 35, 11 (2024), 15528–15542. https://doi.org/10.1109/TNNLS. 2023.3288022
2024
-
[30]
Yapeng Tian, Di Hu, and Chenliang Xu. 2021. Cyclic Co-Learning of Sounding Object Visual Grounding and Sound Separation. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 . Computer Vision Foundation / IEEE, 2745–2754. https://...
2021
-
[31]
Efthymios Tzinis, Scott Wisdom, Aren Jansen, Shawn Hershey, Tal Remez, Dan Ellis, and John R. Hershey. 2021. Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds. In 9th International Conference on Learning Representations, ICLR 2021, Virtual...
2021
-
[32]
Efthymios Tzinis, Scott Wisdom, Tal Remez, and John R. Hershey. 2022. Au- dioScopeV2: Audio-Visual Attention Architectures for Calibrated Open-Domain On-Screen Sound Separation. In Computer Vision - ECCV 2022 - 17th European Conference, Tel A viv, Israel, October 23-27, 2022, ...
2022 doi
-
[33]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, AidanN. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. Neural Information Processing Systems,Neural Information Pro- cessing Systems (Jun 2017). https://proceedings.neurips.c...
2017
-
[34]
Yuxin Ye, Wenming Yang, and Yapeng Tian. 2024. LAVSS: Location-Guided Audio- Visual Spatial Audio Separation. In IEEE/CVF Winter Conference on Applications of Computer Vision, W ACV 2024, Waikoloa, HI, USA, January 3-8, 2024 . IEEE, 5496–5507. https://doi.org/10.1109/WACV57701...
2024
-
[35]
Dong Yu, Morten Kolbæk, Zheng-Hua Tan, and Jesper Jensen. 2017. Permutation invariant training of deep models for speaker-independent multi-talker speech separation. In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2017, New Orleans, LA,...
2017
-
[36]
Yinfeng Yu, Lele Cao, Fuchun Sun, Xiaohong Liu, and Liejun Wang. 2022. Pay Self- Attention to Audio-Visual Navigation. In 33rd British Machine Vision Conference 2022, BMVC 2022, London, UK, November 21-24, 2022 . BMVA Press, 46
2022
-
[37]
Yinfeng Yu, Lele Cao, Fuchun Sun, Chao Yang, Huicheng Lai, and Wenbing Huang
-
[38]
Yinfeng Yu, Changan Chen, Lele Cao, Fangkai Yang, Wenbing Huang, and Fuchun Sun. 2023. Measuring Acoustics with Collaborative Multiple Agents. In The 32nd International Joint Conference on Artificial Intelligence, IJCAI 2023, Macao, 19th- 25th August 2023
2023
-
[39]
Yinfeng Yu, Wenbing Huang, Fuchun Sun, Changan Chen, Yikai Wang, and Xiaohong Liu. 2022. Sound Adversarial Audio-Visual Navigation. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, DGFNet: End-to-End Audio-Visual Source Separation Ba...
2022
-
[40]
Li Yuan, Yunpeng Chen, Tao Wang, Weihao Yu, Yujun Shi, Zihang Jiang, Francis E. H. Tay, Jiashi Feng, and Shuicheng Yan. [n. d.]. Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV) . ht...
2021
-
[41]
Hang Zhao, Orazio Gallo, Iuri Frosio, and Jan Kautz. 2017. Loss Functions for Image Restoration With Neural Networks. IEEE Transactions on Computational Imaging (2017), 47–57. https://doi.org/10.1109/TCI.2016.2644865
2017
-
[42]
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba. 2019. The Sound of Motions. In 2019 IEEE/CVF International Conference on Computer Vision, ICCV 2019, Seoul, Korea (South), October 27 - November 2, 2019 . IEEE, 1735–1744. https: //doi.org/10.1109/ICCV.2019.00182
2019
-
[43]
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. 2018. The Sound of Pixels. In The European Conference on Computer Vision (ECCV). http://arxiv.org/abs/1804.03160
2018 arXiv
-
[44]
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Krähenbühl, and Ishan Misra. 2022. Detecting Twenty-thousand Classes using Image-level Supervision. In ECCV (Lecture Notes in Computer Science, Vol. 13669) . Springer, 350–368. https: //doi.org/10.1007/978-3-031-20077-9_21
2022 doi
-
[45]
Lingyu Zhu and Esa Rahtu. 2020. Visually Guided Sound Source Separation Using Cascaded Opponent Filter Network. In Computer Vision - ACCV 2020 - 15th Asian Conference on Computer Vision, Kyoto, Japan, November 30 - December 4, 2020, Revised Selected Papers, Part VI (Lecture No...
2020 doi
-
[46]
Lingyu Zhu and Esa Rahtu. 2022. Visually Guided Sound Source Separation and Localization using Self-Supervised Motion Representations. In 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). 2171–2181. https: //doi.org/10.1109/wacv51458.2022.00223
2022
-
[2020]
In Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XVIII (Lecture Notes in Computer Science, Vol
Self-supervised Learning of Audio-Visual Objects from Video. In Computer Vision - ECCV 2020 - 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XVIII (Lecture Notes in Computer Science, Vol. 12363) . Springer, 208–224. https://doi.org/10.1007/978-3-0...
2020 doi
-
[2023]
Neural Computation 35, 5 (04 2023), 958–976
Echo-Enhanced Embodied Visual Navigation. Neural Computation 35, 5 (04 2023), 958–976
2023
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.