REVIEW 3 major objections 5 minor 202 references
A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read This survey claims that deep audio-visual correlation learning (AVCL) is best organized by model family and objective function, and that two trends dominate: self-supervised objectives produce more discriminative features, while attention…
desk verdict Useful survey of AV correlation learning, but the table that anchors its main claims mislabels mAP as accuracy; needs a revision before I'd trust the comparisons. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is a two-axis categorization: multimodal encoding models (attention mechanisms, auto-encoders, GANs, and diffusion) on one axis, and objective functions (cross-entropy, correlation loss, adversarial loss, ELBO, hinge, triplet, contrastive, NCE) on the other, applied to the four application groups of sound separation and localization, retrieval, recognition, and generation. The paper's Table 1 is the load-bearing structure: each surveyed method is placed by model type, loss, learning paradigm, metric, and benchmark, which is what makes the cross-application comparison and the two field trends visible. The audio-visual correlation formalization, paired features $x_1, x_2$ projected into a shared or separate common space $S$, defines the problem that every method in the table attacks.
What would settle it
Apply explicit inclusion criteria and re-benchmark a stratified sample of the surveyed methods on one shared protocol with the same datasets, metrics, and training budget; if self-supervised methods do not show more discriminative features than equally trained supervised baselines, the headline trend is false.
Extended reading notes
Core claim
On its own terms, the paper's discovery is taxonomic: it finds that methods aimed at different audio-visual applications share structural properties, so the field can be described by a small set of encoding models and loss families. It synthesizes two field-level trends: (i) self-supervised methods yield more discriminative features by leveraging large volumes of unlabeled data, and (ii) attention-based methods improve alignment and synchronization of audio and visual sequences. It also positions the reviewed methods inside deep knowledge representation, meaning they encode raw signals into abstract, non-explainable embeddings, and argues that injecting human-understandable structured knowledge, such as proxy tasks and pseudo-labels, is the natural route to interpretability and reliability.
Load-bearing premise
The survey's summary trends rest on published performance numbers collected from different datasets, metrics, and training setups, while the paper itself concedes that direct quantitative comparison between them is not possible.
Editorial extensions
If this is right
- A newcomer can choose an AVCL method by looking up model family and loss in the survey's overview table instead of reading the field application by application.
- Self-supervised pretraining on large unlabeled corpora will continue to be the main source of discriminative audio-visual features, with fine-tuning on smaller labeled sets for downstream tasks.
- Attention-based models, especially co-attention and cross-attention, will remain the default mechanism for fine-grained audio-visual alignment and synchronization.
- Without shared benchmarks and metrics, reported results across the surveyed methods cannot be ranked, so standard evaluation protocols are a prerequisite for the field's progress.
- Injecting human-understandable structured knowledge, such as pseudo-labels, proxy tasks, text, or graphs, is the paper's proposed direction toward interpretability and reliability of deep AVCL models.
Reading between the lines
- Because the survey does not run its own experiments, a controlled comparison that varies only the learning paradigm would be needed to confirm that self-supervision, not just scale, produces the reported discriminative advantage.
- The absence of explicit inclusion criteria means the comprehensiveness of the 199-method table is not independently checkable; a preregistered search could shift the reported distribution of model families.
- The paper's observation that attention interaction functions, such as addition versus dot-product, constrain what can be correlated suggests that matching interaction function to representation structure is a design variable future methods should report.
- Graph-based text semantics, mentioned as a way to reduce language-model cost, could also serve as a source of structured knowledge and as a principled way to choose negative samples for contrastive losses in AVCL.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a survey of deep audio-visual correlation learning (AVCL), organizing recent work by feature extraction models (attention, auto-encoders, GANs, diffusion), objective functions, datasets, and downstream tasks. It claims to fill a gap in prior surveys by systematically categorizing methods across application use cases, and it identifies two field trends: self-supervised methods produce more discriminative features, and attention-based methods improve audio-visual synchronization. The paper includes a large comparison table (Table 1), dedicated benchmark tables for audio and video encoders (Tables 2 and 3), a dataset catalog (Table 4), and a discussion of evaluation metrics and research challenges.
Significance. If the survey's organization were reliable, it would be a useful entry point for researchers, especially because it covers a broad corpus (199 references) and offers a taxonomy that spans model families, losses, and datasets. The paper also contributes a useful perspective on structured knowledge and self-supervised proxy tasks as a route toward more interpretable AVCL. The benchmark appendices are a practical resource. However, the quantitative claims that anchor the survey's central trends are undermined by metric mislabeling and by the paper's own admission that direct comparisons are not possible; the survey's value currently lies in its qualitative taxonomy, not in the numerical evidence, and several objective-function equations are garbled. These issues are correctable, and the underlying synthesis remains plausible.
major comments (3)
- [Section 4 and Table 1] The two headline trends are introduced as "As reflected in Table 1, recent research has intensively demonstrated that i) self-supervised learning methods ... and ii) attention-based methods ...". However, Table 1 is not internally reliable for that role: the AudioSet rows for refs [46], [44], and [62] are labeled "ACC" with values 50.4, 51.8, and 53.3, whereas Table 2 reports the same methods as mean Average Precision (mAP) with values 0.504, 0.518, and 0.533. ACC and mAP are different metrics, and the table also mixes WER, SDR, FID, Recall@10, AUC, and ACC across different datasets and protocols. In addition, Section 4 concedes that "it is not possible to make a direct quantitative comparison between them". The table should be relabeled, or preferably reframed as an index of representative methods rather than as quantitative evidence for the trends; the trend statements should be presented as qualitative observations consistent with the cited works, not as results demonstrated by Table 1.
- [Section 3.2, Eqs. (5), (12), (13), (16)] Several objective-function formulas are garbled. Eq. (12) is written as the sum of two identical terms, max(0, cos(c,i)+α) + max(0, cos(c,i)+α), which cannot be the intended hinge loss. Eq. (13) has inconsistent index use: M_ij appears in a sum over j, but the numerator and denominator use cos(ii) and cos(ij) without clearly separating the positive and negative pairs, and L_va is written with the same cos(ij) structure rather than the transposed comparison. Eq. (16) uses cos(x_a, x_v) inside the sum over negative samples, where it should be cos(x_a, x') for x' in the negative set. Eq. (5) defines A_ij = 1/2 cos(u_a, u_v), which does not depend on i and j as written. These errors compromise the reference value of the objective-function catalog, which is a central component of the survey.
- [Section 1.2 and Section 4] The survey does not state its search strategy, inclusion criteria, or temporal coverage for the 199 cited works, even though it claims to offer a "comprehensive study" that "fills this gap". Without this information, the representativeness and completeness of the review cannot be independently checked. The authors should add a short methodology paragraph describing the databases searched, the keywords used, the inclusion/exclusion criteria, and the period covered, or explicitly position the survey as a structured narrative review rather than a systematic one.
minor comments (5)
- [Abstract and Section 1] The abstract contains a sentence fragment: "The rapid growth of Deep Learning propelled the development of proposals that process audio-visual data and can be observed in the number of proposals in the past years. Thus encouraging the development of a comprehensive survey." This should be rewritten as a complete sentence.
- [Section 3.1.2 and Section 3.2] The itemized lists in Section 3.1.2 and Section 3.2 contain duplicate numbering: two items are labeled "vi" in Section 3.1.2, and the objective-function list in Section 3.2 uses "v" three times (Correlation Loss, Adversarial Loss, ELBO) in an order that does not match the subsections. Renumber the lists for clarity.
- [Section 3.2, Eq. (10)] Eq. (10) writes KL(q(Z|x)||p(Z|x)) = log(p(x)) - E_{Z∼q}[p_θ(x,Z)/q_θ(Z|x)], but the expectation should contain the log ratio, i.e., E[log(p_θ(x,Z)/q_θ(Z|x))]. The missing logarithm makes the identity incorrect.
- [Table 3] One row in Table 3 reads "VicTR† 72.4 51 RGB 32 [67]" with the second number likely a typo for 95.8, given the later row "VicTR† 70.7 95.8 RGB 32 [67]" and the corresponding paper. Please verify all entries against the source.
- [Section D.5] The description of SDR is inaccurate: SDR (signal-to-distortion ratio) measures the ratio of desired signal to distortion, not the Euclidean distance between representations or the amount of background noise. The text conflates SDR with other distance-based metrics.
Circularity Check
The survey's comparative claims are summaries of external literature; no internal derivation is circular.
full rationale
This is a survey paper; it makes no first-principles prediction and fits no parameter. Its load-bearing content is the taxonomy in Section 3 and the two trends in Sections 4 and 5, which are explicitly anchored to Table 1, a compilation of published results from external papers. The paper nowhere defines its target conclusions in terms of its own prior equations, and it does not cite its own work as the sole warrant for a forced conclusion. Self-citations such as [125], [173], [176], [177], [178], and [181] appear as contextual references or as example methods in Table 1, but the survey's claims do not reduce to those citations: [176] and [181] are cited as instances of correlation-loss and ELBO objectives, not as the justification for the field-level trends. The known limitation that Table 1 mixes ACC, mAP, SDR, WER, FID, and AUC across protocols, and the paper's own concession in Section 4 that 'it is not possible to make a direct quantitative comparison between them,' undermines the quantitative evidential weight of those trends, but that is a correctness and comparability concern rather than circularity by construction. No equation in the paper is equivalent to its input by definition, and no fitted value is renamed as a prediction.
Assumptions & free parameters
assumptions (4)
- domain assumption The set of 199 surveyed works is representative of the AVCL literature.
- domain assumption Published performance numbers across Tables 1-3 are comparable enough to support qualitative SOTA statements.
- domain assumption The taxonomy axes (model type, loss function, learning paradigm) are a valid way to organize AVCL methods.
- standard math Standard objective function mathematics (cross-entropy, ELBO, InfoNCE, hinge and contrastive losses) is as stated.
Cite this review
Pith. "Pith review of A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning." pith.science (2026). https://pith.science/paper/K4O2O5FH
@misc{pith2026241200049,
author = {Pith},
title = {Pith review of: A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/K4O2O5FH}},
note = {Machine review of arXiv:2412.00049}
}
read the original abstract
Audio-visual correlation learning aims to capture and understand natural phenomena between audio and visual data. The rapid growth of Deep Learning propelled the development of proposals that process audio-visual data and can be observed in the number of proposals in the past years. Thus encouraging the development of a comprehensive survey. Besides analyzing the models used in this context, we also discuss some tasks of definition and paradigm applied in AI multimedia. In addition, we investigate objective functions frequently used and discuss how audio-visual data is exploited in the optimization process, i.e., the different methodologies for representing knowledge in the audio-visual domain. In fact, we focus on how human-understandable mechanisms, i.e., structured knowledge that reflects comprehensible knowledge, can guide the learning process. Most importantly, we provide a summarization of the recent progress of Audio-Visual Correlation Learning (AVCL) and discuss the future research directions.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[46]
Liu, Andrew Rouditchenko, and James Glass
Yuan Gong, Alexander H. Liu, Andrew Rouditchenko, and James Glass. 2022. UAVM: Towards Unifying Audio and Visual Models.IEEE Signal Processing Letters 29 (2022), 2437–2441. https://doi.org/10.1109/lsp.2022.3224688
arXiv 2022
-
[44]
Mariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu, Mario Lucic, Cordelia Schmid, and Anurag Arnab. 2024. Audiovisual Masked Autoencoders. arXiv:2212.05922 [cs.CV] https://arxiv.org/abs/2212.05922
arXiv 2024
-
[62]
Po-Yao Huang, Vasu Sharma, Hu Xu, Chaitanya Ryali, Haoqi Fan, Yanghao Li, Shang-Wen Li, Gargi Ghosh, Jitendra Malik, and Christoph Feichtenhofer. 2023. MAViL: Masked Audio-Video Learners. arXiv:2212.08071 [cs.CV] https://arxiv.org/abs/2212.08071
arXiv 2023
-
[1]
Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan
-
[2]
Self-supervised object detection from audio-visual correspondence
Triantafyllos Afouras, Yuki M. Asano, Francois Fagan, Andrea Vedaldi, and Florian Metze. 2021. Self-supervised object detection from audio-visual correspondence. https://doi.org/10.48550/ARXIV.2104.06401
work page Pith review arXiv doi:10.48550/arxiv.2104.06401 2021
-
[3]
Triantafyllos Afouras, Andrew Owens, Joon Son Chung, and Andrew Zisserman. 2020. Self-supervised Learning of Audio-Visual Objects from Video. In Computer Vision – ECCV 2020 , Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (Eds.). Springer International Publishing, Cham, 208–224
2020
-
[4]
Hassan Akbari, Liangzhe Yuan, Rui Qian, Wei-Hong Chuang, Shih-Fu Chang, Yin Cui, and Boqing Gong. 2021. VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text. arXiv:2104.11178 [cs.CV] https://arxiv.org/abs/2104.11178
arXiv 2021
-
[5]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski,...
arXiv 2022
Show all 202 references
-
[6]
Jean-Baptiste Alayrac, Adria Recasens, Rosalia Schneider, Relja Arandjelović, Jason Ramapuram, Jeffrey De Fauw, Lucas Smaira, Sander Dieleman, and Andrew Zisserman. 2020. Self-Supervised MultiModal Versatile Networks. In Advances in Neural Information Processing Systems , H. L...
2020
-
[7]
Humam Alwassel, Dhruv Mahajan, Bruno Korbar, Lorenzo Torresani, Bernard Ghanem, and Du Tran. 2020. Self-Supervised Learning by Cross-Modal Audio-Video Clustering. arXiv:1911.12667 [cs.CV] https://arxiv.org/abs/1911.12667
2020 arXiv
-
[8]
Tenglong Ao, Zeyi Zhang, and Libin Liu. 2023. GestureDiffuCLIP: Gesture Diffusion Model with CLIP Latents. ACM Trans. Graph. 42, 4, Article 42 (jul 2023), 18 pages. https://doi.org/10.1145/3592097
2023 doi
-
[9]
Relja Arandjelovic and Andrew Zisserman. 2017. Look, listen and learn. In Proc. of the IEEE Int. Conf. on Computer Vision . 609–617
2017
-
[10]
Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021. ViViT: A Video Vision Transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 6836–6846
2021
-
[11]
Christos Athanasiadis, Enrique Hortal, and Stylianos Asteriadis. 2020. Audio-visual domain adaptation using conditional semi-supervised generative adversarial networks. Neurocomputing 397 (2020), 331–344
2020
-
[12]
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2017. Multimodal Machine Learning: A Survey and Taxonomy. arXiv:1705.09406 [cs.LG] https://arxiv.org/abs/1705.09406
2017 arXiv
-
[13]
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei. 2022. BEiT: BERT Pre-Training of Image Transformers. arXiv:2106.08254 [cs.CV] https: //arxiv.org/abs/2106.08254
2022 arXiv
-
[14]
Yue Cao, Mingsheng Long, et al. 2016. Correlation autoencoder hashing for supervised cross-modal search. In Proc. of the 2016 ACM on Int. Conf. on Multimedia Retrieval. 197–204
2016
-
[15]
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. 2021. Unsupervised Learning of Visual Features by Contrasting Cluster Assignments. arXiv:2006.09882 [cs.CV] https://arxiv.org/abs/2006.09882
2021 arXiv
-
[16]
Joao Carreira and Andrew Zisserman. 2017. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2017
-
[17]
Honglie Chen, Weidi Xie, Triantafyllos Afouras, Arsha Nagrani, Andrea Vedaldi, and Andrew Zisserman. 2021. Localizing Visual Sounds the Hard Way. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition . 16867–16876
2021
-
[18]
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2022. HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection. arXiv:2202.00874 [cs.SD] https://arxiv.org/abs/2202.00874
2022 arXiv
-
[19]
Sihan Chen, Xingjian He, Longteng Guo, Xinxin Zhu, Weining Wang, Jinhui Tang, and Jing Liu. 2023. VALOR: Vision-Audio-Language Omni- Perception Pretraining Model and Dataset. arXiv:2304.08345 [cs.LG] https://arxiv.org/abs/2304.08345
2023 arXiv
-
[20]
Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Daniel Tompkins, Zhuo Chen, and Furu Wei. 2022. BEATs: Audio Pre-Training with Acoustic Tokenizers. arXiv:2212.09058 [eess.AS] https://arxiv.org/abs/2212.09058
2022 arXiv
-
[21]
Shizhe Chen, Yida Zhao, Qin Jin, and Qi Wu. 2020. Fine-grained video-text retrieval with hierarchical graph reasoning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10638–10647
2020
-
[22]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A Simple Framework for Contrastive Learning of Visual Representa- tions. arXiv:2002.05709 [cs.LG] https://arxiv.org/abs/2002.05709
2020 arXiv
-
[23]
Ying Cheng, Ruize Wang, Zhihao Pan, Rui Feng, and Yuejie Zhang. 2020. Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning. In Proc. of the 28th ACM Int. Conf. on Multimedia . 3884–3892
2020
-
[24]
Chung-Cheng Chiu, James Qin, Yu Zhang, Jiahui Yu, and Yonghui Wu. 2022. Self-supervised learning with random-projection quantizer for speech recognition. In International Conference on Machine Learning . PMLR, 3915–3924
2022
-
[25]
Jeongsoo Choi, Joanna Hong, and Yong Man Ro. 2023. DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker Embedding. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 7812–7821. 2024-12-03 02:50. Page 29 of 1–44. Manuscript sub...
2023
-
[26]
Nieves Crasto, Philippe Weinzaepfel, Karteek Alahari, and Cordelia Schmid. 2019. MARS: Motion-Augmented RGB Stream for Action Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2019
-
[27]
Ishan Dave, Rohit Gupta, Mamshad Nayeem Rizve, and Mubarak Shah. 2022. TCLR: Temporal contrastive learning for video representation. Computer Vision and Image Understanding 219 (2022), 103406. https://doi.org/10.1016/j.cviu.2022.103406
2022
-
[28]
Julien Denize, Jaonary Rabarisoa, Astrid Orcesi, and Romain Hérault. 2023. Similarity contrastive estimation for image and video soft contrastive self-supervised learning. Machine Vision and Applications 34, 6 (2023). https://doi.org/10.1007/s00138-023-01444-9
2023 doi
-
[29]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805 [cs.CL] https://arxiv.org/abs/1810.04805
2019 arXiv
-
[30]
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Radford, and Ilya Sutskever. 2020. Jukebox: A Generative Model for Music. arXiv:2005.00341 [eess.AS] https://arxiv.org/abs/2005.00341
2020 arXiv
-
[31]
Prafulla Dhariwal and Alexander Nichol. 2021. Diffusion Models Beat GANs on Image Synthesis. In Advances in Neural Information Processing Systems, M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Vaughan (Eds.), Vol. 34. Curran Associates, Inc., 8780–8794. ht...
2021
-
[32]
Ali Diba, Mohsen Fayyaz, Vivek Sharma, Manohar Paluri, Jürgen Gall, Rainer Stiefelhagen, and Luc Van Gool. 2020. Large Scale Holistic Video Understanding. In Computer Vision – ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part V (Glasgow, U...
2020 doi
-
[33]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognit...
2021 arXiv
-
[34]
Bin Duan, Hao Tang, Wei Wang, Ziliang Zong, Guowei Yang, and Yan Yan. 2021. Audio-Visual Event Localization via Recursive Fusion by Joint Co-Attention. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV) . 4013–4022
2021
-
[35]
Haodong Duan, Yue Zhao, Yuanjun Xiong, Wentao Liu, and Dahua Lin. 2020. Omni-Sourced Webly-Supervised Learning for Video Recognition. In Computer Vision – ECCV 2020 , Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (Eds.). Springer International Publishing, C...
2020
-
[36]
Fanzeres and Climent Nadeu
Leonardo A. Fanzeres and Climent Nadeu. 2022. Sound-to-Imagination: An Exploratory Study on Unsupervised Crossmodal Translation Using Diverse Audiovisual Data. arXiv:2106.01266 [cs.SD] https://arxiv.org/abs/2106.01266
2022 arXiv
-
[37]
Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Girshick, and Kaiming He. 2021. A Large-Scale Study on Unsupervised Spatiotemporal Representation Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 3299–3309
2021
-
[38]
Christoph Feichtenhofer, Axel Pinz, and Andrew Zisserman. 2016. Convolutional Two-Stream Network Fusion for Video Action Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2016
-
[39]
Tenenbaum, and Antonio Torralba
Chuang Gan, Deng Huang, Peihao Chen, Joshua B. Tenenbaum, and Antonio Torralba. 2020. Foley Music: Learning to Generate Music from Videos. In Computer Vision – ECCV 2020 , Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frahm (Eds.). Springer International Publishi...
2020
-
[40]
Tenenbaum, and Antonio Torralba
Chuang Gan, Deng Huang, Hang Zhao, Joshua B. Tenenbaum, and Antonio Torralba. 2020. Music Gesture for Visual Sound Separation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2020
-
[41]
Ruohan Gao, Rogerio Feris, and Kristen Grauman. 2018. Learning to Separate Object Sounds by Watching Unlabeled Video. arXiv:1804.01665 [cs.CV] https://arxiv.org/abs/1804.01665
2018 arXiv
-
[42]
Ruohan Gao and Kristen Grauman. 2019. Co-Separating Sounds of Visual Objects. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2019
-
[43]
Gemmeke, Daniel P
Jort F. Gemmeke, Daniel P. W. Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R. Channing Moore, Manoj Plakal, and Marvin Ritter. 2017. Audio Set: An ontology and human-labeled dataset for audio events. In2017 IEEE International Conference on Acoustics, Speech and Signal Pr...
2017
-
[45]
Yuan Gong, Yu-An Chung, and James Glass. 2021. AST: Audio Spectrogram Transformer. arXiv:2104.01778 [cs.SD] https://arxiv.org/abs/2104.01778
2021 arXiv
-
[47]
Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James Glass
Yuan Gong, Andrew Rouditchenko, Alexander H. Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James Glass. 2023. Contrastive Audio-Visual Masked Autoencoder. arXiv:2210.07839 [cs.MM] https://arxiv.org/abs/2210.07839
2023 arXiv
-
[48]
Shreyank N Gowda, Marcus Rohrbach, and Laura Sevilla-Lara. 2021. SMART Frame Selection for Action Recognition. Proceedings of the AAAI Conference on Artificial Intelligence 35, 2 (May 2021), 1451–1459. https://doi.org/10.1609/aaai.v35i2.16235
2021 doi
-
[49]
Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko
Jean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec, Pierre H. Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Daniel Guo, Mohammad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, Rémi Munos, and Michal Valko. 2020. Bootstrap your own l...
2020 arXiv
-
[50]
Shuyang Gu, Dong Chen, Jianmin Bao, Fang Wen, Bo Zhang, Dongdong Chen, Lu Yuan, and Baining Guo. 2022. Vector Quantized Diffusion Model for Text-to-Image Synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 10696–10706. Manus...
2022
-
[51]
Michael Hahn. 2020. Theoretical limitations of self-attention in neural sequence models. Trans. of the Association for Computational Linguistics 8 (2020), 156–171
2020
-
[52]
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. 2021. Masked Autoencoders Are Scalable Vision Learners. arXiv:2111.06377 [cs.CV] https://arxiv.org/abs/2111.06377
2021 arXiv
-
[53]
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Momentum Contrast for Unsupervised Visual Representation Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2020
-
[54]
Xiangteng He, Yuxin Peng, and Liu Xie. 2019. A New Benchmark and Approach for Fine-grained Cross-media Retrieval. In Proceedings of the 27th ACM International Conference on Multimedia (Nice, France) (MM ’19). Association for Computing Machinery, New York, NY, USA, 1740–1748. h...
2019
-
[55]
Jongkwang Hong, Bora Cho, Yong Won Hong, and Hyeran Byun. 2019. Contextual Action Cues from Camera Sensor for Multi-Stream Action Recognition. Sensors (Basel, Switzerland) 19, 6 (2019), 1382. https://doi.org/10.3390/s19061382
2019 doi
-
[56]
Shota Horiguchi, Naoyuki Kanda, and Kenji Nagamatsu. 2018. Face-Voice Matching using Cross-modal Embeddings. In Proceedings of the 26th ACM International Conference on Multimedia (Seoul, Republic of Korea) (MM ’18). Association for Computing Machinery, New York, NY, USA, 1011–...
2018
-
[57]
Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units. IEEE/ACM Transactions on Audio, Speech, and Language Processin...
2021
-
[58]
Di Hu, Zheng Wang, Haoyi Xiong, Dong Wang, Feiping Nie, and Dejing Dou. 2020. Curriculum Audiovisual Learning. arXiv:2001.09414 [cs.CV] https://arxiv.org/abs/2001.09414
2020 arXiv
-
[59]
Ruihan Hu, Songbing Zhou, Zhi Ri Tang, Sheng Chang, Qijun Huang, Yisen Liu, Wei Han, and Edmond Q. Wu. 2021. DMMAN: A two-stage audio–visual fusion framework for sound separation and event localization. Neural Networks 133 (2021), 229–239. https://doi.org/10.1016/j.neunet. 2020.10.003
2021 doi
-
[60]
Xixi Hu, Ziyang Chen, and Andrew Owens. 2022. Mix and Localize: Localizing Sound Sources in Mixtures. arXiv:2211.15058 [cs.CV] https: //arxiv.org/abs/2211.15058
2022 arXiv
-
[61]
Jingjia Huang, Yinan Li, Jiashi Feng, Xinglong Wu, Xiaoshuai Sun, and Rongrong Ji. 2022. Clover: Towards A Unified Video-Language Alignment and Fusion Model. arXiv:2207.07885 [cs.CV] https://arxiv.org/abs/2207.07885
2022 arXiv
-
[63]
Gabriel Ilharco, Yuan Zhang, and Jason Baldridge. 2019. Large-scale representation learning from visually grounded untranscribed speech. arXiv:1909.08782 [cs.CV] https://arxiv.org/abs/1909.08782
2019 arXiv
-
[64]
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira. 2021. Perceiver: General Perception with Iterative Attention. arXiv:2103.03206 [cs.CV] https://arxiv.org/abs/2103.03206
2021 arXiv
-
[65]
Aren Jansen, Daniel PW Ellis, Shawn Hershey, R Channing Moore, Manoj Plakal, Ashok C Popat, and Rif A Saurous. 2020. Coincidence, categorization, and consolidation: Learning to recognize sounds with minimal supervision. In ICASSP 2020-2020 IEEE Int. Conf. on Acoustics, Speech ...
2020
-
[66]
Aren Jansen, Manoj Plakal, Richard Channing Moore, Shawn Hershey, Ratheet Pandya, Ryan Rifkin, Jiayang Liu, and Daniel Ellis. 2020. Unsupervised Learning of Semantic Audio Representations
2020
-
[67]
Kumara Kahatapitiya, Anurag Arnab, Arsha Nagrani, and Michael S. Ryoo. 2024. VicTR: Video-conditioned Text Representations for Activity Recognition. arXiv:2304.02560 [cs.CV] https://arxiv.org/abs/2304.02560
2024 arXiv
-
[68]
Esat Kalfaoglu, Sinan Kalkan, and A
M. Esat Kalfaoglu, Sinan Kalkan, and A. Aydin Alatan. 2020. Late Temporal Modeling in 3D CNN Architectures with BERT for Action Recognition. In Computer Vision – ECCV 2020 Workshops , Adrien Bartoli and Andrea Fusiello (Eds.). Springer International Publishing, Cham, 731–747
2020
-
[69]
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang, Wenwu Wang, and Mark D Plumbley. 2020. Panns: Large-scale pretrained audio neural networks for audio pattern recognition. IEEE/ACM Trans. on Audio, Speech, and Language Processing 28 (2020), 2880–2894
2020
-
[70]
Bruno Korbar, Du Tran, and Lorenzo Torresani. 2019. SCSampler: Sampling Salient Clips From Video for Efficient Action Recognition. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2019
-
[71]
Khaled Koutini, Jan Schlüter, Hamid Eghbal-zadeh, and Gerhard Widmer. 2022. Efficient Training of Audio Transformers with Patchout. In Interspeech 2022, 23rd Annual Conference of the International Speech Communication Association, Incheon, Korea, 18-22 September 2022 . ISCA, 2...
2022 doi
-
[72]
Jiyoung Lee, Joon Son Chung, and Soo-Whan Chung. 2023. Imaginary voice: Face-styled diffusion model for text-to-speech. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 1–5
2023
-
[73]
Jun-Tae Lee, Mihir Jain, Hyoungwoo Park, and Sungrack Yun. 2021. Cross-Attentional Audio-Visual Fusion for Weakly-Supervised Action Localization. In International Conference on Learning Representations . https://openreview.net/forum?id=hWr3e3r-oH5
2021
-
[74]
Jinyu Li, Abdelrahman Mohamed, Geoffrey Zweig, and Yifan Gong. 2015. LSTM time and frequency recurrence for automatic speech recognition. In 2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU) . 187–191. https://doi.org/10.1109/ASRU.2015.7404793
2015
-
[75]
Ross, and Angjoo Kanazawa
Ruilong Li, Shan Yang, David A. Ross, and Angjoo Kanazawa. 2021. AI Choreographer: Music Conditioned 3D Dance Generation With AIST++. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 13401–13412. 2024-12-03 02:50. Page 31 of 1–44. Manuscript ...
2021
-
[76]
Yinxiao Li, Zhichao Lu, Xuehan Xiong, and Jonathan Huang. 2022. PERF-Net: Pose Empowered RGB-Flow Net. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). 513–522
2022
-
[77]
Yingming Li, Ming Yang, and Zhongfei Zhang. 2017. A Survey of Multi-View Representation Learning. IEEE Trans. on Knowledge and Data Engineering 31, 10 (2017), 1863–1883. https://doi.org/10.1109/tkde.2018.2872063 arXiv:1610.01206
2017
-
[78]
Zhaohui Li, Haitao Wang, and Xinghua Jiang. 2023. AudioFormer: Audio Transformer learns audio feature representations from discrete acoustic codes. arXiv:2308.07221 [cs.SD] https://arxiv.org/abs/2308.07221
2023 arXiv
-
[79]
Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2023. Foundations and Trends in Multimodal Machine Learning: Principles, Challenges, and Open Questions. arXiv:2209.03430 [cs.LG] https://arxiv.org/abs/2209.03430
2023 arXiv
-
[80]
Jie Lin, Zekun Mu, Tianqing Zhao, Hanlin Zhang, Xinyu Yang, and Peng Zhao. 2023. Action density based frame sampling for human action recognition in videos. Journal of Visual Communication and Image Representation 90 (2023), 103740. https://doi.org/10.1016/j.jvcir.2022.103740
2023
-
[81]
Wei Lin, Leonid Karlinsky, Nina Shvetsova, Horst Possegger, Mateusz Kozinski, Rameswar Panda, Rogerio Feris, Hilde Kuehne, and Horst Bischof
-
[82]
Yan-Bo Lin and Yu-Chiang Frank Wang. 2020. Audiovisual transformer with instance attention for audio-visual event localization. In Proc. of the Asian Conf. on Computer Vision
2020
-
[83]
Yan-Bo Lin and Yu-Chiang Frank Wang. 2021. Exploiting Audio-Visual Consistency with Partial Supervision for Spatial Audio Generation. arXiv:2105.00708 [cs.SD] https://arxiv.org/abs/2105.00708
2021 arXiv
-
[84]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv:1907.11692 [cs.CL] https://arxiv.org/abs/1907.11692
2019 arXiv
-
[85]
Yun Liu, Xiaoming Zhang, Feiran Huang, Bo Zhang, and Zhoujun Li. 2022. Cross-Attentional Spatio-Temporal Semantic Graph Networks for Video Question Answering. IEEE Transactions on Image Processing 31 (2022), 1684–1696. https://doi.org/10.1109/tip.2022.3142526
2022
-
[86]
Dezhao Luo, Chang Liu, Yu Zhou, Dongbao Yang, Can Ma, Qixiang Ye, and Weiping Wang. 2020. Video Cloze Procedure for Self-Supervised Spatio-Temporal Learning. arXiv:2001.00294 [cs.CV] https://arxiv.org/abs/2001.00294
2020 arXiv
-
[87]
Chenyang Lyu, Wenxi Li, Tianbo Ji, Longyue Wang, Liting Zhou, Cathal Gurrin, Linyi Yang, Yi Yu, Yvette Graham, and Jennifer Foster. 2023. Graph- Based Video-Language Learning with Multi-Grained Audio-Visual Alignment. InProceedings of the 31st ACM International Conference on M...
2023
-
[88]
Shuang Ma, Zhaoyang Zeng, Daniel McDuff, and Yale Song. 2021. Active Contrastive Learning of Audio-Visual Video Representations. arXiv:2009.09805 [cs.LG] https://arxiv.org/abs/2009.09805
2021 arXiv
-
[89]
Shuang Ma, Zhaoyang Zeng, Daniel McDuff, and Yale Song. 2021. Contrastive Self-Supervised Learning of Global-Local Audio-Visual Representations. https://openreview.net/forum?id=Py4VjN6V2JX
2021
-
[90]
Frank, and Mirjam Ernestus
Danny Merkx, Stefan L. Frank, and Mirjam Ernestus. 2019. Language learning using speech to image retrieval. arXiv:1909.03795 [cs.CL] https://arxiv.org/abs/1909.03795
2019 arXiv
-
[91]
Shaobo Min, Qi Dai, Hongtao Xie, Chuang Gan, Yongdong Zhang, and Jingdong Wang. 2021. Cross-Modal Attention Consistency for Video-Audio Unsupervised Learning. arXiv:2106.06939 [cs.CV] https://arxiv.org/abs/2106.06939
2021 arXiv
-
[92]
Schuller, and Maja Pantic
Rodrigo Mira, Konstantinos Vougioukas, Pingchuan Ma, Stavros Petridis, Björn W. Schuller, and Maja Pantic. 2022. End-to-End Video-To-Speech Synthesis using Generative Adversarial Networks. https://doi.org/10.1109/TCYB.2022.3162495 arXiv:2104.13332 [cs.LG]
2022
-
[93]
Mathew Monfort, SouYoung Jin, Alexander Liu, et al. 2021. Spoken moments: Learning joint audio-visual representations from video descriptions. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition . 14871–14881
2021
-
[94]
Pedro Morgado, Yi Li, and Nuno Vasconcelos. 2020. Learning Representations from Audio-Visual Spatial Alignment. arXiv:2011.01819 [cs.CV] https://arxiv.org/abs/2011.01819
2020 arXiv
-
[95]
Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun. 2022. Attention Bottlenecks for Multimodal Fusion. arXiv:2107.00135 [cs.CV] https://arxiv.org/abs/2107.00135
2022 arXiv
-
[96]
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim. 2017. Dual attention networks for multimodal reasoning and matching. In Proc. of the IEEE Conf. on computer vision and pattern recognition . 299–307
2017
-
[97]
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng. 2011. Multimodal deep learning. In Proceedings of the 28th international conference on machine learning (ICML-11) . 689–696
2011
-
[98]
O., Melo Victor H
Bastos Igor L. O., Melo Victor H. C., and William Robson Schwartz. 2020. Bubblenet: A Disperse Recurrent Structure To Recognize Activities. 2020 IEEE International Conference on Image Processing (ICIP) 00 (2020), 2216–2220. https://doi.org/10.1109/icip40778.2020.9190769
2020
-
[99]
Henriques, Geoffrey Zweig, and Andrea Vedaldi
Mandela Patrick, Yuki Asano, Polina Kuznetsova, Ruth Fong, Joao F. Henriques, Geoffrey Zweig, and Andrea Vedaldi. 2021. Multi-modal Self-Supervision from Generalized Data Transformations. https://openreview.net/forum?id=mgVbI13p96
2021
-
[100]
Yuxin Peng and Jinwei Qi. 2019. CM-GANs: Cross-modal Generative Adversarial Networks for Common Representation Learning. ACM Trans. Multimedia Comput. Commun. Appl. 15, 1, Article 22 (feb 2019), 24 pages. https://doi.org/10.1145/3284750
2019 doi
-
[101]
Hai Pham, Paul Pu Liang, Thomas Manzini, Louis-Philippe Morency, and Barnabás Póczos. 2019. Found in translation: Learning robust joint representations by cyclic translations between modalities. In Proc. of the AAAI Conf. on Artificial Intelligence , Vol. 33. 6892–6899
2019
-
[102]
AJ Piergiovanni and Michael S. Ryoo. 2019. Representation Flow for Action Recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . Manuscript submitted to ACM 2024-12-03 02:50. Page 32 of 1–44. Deep Audio-Visual Correlation Lea...
2019
-
[103]
Laure Pretet, Gael Richard, and Geoffroy Peeters. 2021. Cross-Modal Music-Video Recommendation: A Study of Design Choices. arXiv:2104.14799 [cs.MM] https://arxiv.org/abs/2104.14799
2021 arXiv
-
[104]
Hendrik Purwins, Bo Li, Tuomas Virtanen, Jan Schlüter, Shuo-Yiin Chang, and Tara Sainath. 2019. Deep learning for audio signal processing. IEEE Journal of Selected Topics in Signal Processing 13, 2 (2019), 206–219
2019
-
[105]
Rui Qian, Yeqing Li, Zheng Xu, Ming-Hsuan Yang, Serge Belongie, and Yin Cui. 2022. Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models. arXiv:2207.07646 [cs.CV] https://arxiv.org/abs/2207.07646
2022 arXiv
-
[106]
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, and Yin Cui. 2021. Spatiotemporal Contrastive Video Representation Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 6964–6974
2021
-
[107]
Zhaofan Qiu, Ting Yao, Chong-Wah Ngo, Xinmei Tian, and Tao Mei. 2019. Learning Spatio-Temporal Representation With Local and Global Diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
2019
-
[108]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learning Transferable Visual Models From Natural Language Supervision. arXiv:2103.000...
2021 arXiv
-
[109]
Vandana Rajan, Alessio Brutti, and Andrea Cavallaro. 2021. Robust Latent Representations Via Cross-Modal Translation and Alignment. In ICASSP 2021-2021 IEEE Int. Conf. on Acoustics, Speech and Signal Processing . IEEE, 4315–4319
2021
-
[110]
Janani Ramaswamy and Sukhendu Das. 2020. See the sound, hear the pixels. In Proc. of the IEEE/CVF Winter Conf. on Applications of Computer Vision. 2970–2979
2020
-
[111]
Aditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray, Chelsea Voss, Alec Radford, Mark Chen, and Ilya Sutskever. 2021. Zero-shot text-to-image generation. In International Conference on Machine Learning . PMLR, 8821–8831
2021
-
[112]
Karthik Ramesh, Chao Xing, Wupeng Wang, Dong Wang, and Xiao Chen. 2021. Vset: A Multimodal Transformer for Visual Speech Enhancement. ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 00 (2021), 6658–6662. https://doi.org/10.1...
2021
-
[113]
Adrià Recasens, Pauline Luc, Jean-Baptiste Alayrac, Luyu Wang, Ross Hemsley, Florian Strub, Corentin Tallec, Mateusz Malinowski, Viorica Patraucean, Florent Altché, Michal Valko, Jean-Bastien Grill, Aäron van den Oord, and Andrew Zisserman. 2021. Broaden Your Views for Self-Su...
2021 arXiv
-
[114]
Colorado J Reed, Xiangyu Yue, Ani Nrusimha, Sayna Ebrahimi, Vivek Vijaykumar, Richard Mao, Bo Li, Shanghang Zhang, Devin Guillory, Sean Metzger, Kurt Keutzer, and Trevor Darrell. 2022. Self-Supervised Pretraining Improves Self-Supervised Pretraining. In Proceedings of the IEEE...
2022
-
[115]
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 10684–10695
2022
-
[116]
Andrew Rouditchenko, Angie Boggust, David Harwath, Brian Chen, Dhiraj Joshi, Samuel Thomas, Kartik Audhkhasi, Hilde Kuehne, Rameswar Panda, Rogerio Feris, Brian Kingsbury, Michael Picheny, Antonio Torralba, and James Glass. 2021. AVLnet: Learning Audio-Visual Language Represen...
2021 arXiv
-
[117]
Ludan Ruan, Yiyang Ma, Huan Yang, Huiguo He, Bei Liu, Jianlong Fu, Nicholas Jing Yuan, Qin Jin, and Baining Guo. 2023. Mm-diffusion: Learning multi-modal diffusion models for joint audio and video generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat...
2023
-
[118]
Mostafa Sadeghi, Simon Leglaive, Xavier Alameda-Pineda, Laurent Girin, and Radu Horaud. 2020. Audio-Visual Speech Enhancement Using Conditional Variational Auto-Encoders. IEEE/ACM Transactions on Audio, Speech, and Language Processing 28 (2020), 1788–1800. https://doi.org/10. ...
2020
-
[119]
Aaqib Saeed, David Grangier, and Neil Zeghidour. 2021. Contrastive learning of general-purpose audio representations. In ICASSP 2021-2021 IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 3875–3879
2021
-
[120]
Tara N Sainath and Bo Li. 2016. Modeling Time-Frequency Patterns with LSTM vs. Convolutional Architectures for LVCSR Tasks.. In Interspeech. 813–817
2016
-
[121]
Pritam Sarkar and Ali Etemad. 2022. Self-Supervised Audio-Visual Representation Learning with Relaxed Cross-Modal Synchronicity. arXiv:2111.05329 [cs.CV] https://arxiv.org/abs/2111.05329
2022 arXiv
-
[122]
Pritam Sarkar and Ali Etemad. 2023. XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning. arXiv:2211.13929 [cs.CV] https://arxiv.org/abs/2211.13929
2023 arXiv
-
[123]
Florian Schmid, Khaled Koutini, and Gerhard Widmer. 2023. Efficient Large-scale Audio Tagging via Transformer-to-CNN Knowledge Distillation. arXiv:2211.04772 [cs.SD] https://arxiv.org/abs/2211.04772
2023 arXiv
-
[124]
Sanghyun Seo, Sanghyuck Na, and Juntae Kim. 2020. HMTL: Heterogeneous Modality Transfer Learning for Audio-Visual Sentiment Analysis. IEEE Access 8 (2020), 140426–140437
2020
-
[125]
Rajiv Ratn Shah, Yi Yu, and Roger Zimmermann. 2014. Advisor: Personalized video soundtrack recommendation by late fusion with heuristic rankings. In Proc. of the 22nd ACM Int. Conf. on Multimedia . 607–616
2014
-
[126]
Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, and Abdelrahman Mohamed. 2022. Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction. arXiv:2201.02184 [eess.AS] https://arxiv.org/abs/2201.02184 2024-12-03 02:50. Page 33 of 1–44. Manuscript submitted to ...
2022 arXiv
-
[127]
Zhaofeng Shi. 2021. A Survey on Audio Synthesis and Audio-Visual Multimodal Processing. arXiv:2108.00443 [eess.AS] https://arxiv.org/abs/2108. 00443
2021 arXiv
-
[128]
Karen Simonyan and Andrew Zisserman. 2014. Two-Stream Convolutional Networks for Action Recognition in Videos. arXiv:1406.2199 [cs.CV] https://arxiv.org/abs/1406.2199
2014 arXiv
-
[129]
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. 2015. Deep unsupervised learning using nonequilibrium thermody- namics. In International conference on machine learning . PMLR, 2256–2265
2015
-
[130]
Jonathan Stroud, David Ross, Chen Sun, Jia Deng, and Rahul Sukthankar. 2020. D3D: Distilled 3D Networks for Video Action Recognition. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV)
2020
-
[131]
Jabeen Summaira, Xi Li, Amin Muhammad Shoib, Songyuan Li, and Jabbar Abdul. 2021. Recent Advances and Trends in Multimodal Deep Learning: A Review. arXiv:2105.11087 [cs.CV] https://arxiv.org/abs/2105.11087
2021 arXiv
-
[132]
Li, Mingkui Tan, and Chuang Gan
Xinyu Sun, Peihao Chen, Liangwei Chen, Changhao Li, Thomas H. Li, Mingkui Tan, and Chuang Gan. 2023. Masked Motion Encoding for Self-Supervised Video Representation Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 2235–2245
2023
-
[133]
Li Tao, Xueting Wang, and Toshihiko Yamasaki. 2020. Self-supervised Video Representation Learning Using Inter-intra Contrastive Framework. https://doi.org/10.1145/3394171.3413694 arXiv:2008.02531 [cs.CV]
2020
-
[134]
Yapeng Tian, Jing Shi, Bochen Li, Zhiyao Duan, and Chenliang Xu. 2018. Audio-Visual Event Localization in Unconstrained Videos. arXiv:1803.08842 [cs.CV] https://arxiv.org/abs/1803.08842
2018 arXiv
-
[135]
Martine Toering, Ioannis Gatopoulos, Maarten Stol, and Vincent Tao Hu. 2022. Self-Supervised Video Representation Learning With Cross-Stream Prototypical Contrasting. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV) . 108–118
2022
-
[136]
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022. VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training. arXiv:2203.12602 [cs.CV] https://arxiv.org/abs/2203.12602
2022 arXiv
-
[137]
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015. Learning Spatiotemporal Features With 3D Convolutional Networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)
2015
-
[138]
Du Tran, Heng Wang, Lorenzo Torresani, Jamie Ray, Yann LeCun, and Manohar Paluri. 2018. A Closer Look at Spatiotemporal Convolutions for Action Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2018
-
[139]
Tyagi and C
V. Tyagi and C. Wellekens. 2005. On desensitizing the Mel-cepstrum to spurious spectral components for robust speech recognition. In Proceedings. (ICASSP ’05). IEEE International Conference on Acoustics, Speech, and Signal Processing, 2005. , Vol. 1. I/529–I/532 Vol. 1. https:...
2005 arXiv
-
[140]
Aaron van den Oord, Oriol Vinyals, and koray kavukcuoglu. 2017. Neural Discrete Representation Learning. In Advances in Neural Information Processing Systems, I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran As...
2017
-
[141]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2023. Attention Is All You Need. arXiv:1706.03762 [cs.CL] https://arxiv.org/abs/1706.03762
2023 arXiv
-
[142]
Sergey Verbitskiy, Vladimir Berikov, and Viacheslav Vyshegorodtsev. 2022. ERANNs: Efficient Residual Audio Neural Networks for Audio Pattern Recognition. https://doi.org/10.1016/j.patrec.2022.07.012 arXiv:2106.01621 [cs.SD]
2022 arXiv
-
[143]
Bokun Wang, Yang Yang, Xing Xu, Alan Hanjalic, and Heng Tao Shen. 2017. Adversarial cross-modal retrieval. In Proc. of the 25th ACM Int. Conf. on Multimedia. 154–162
2017
-
[144]
Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yinan He, Yi Wang, Yali Wang, and Yu Qiao. 2023. VideoMAE V2: Scaling Video Masked Autoencoders With Dual Masking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 14549–14560
2023
-
[145]
Luyu Wang, Kazuya Kawakami, and Aaron van den Oord. 2020. Contrastive Predictive Coding of Audio with an Adversary. InProc. Interspeech
2020
-
[146]
Lei Wang and Piotr Koniusz. 2021. Self-supervising Action Recognition by Statistical Moment and Subspace Descriptors. In Proceedings of the 29th ACM International Conference on Multimedia (Virtual Event, China) (MM ’21). Association for Computing Machinery, New York, NY, USA, ...
2021
-
[147]
Lei Wang, Piotr Koniusz, and Du Q. Huynh. 2019. Hallucinating IDT Descriptors and I3D Optical Flow Features for Action Recognition With CNNs. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2019
-
[148]
Luyu Wang, Pauline Luc, Adria Recasens, Jean-Baptiste Alayrac, and Aaron van den Oord. 2021. Multimodal Self-Supervised Learning of General Audio Representations. arXiv:2104.12807 [cs.SD] https://arxiv.org/abs/2104.12807
2021 arXiv
-
[149]
Luyu Wang and Aaron van den Oord. 2021. Multi-Format Contrastive Learning of Audio Representations. arXiv:2103.06508 [cs.SD] https: //arxiv.org/abs/2103.06508
2021 arXiv
-
[150]
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2019. Temporal Segment Networks for Action Recognition in Videos. IEEE Transactions on Pattern Analysis and Machine Intelligence 41, 11 (2019), 2740–2755. https://doi.org/10.1109/TPAMI. 201...
2019
-
[151]
Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, and Lu Yuan. 2022. BEVT: BERT Pretraining of Video Transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 14733–14743. M...
2022
-
[152]
Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Lu Yuan, and Yu-Gang Jiang. 2023. Masked Video Distillation: Rethinking Masked Feature Modeling for Self-Supervised Video Representation Learning. In Proceedings of the IEEE/CVF Conference on Computer ...
2023
-
[153]
Shaonan Wang, Jiajun Zhang, and Chengqing Zong. 2018. Associative Multichannel Autoencoder for Multimodal Word Representation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing , Ellen Riloff, David Chiang, Julia Hockenmaier, and Jun’ich...
2018 doi
-
[154]
Wenhui Wang, Hangbo Bao, Li Dong, Johan Bjorck, Zhiliang Peng, Qiang Liu, Kriti Aggarwal, Owais Khan Mohammed, Saksham Singhal, Subhojit Som, and Furu Wei. 2022. Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks. arXiv:2208.10442 [cs.CV] ht...
2022 arXiv
-
[155]
Xiaolong Wang, Ali Farhadi, and Abhinav Gupta. 2016. Actions˜ Transformations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2016
-
[156]
Xiaohan Wang, Linchao Zhu, Fei Wu, and Yi Yang. 2023. A Differentiable Parallel Sampler for Efficient Video Classification.ACM Trans. Multimedia Comput. Commun. Appl. 19, 3, Article 112 (feb 2023), 18 pages. https://doi.org/10.1145/3569584
2023 doi
-
[157]
Xiaohan Wang, Linchao Zhu, Yu Wu, and Yi Yang. 2023. Symbiotic Attention for Egocentric Action Recognition With Object-Centric Alignment. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 6 (2023), 6605–6617. https://doi.org/10.1109/tpami.2020.3015894
2023
-
[158]
Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. 2023. Masked Feature Prediction for Self-Supervised Visual Pre-Training. arXiv:2112.09133 [cs.CV] https://arxiv.org/abs/2112.09133
2023 arXiv
-
[159]
Yake Wei, Di Hu, Yapeng Tian, and Xuelong Li. 2022. Learning in Audio-visual Context: A Review, Analysis, and New Perspective. arXiv:2208.09579 [cs.CV] https://arxiv.org/abs/2208.09579
2022 arXiv
-
[160]
Hao Wu, Jiayuan Mao, Yufeng Zhang, et al. 2019. Unified visual-semantic embeddings: Bridging vision and language with structured meaning representations. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition . 6609–6618
2019
-
[161]
Wenhao Wu, Zhun Sun, and Wanli Ouyang. 2023. Revisiting Classifier: Transferring Vision-Language Models for Video Recognition. arXiv:2207.01297 [cs.CV] https://arxiv.org/abs/2207.01297
2023 arXiv
-
[162]
Yu Wu and Yi Yang. 2021. Exploring Heterogeneous Clues for Weakly-Supervised Audio-Visual Video Parsing. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 1326–1335
2021
-
[163]
Yu Wu, Linchao Zhu, Yan Yan, and Yi Yang. 2019. Dual Attention Matching for Audio-Visual Event Localization. InProceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)
2019
-
[164]
Zuxuan Wu, Ting Yao, Yanwei Fu, and Yu-Gang Jiang. 2017. Deep learning for video classification and captioning . Association for Computing Machinery and Morgan & Claypool, 3–29. https://doi.org/10.1145/3122865.3122867
2017
-
[165]
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. 2021. VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding. arXiv:2109.14084 [cs.CV] https://arxiv.org/abs/2109.14084
2021 arXiv
-
[166]
Haoming Xu, Runhao Zeng, Qingyao Wu, Mingkui Tan, and Chuang Gan. 2020. Cross-Modal Relation-Aware Networks for Audio-Visual Event Localization. In Proceedings of the 28th ACM International Conference on Multimedia (Seattle, WA, USA) (MM ’20). Association for Computing Machine...
2020
-
[167]
Peng Xu, Xiatian Zhu, and David A. Clifton. 2023. Multimodal Learning With Transformers: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 10 (2023), 12113–12132. https://doi.org/10.1109/TPAMI.2023.3275156
2023
-
[168]
Xing Xu, Li He, Huimin Lu, Lianli Gao, and Yanli Ji. 2019. Deep adversarial metric learning for cross-modal retrieval. World Wide Web 22, 2 (2019), 657–672
2019
-
[169]
Xing Xu, Kaiyi Lin, Yang Yang, Alan Hanjalic, and Heng Tao Shen. 2020. Joint Feature Synthesis and Embedding: Adversarial Cross-modal Retrieval Revisited. IEEE Trans. on Pattern Analysis and Machine Intelligence (2020)
2020
-
[170]
Hanyu Xuan, Lei Luo, Zhenyu Zhang, Jian Yang, and Yan Yan. 2021. Discriminative Cross-Modality Attention Network for Temporal Inconsistent Audio-Visual Event Localization. IEEE Transactions on Image Processing 30 (2021), 7878–7888. https://doi.org/10.1109/tip.2021.3106814
2021
-
[171]
Yi Yang, Yueting Zhuang, and Yunhe Pan. 2021. Multiple knowledge representation for big data artificial intelligence: framework, applications, and case studies. Frontiers of Information Technology & Electronic Engineering 22, 12 (2021), 1551–1558. https://doi.org/10.1631/fitee.2100463
2021 doi
-
[172]
Qinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu, Qi Qian, Ji Zhang, and Fei Huang. 2022. HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training. arXiv:2212.14546 [cs.CV] https://arxiv.org/abs/2212.14546
2022 arXiv
-
[173]
Yi Yu, Zhijie Shen, and Roger Zimmermann. 2012. Automatic music soundtrack generation for outdoor videos from contextual sensor information. In Proc. of the 20th ACM Int. Conf. on Multimedia . 1377–1378
2012
-
[174]
Joe Yue-Hei Ng, Matthew Hausknecht, Sudheendra Vijayanarasimhan, Oriol Vinyals, Rajat Monga, and George Toderici. 2015. Beyond Short Snippets: Deep Networks for Video Classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2015
-
[175]
Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, and Stéphane Deny. 2021. Barlow twins: Self-supervised learning via redundancy reduction. In International Conference on Machine Learning . PMLR, 12310–12320
2021
-
[176]
Donghuo Zeng, Jianming Wu, Gen Hattori, Yi Yu, and Rong Xu. 2021. Learning Explicit and Implicit Latent Common Spaces for Audio-Visual Cross-Modal Retrieval. arXiv:2110.13556 [cs.MM] https://arxiv.org/abs/2110.13556
2021 arXiv
-
[177]
Donghuo Zeng, Yi Yu, and Keizo Oyama. 2018. Audio-visual embedding for cross-modal music video retrieval through supervised deep CCA. In 2018 IEEE Int. Symposium on Multimedia (ISM) . IEEE, 143–150. 2024-12-03 02:50. Page 35 of 1–44. Manuscript submitted to ACM 36 Luís Vilaça, et al
2018
-
[178]
Donghuo Zeng, Yi Yu, and Keizo Oyama. 2020. Deep Triplet Neural Networks with Cluster-CCA for Audio-Visual Cross-Modal Retrieval. ACM Trans. Multimedia Comput. Commun. Appl. 16, 3, Article 76 (jul 2020), 23 pages. https://doi.org/10.1145/3387164
2020 doi
-
[179]
S. Zha, F. Luisier, W. Andrews, N. Srivastava, and R. Salakhutdinov. 2015. Exploiting Image-trained CNN Architectures for Unconstrained Video Classification. 26th British Machine Vision Conference (BMVC 15) , 60.1–60.13
2015
-
[180]
Jingran Zhang, Fumin Shen, Xing Xu, and Heng Tao Shen. 2019. Cooperative Cross-Stream Network for Discriminative Action Representation. arXiv:1908.10136 [cs.CV] https://arxiv.org/abs/1908.10136
2019 arXiv
-
[181]
Jiwei Zhang, Yi Yu, Suhua Tang, Jianming Wu, and Wei Li. 2021. Variational Autoencoder with CCA for Audio-Visual Cross-Modal Retrieval. arXiv:2112.02601 [cs.IR] https://arxiv.org/abs/2112.02601
2021 arXiv
-
[182]
Yanyi Zhang, Xinyu Li, Chunhui Liu, Bing Shuai, Yi Zhu, Biagio Brattoli, Hao Chen, Ivan Marsic, and Joseph Tighe. 2021. VidTr: Video Transformer Without Convolutions. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 13577–13587
2021
-
[183]
Yuanyuan Zhang, Zi-Rui Wang, and Jun Du. 2019. Deep fusion: An attention guided factorized bilinear pooling for audio-video emotion recognition. In 2019 Int. Joint Conf. on Neural Networks . IEEE, 1–8
2019
-
[184]
Hang Zhao, Chuang Gan, Wei-Chiu Ma, and Antonio Torralba. 2019. The sound of motions. In Proc. of the IEEE/CVF Int. Conf. on Computer Vision . 1735–1744
2019
-
[185]
Hang Zhao, Chuang Gan, Andrew Rouditchenko, Carl Vondrick, Josh McDermott, and Antonio Torralba. 2018. The Sound of Pixels. arXiv:1804.03160 [cs.CV] https://arxiv.org/abs/1804.03160
2018 arXiv
-
[186]
Liangli Zhen, Peng Hu, Xu Wang, and Dezhong Peng. 2019. Deep supervised cross-modal retrieval. In Proc. of the IEEE/CVF Conf. on Computer Vision and Pattern Recognition . 10394–10403
2019
-
[187]
Aihua Zheng, Menglan Hu, Bo Jiang, Yan Huang, Yan Yan, and Bin Luo. 2022. Adversarial-Metric Learning for Audio-Visual Cross-Modal Matching. IEEE Transactions on Multimedia 24 (2022), 338–351. https://doi.org/10.1109/TMM.2021.3050089
2022
-
[188]
Lecheng Zheng, Yu Cheng, Hongxia Yang, Nan Cao, and Jingrui He. 2021. Deep Co-Attention Network for Multi-View Subspace Learning. In Proceedings of the Web Conference 2021 (Ljubljana, Slovenia) (WWW ’21). Association for Computing Machinery, New York, NY, USA, 1528–1539. https...
2021
-
[189]
Yuan Zhi, Zhan Tong, Limin Wang, and Gangshan Wu. 2021. MGSampler: An Explainable Sampling Strategy for Video Action Recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 1513–1522
2021
-
[190]
Jinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen, Cihang Xie, Alan Yuille, and Tao Kong. 2022. iBOT: Image BERT Pre-Training with Online Tokenizer. arXiv:2111.07832 [cs.CV] https://arxiv.org/abs/2111.07832
2022 arXiv
-
[191]
Bin Zhu, Bin Lin, Munan Ning, Yang Yan, Jiaxi Cui, HongFa Wang, Yatian Pang, Wenhao Jiang, Junwu Zhang, Zongwei Li, Wancai Zhang, Zhifeng Li, Wei Liu, and Li Yuan. 2024. LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment. arXi...
2024 arXiv
-
[192]
Hao Zhu, Hao Zhu, Mandi Luo, Rui Wang, Rui Wang, Aihua Zheng, and Ran He. 2021. Deep Audio-visual Learning: A Survey. Int. Journal of Automation and Computing 18, 3 (2021), 351–376. https://doi.org/10.1007/s11633-021-1293-0
2021 doi
-
[193]
Lingting Zhu, Xian Liu, Xuanyu Liu, Rui Qian, Ziwei Liu, and Lequan Yu. 2023. Taming diffusion models for audio-driven co-speech gesture generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10544–10553
2023
-
[194]
Lingyu Zhu and Esa Rahtu. 2020. Visually Guided Sound Source Separation using Cascaded Opponent Filter Network. In Proceedings of the Asian Conference on Computer Vision (ACCV)
2020
-
[195]
Lingyu Zhu and Esa Rahtu. 2021. Leveraging Category Information for Single-Frame Visual Sound Source Separation. In 2021 9th European Workshop on Visual Information Processing . IEEE, 1–6
2021
-
[196]
Wangjiang Zhu, Jie Hu, Gang Sun, Xudong Cao, and Yu Qiao. 2016. A Key Volume Mining Deep Framework for Action Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
2016
-
[197]
Ye Zhu, Kyle Olszewski, Yu Wu, Panos Achlioptas, Menglei Chai, Yan Yan, and Sergey Tulyakov. 2022. Quantized GAN for Complex Music Generation from Dance Videos. arXiv:2204.00604 [cs.CV] https://arxiv.org/abs/2204.00604
2022 arXiv
-
[198]
Ye Zhu, Yu Wu, Hugo Latapie, Yi Yang, and Yan Yan. 2021. Learning Audio-Visual Correlations from Variational Cross-Modal Generation. arXiv:2102.03424 [cs.CV] https://arxiv.org/abs/2102.03424
2021 arXiv
-
[199]
VIDEO FRAME SEQUENCE ANALYSIS
Ye Zhu, Yu Wu, Kyle Olszewski, Jian Ren, Sergey Tulyakov, and Yan Yan. 2023. Discrete Contrastive Diffusion for Cross-Modal Music and Image Generation. arXiv:2206.07771 [cs.CV] https://arxiv.org/abs/2206.07771 Manuscript submitted to ACM 2024-12-03 02:50. Page 36 of 1–44. Deep...
2023 arXiv
-
[2016]
arXiv:1609.08675 [cs.CV] https://arxiv.org/abs/1609.08675
YouTube-8M: A Large-Scale Video Classification Benchmark. arXiv:1609.08675 [cs.CV] https://arxiv.org/abs/1609.08675
-
[2020]
https://doi.org/10.21437/Interspeech.2020-1891
826–830. https://doi.org/10.21437/Interspeech.2020-1891
2020 doi
-
[2023]
arXiv:2303.08914 [cs.CV] https://arxiv.org/abs/2303.08914
MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language Knowledge. arXiv:2303.08914 [cs.CV] https://arxiv.org/abs/2303.08914
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.