REVIEW 6 major objections 5 minor 66 references
KDC-MAE: Knowledge Distilled Contrastive Mask Auto-Encoder
T0 review · 6 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that adding a symmetric KL self-distillation loss between two complementary masked views of the same audio-video clip improves joint audio-visual representation learning over the contrastive masked autoencoder baseline.
desk verdict The paper has a good idea—complementary-mask self-distillation on CAV-MAE—but the mask recipe as written is arithmetically impossible, and the empirical support is weaker than the prose suggests. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the complementary dual-mask pair combined with a symmetric KL divergence on origin-corrected, mean-pooled joint embeddings. For each modality, two masks are generated so that the visible token sets are disjoint (M1 ∩ M2 = ∅) and each exposes 25% of the tokens at the 75% masking ratio; the two masked views pass through weight-shared audio, video, and joint encoders, and the loss Lkd = (D(p1||p2) + D(p2||p1))/2 is backpropagated only through the encoders. To make KL applicable, the pooled embedding is shifted by its minimum to be non-negative and then linearly normalized to a probability distribution, deliberately avoiding softmax. This mechanism is what gives the paper's 'modular correspondence': the same modality content should yield the same pooled code regardless of which visible subset is fed in.
What would settle it
Train KDC-MAE with the same complementary masks and reconstruction and contrastive losses, but replace the distillation target with a fixed or randomly shuffled pooled embedding while keeping all hyperparameters identical. If downstream accuracy stays at the reported level, the KL term is not the cause of the gains; if accuracy falls back toward the CAV-MAE baseline, the mutual-agreement loss is doing the work.
Extended reading notes
Core claim
The central claim is that making the encoder mask-agnostic improves audio-visual representation learning. Concretely, an encoder trained so that two disjoint subsets of visible patches from the same clip yield almost identical pooled joint embeddings—enforced by a symmetric KL divergence propagated only through the encoder—learns representations that transfer better to downstream tasks than an encoder trained only with reconstruction and contrastive losses. Because the decoder is discarded for downstream use, the paper argues that the encoder alone should carry the invariance, and the two complementary views serve as mutual ground truth in the absence of labels. The reported finetuning results support the claim: 64.23% versus 63.89% on VGG-Sound audio-visual classification, 41.03% versus 39.61% on AudioSet-20k audio-visual classification, and further gains on Kinetics-400 and ILSVRC in several configurations.
Load-bearing premise
The load-bearing premise is that two disjoint 25%-visible views of the same clip should produce nearly identical pooled joint embeddings, and that enforcing this agreement during pretraining transfers to better performance on fully visible new clips.
Editorial extensions
If this is right
- On audio-visual classification, KDC-MAE outperforms CAV-MAE on VGG-Sound (64.23% versus 63.89%) and on AudioSet-20k (41.03% versus 39.61%), and the paper reports gains in most audio-only and video-only finetuning configurations.
- After pretraining on Kinetics-400, the no-overlap dual-mask model reaches 71.90% on Kinetics action recognition and 74.76% on ILSVRC classification, above the CAV-MAE baselines of 71.69% and 73.73%.
- Complementary (non-overlapping) masks help more when the downstream annotation is audio-centric, while overlapping masks help more for video-centric annotation, because complementary masks expose more video tokens and can overfit when video is the label source.
- Using three parallel masked streams degrades performance relative to two; the paper attributes this to overexposure of data and overfitting, so two streams are the recommended configuration.
- On VGG-Sound, the paper reports that the proposed variants improve audio-to-visual retrieval and inpainting over CAV-MAE, while remaining competitive on localization; visual-to-audio retrieval metrics are mixed or lower.
Reading between the lines
- A direct way to test whether the mutual-agreement objective is doing the work is to evaluate the pretrained encoder on heavily occluded or randomly cropped clips: if the distillation loss truly creates mask-agnostic embeddings, KDC-MAE should degrade less than CAV-MAE under such input corruption.
- Because the KL loss operates on pooled embeddings, it could in principle encourage the encoder to discard patch-level information that pooling would wash out anyway; measuring performance on dense or patch-level downstream tasks, or removing the distillation term while keeping the dual-mask setup, would separate the effect of distillation from the effect of the extra augmented view.
- The shift-and-normalize projection to a probability distribution is not theoretically justified in the paper; replacing it with softmax or a learned projection head is a natural ablation that would show whether the specific projection matters for the gains.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes KDC-MAE, a self-supervised audio-visual pretraining method that combines a masked autoencoder reconstruction loss, a contrastive audio-visual loss, and a symmetric KL-divergence self-distillation loss between the joint embeddings of two masked views. The two views are generated by a proposed 'complementary masking' strategy intended to force the encoder to be mask-agnostic. The authors compare against CAV-MAE on VGG-Sound, AudioSet-20k, Kinetics-400, and ImageNet for classification, retrieval, inpainting, and localization, and report small accuracy gains in several configurations, which they attribute to the distillation objective and complementary masking.
Significance. If the central claims hold, the paper contributes a modest but potentially useful combination of the three main SSL paradigms for audio-visual representation learning, and it provides a broad set of ablations across datasets and downstream tasks. The breadth of experiments—including classification, retrieval, inpainting, and localization, and comparisons with adaptive masking and three-stream variants—is a strength. However, the manuscript is currently not solidly supported: the complementary-mask construction as written is arithmetically impossible, several tables contain mutually inconsistent numbers for the same configurations, and some written conclusions contradict the paper's own results. The claimed 'modular correspondence' is enforced by the loss by construction rather than discovered, so that interpretive claim needs to be reframed.
major comments (6)
- [Section 2.2] The complementary mask construction is arithmetically inconsistent as stated. The text defines M1 and M2 as sets of unmasked tokens with M1 ∩ M2 = ∅, then requires |M1| = |M2| = mr · |Ua| with mr = 0.75 and |Ua| = 512; two disjoint subsets of size 384 cannot both fit in a 512-element set, because the residue Ua − M1 has only 128 elements. The same contradiction holds for video with |Uv| = 196 and mr · |Uv| = 147. If 'mask' instead means the masked tokens, then the two 75% masks necessarily overlap by construction. The paper states that 75% masking leaves 128 audio and 49 video visible patches, so the size formula appears to be off by a factor of three. This makes the proposed 'No Overlap' mechanism undefined and prevents reproduction of the central contribution.
- [Section 2.3, Eq. (5)] The claim that the encoder learns 'modular correspondence' from the distillation loss is not independently supported. Lkd(p1,p2) is the symmetric KL divergence between the pooled joint embeddings of the two masked views and is minimized by making the two embeddings agree. Agreement between views is therefore enforced by construction and cannot be cited as evidence that a 'mutual ground truth' was discovered. The downstream finetuning results use independent labels and provide some grounding for the accuracy claim, but the interpretive conclusion about modular correspondence should be either removed or tested by probing the learned representations on tasks that do not optimize the same agreement objective.
- [Section 3.3, Table 2] The hyperparameter and model selection procedure raises fair-comparison concerns. λkd is swept on the same VGG-Sound and AudioSet-20k finetuning sets used to report the final results, and the final configuration is chosen as the best in that sweep; no held-out validation is described. In addition, the row 'Proposed No Overlap Dual Mask λkd=5' gives better AS-20k A-V accuracy (41.62) than the selected λkd=10 row (41.03), so the selection rule is unclear. Please specify how many seeds were run, whether the test sets were used for selection, and why λkd=10 is the headline configuration.
- [Tables 1 and 5; Section 3.3] The numbers for the same configurations are not consistent across tables. Table 1 reports CAV-MAE on AS-20k FT A-V as 39.61 and AS-20k FT V-only as 31.00, while Table 5 reports 37.02 and 28.86 for the same CAV-MAE AS-20k rows. Table 5 also shows no-overlap dual mask AS-20k (A only) at 37.01, below CAV-MAE at 37.21, and Kinetics-400 (A V) with overlap at 71.53, below CAV-MAE at 71.55, which contradicts the sentence 'classification accuracy only improves when using the proposed dual mask method in all datasets.' Please reconcile these numbers and adjust the claims accordingly.
- [Table 6] The abstract's claim of 'better learning ... over multiple tasks' is not supported by Table 6 for the headline 'No Overlap' configuration. Compared with CAV-MAE, No Overlap is worse on audio→visual retrieval (R@1 0.0506 vs 0.0580), visual→audio retrieval (R@1 0.0416 vs 0.0662), and visual sound source localization (Avg Cos Sim 0.2798 vs 0.2884); only inpainting loss improves. The sentence 'performs better in the case of audio-to-visual retrieval' is true only for the 'With Overlap' row. Please report the results for the final configuration consistently and temper the task-level claims.
- [Section 2.4] The statement that Lc and Lkd affect only 'later two' components is inconsistent with Eq. (1), which computes the contrastive loss on mean-pooled outputs of the modality-specific encoders (cv_i, ca_j). If Lc and Lkd intentionally stop gradients before the audio/video encoders, that architectural detail is missing; if not, the sentence is wrong. This affects which modules are trained by each objective and is important for reproducing the method.
minor comments (5)
- [Section 2.2] The text contains 'duel head MAE', which should be 'dual head MAE'.
- [Section 2.3] The text contains an unresolved 'refer Fig. ??' reference; Figure 2 is presumably intended.
- [Section 3.1] The dataset name is misspelled as 'ILSRVC' and should be 'ILSVRC'; capitalization of 'VGGsound' is also inconsistent.
- [Table 2 text] The definitions of overlapped and non-overlapped masks give the same condition for both: 'Mi ∩ Mj = φ ∀ i,j, i≠j' appears twice; the overlapping case should presumably have Mi ∩ Mj ≠ φ.
- [Section 3.3] No error bars or multiple-seed results are reported for the finetuning accuracy numbers, several of which differ by less than 0.5 percentage points; this should be addressed in the final version.
Circularity Check
No significant circularity: the claimed gains are measured on independent downstream labels, not on the self-distillation objective itself.
full rationale
KDC-MAE's central claim is that adding symmetric-KL self-distillation between two masked views of the same audio-visual clip improves SSL representations. The evidence for that claim is downstream finetuning accuracy on VGG-Sound, AudioSet-20k, Kinetics-400 and ILSVRC, plus retrieval/inpainting/localization numbers; these labels and tasks are external to the pretraining objective. The Lkd term indeed makes cross-view agreement a training objective by construction, but the paper never offers cross-view agreement itself as an experimental result or as validation; it is the loss being minimized. The Section 2.3 statement 'due to lack of ground truth in L′, mutual ground truth is the only option left' is a motivation for choosing self-distillation, not a circular derivation of the reported downstream gains. Hyperparameters such as λc=0.01 and the ablated λkd values are tuned or inherited from prior work, but the headline classification numbers are measured against held-out labels, so they are not forced by the loss definition. There are no load-bearing self-citations by the present authors: CAV-MAE [16] and DML [65] are external prior work. A separate, non-circular correctness problem exists in Section 2.2: two disjoint 'sets of unmasked tokens' of size mr·|Ua| cannot both fit in |Ua|=512 when mr=0.75, nor in |Uv|=196; this is an internal inconsistency that affects reproducibility, but it is not a circularity and therefore does not raise the circularity score.
Assumptions & free parameters
free parameters (4)
- lambda_c (contrastive loss weight) =
0.01
- lambda_kd (self-distillation loss weight) =
10 in final proposed config; 5 and 15 also tested
- masking_ratio =
0.75
- temperature tau in contrastive loss =
not specified
assumptions (5)
- ad hoc to paper The encoder F should satisfy F(x1) approximately equal to F(x2) for disjoint subsets x1, x2 of the same input, and this can be learned without ground truth via mutual agreement.
- ad hoc to paper Origin-shifted, linearly normalized mean-pooled embeddings are valid probability distributions for KL divergence.
- domain assumption Audio and video of the same 10-second clip correspond, and a single video frame represents the visual content.
- ad hoc to paper Symmetric KL divergence between the two views is a valid self-distillation objective that improves mask-agnostic learning.
- standard math Standard transformer layers, InfoNCE loss, and KL divergence properties behave as in the cited literature.
Cite this review
Pith. "Pith review of KDC-MAE: Knowledge Distilled Contrastive Mask Auto-Encoder." pith.science (2026). https://pith.science/paper/KTMMPBSV
@misc{pith2026241112270,
author = {Pith},
title = {Pith review of: KDC-MAE: Knowledge Distilled Contrastive Mask Auto-Encoder},
year = {2026},
howpublished = {\url{https://pith.science/paper/KTMMPBSV}},
note = {Machine review of arXiv:2411.12270}
}
read the original abstract
In this work, we attempted to extend the thought and showcase a way forward for the Self-supervised Learning (SSL) learning paradigm by combining contrastive learning, self-distillation (knowledge distillation) and masked data modelling, the three major SSL frameworks, to learn a joint and coordinated representation. The proposed technique of SSL learns by the collaborative power of different learning objectives of SSL. Hence to jointly learn the different SSL objectives we proposed a new SSL architecture KDC-MAE, a complementary masking strategy to learn the modular correspondence, and a weighted way to combine them coordinately. Experimental results conclude that the contrastive masking correspondence along with the KD learning objective has lent a hand to performing better learning for multiple modalities over multiple tasks.
Figures
Reference graph
Works this paper leans on
-
[1]
Model selection and multi- model inference
D Anderson and K Burnham. Model selection and multi- model inference. Second. NY: Springer-Verlag, 63(2020):10,
work page 2020
-
[2]
Adamae: Adaptive masking for efficient spatiotempo- ral learning with masked autoencoders
Wele Gedara Chaminda Bandara, Naman Patel, Ali Gho- lami, Mehdi Nikkhah, Motilal Agrawal, and Vishal M Pa- tel. Adamae: Adaptive masking for efficient spatiotempo- ral learning with masked autoencoders. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14507–14517, 2023. 7
work page 2023
-
[3]
Self- supervised learning across domains
Silvia Bucci, Antonio D’Innocente, Yujun Liao, Fabio M Carlucci, Barbara Caputo, and Tatiana Tommasi. Self- supervised learning across domains. IEEE Transactions on Pattern Analysis and Machine Intelligence , 44(9):5516– 5528, 2021. 1
work page 2021
-
[4]
How to understand masked autoencoders
Shuhao Cao, Peng Xu, and David A Clifton. How to understand masked autoencoders. arXiv preprint arXiv:2202.03670, 2022. 3
arXiv 2022
-
[5]
Domain generalization by solving jigsaw puzzles
Fabio M Carlucci, Antonio D’Innocente, Silvia Bucci, Bar- bara Caputo, and Tatiana Tommasi. Domain generalization by solving jigsaw puzzles. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2229–2238, 2019. 1
2019
-
[6]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021. 1, 4
work page 2021
-
[7]
Krishna Chaitanya, Ertunc Erdil, Neerav Karani, and Ender Konukoglu. Contrastive learning of global and local fea- tures for medical image segmentation with limited annota- tions. Advances in neural information processing systems , 33:12546–12558, 2020. 1
work page 2020
-
[8]
Vggsound: A large-scale audio-visual dataset
Honglie Chen, Weidi Xie, Andrea Vedaldi, and Andrew Zisserman. Vggsound: A large-scale audio-visual dataset. In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages 721–725. IEEE, 2020. 5
work page 2020
Show all 66 references
-
[9]
Sdae: Self- distillated masked autoencoder
Yabo Chen, Yuchen Liu, Dongsheng Jiang, Xiaopeng Zhang, Wenrui Dai, Hongkai Xiong, and Qi Tian. Sdae: Self- distillated masked autoencoder. In European Conference on Computer Vision, pages 108–124. Springer, 2022. 1
2022
-
[10]
Similarity contrastive estima- tion for self-supervised soft contrastive learning
Julien Denize, Jaonary Rabarisoa, Astrid Orcesi, Romain H´erault, and St ´ephane Canu. Similarity contrastive estima- tion for self-supervised soft contrastive learning. InProceed- ings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 2706–2716, 2023. 1
2023
-
[11]
An image is worth 16x16 words: Trans- formers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, et al. An image is worth 16x16 words: Trans- formers for image recognition at scale. arXiv preprint a...
2010 arXiv
-
[12]
A large-scale study on unsupervised spatiotemporal representation learning
Christoph Feichtenhofer, Haoqi Fan, Bo Xiong, Ross Gir- shick, and Kaiming He. A large-scale study on unsupervised spatiotemporal representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3299–3309, 2021. 1
2021
-
[13]
Masked autoencoders as spatiotemporal learners
Christoph Feichtenhofer, Yanghao Li, Kaiming He, et al. Masked autoencoders as spatiotemporal learners. Advances in neural information processing systems, 35:35946–35958,
-
[14]
Audio set: An ontology and human- labeled dataset for audio events
Jort F Gemmeke, Daniel PW Ellis, Dylan Freedman, Aren Jansen, Wade Lawrence, R Channing Moore, Manoj Plakal, and Marvin Ritter. Audio set: An ontology and human- labeled dataset for audio events. In 2017 IEEE interna- tional conference on acoustics, speech and signal processin...
2017
-
[15]
Audiovisual masked autoencoders
Mariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu, Mario Lucic, Cordelia Schmid, and Anurag Arnab. Audiovisual masked autoencoders. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 16144–16154, 2023. 1
2023
-
[16]
Contrastive audio-visual masked autoencoder.arXiv preprint arXiv:2210.07839, 2022
Yuan Gong, Andrew Rouditchenko, Alexander H Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James Glass. Contrastive audio-visual masked autoencoder.arXiv preprint arXiv:2210.07839, 2022. 1, 2, 4
2022 arXiv
-
[17]
Bootstrap your own latent-a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altch ´e, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Ghesh- laghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. Advances in neur...
2020
-
[18]
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyv ¨arinen. Noise-contrastive estimation: A new estimation principle for unnormalized statistical models. In Proceedings of the thirteenth inter- national conference on artificial intelligence and statistics , pages 297–304. JMLR Workshop and Conferen...
2010
-
[19]
Self- supervised co-training for video representation learning
Tengda Han, Weidi Xie, and Andrew Zisserman. Self- supervised co-training for video representation learning. Ad- vances in Neural Information Processing Systems, 33:5679– 5690, 2020. 1
2020
-
[20]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 16000– 16009, 2022. 1, 4
2022
-
[21]
Mgmae: Motion guided masking for video masked autoencoding
Bingkun Huang, Zhiyu Zhao, Guozhen Zhang, Yu Qiao, and Limin Wang. Mgmae: Motion guided masking for video masked autoencoding. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , pages 13493– 13504, 2023. 1
2023
-
[22]
Let there be color! joint end-to-end learning of global and local image priors for automatic image colorization with simulta- neous classification
Satoshi Iizuka, Edgar Simo-Serra, and Hiroshi Ishikawa. Let there be color! joint end-to-end learning of global and local image priors for automatic image colorization with simulta- neous classification. ACM Transactions on Graphics (ToG), 35(4):1–11, 2016. 1
2016
-
[23]
A survey on contrastive self-supervised learning.Technologies, 9(1):2,
Ashish Jaiswal, Ashwin Ramesh Babu, Mohammad Zaki Zadeh, Debapriya Banerjee, and Fillia Makedon. A survey on contrastive self-supervised learning.Technologies, 9(1):2,
-
[24]
The kinetics hu- man action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, 9 Tim Green, Trevor Back, Paul Natsev, et al. The kinetics hu- man action video dataset. arXiv preprint arXiv:1705.06950,
-
[25]
Learning image representations by completing dam- aged jigsaw puzzles
Dahun Kim, Donghyeon Cho, Donggeun Yoo, and In So Kweon. Learning image representations by completing dam- aged jigsaw puzzles. In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV) , pages 793–802. IEEE, 2018. 1
2018
-
[26]
Temporally coherent embeddings for self-supervised video representation learning
Joshua Knights, Ben Harwood, Daniel Ward, Anthony Van- derkop, Olivia Mackenzie-Ross, and Peyman Moghadam. Temporally coherent embeddings for self-supervised video representation learning. In 2020 25th International Con- ference on Pattern Recognition (ICPR) , pages 8914–8921....
2020
-
[27]
Self-supervised video similarity learning
Giorgos Kordopatis-Zilos, Giorgos Tolias, Christos Tzelepis, Ioannis Kompatsiaris, Ioannis Patras, and Symeon Pa- padopoulos. Self-supervised video similarity learning. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 4755–4765, 2023. 1
2023
-
[28]
Learning representations for automatic colorization
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich. Learning representations for automatic colorization. In Computer Vision–ECCV 2016: 14th Eu- ropean Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14 , pages 577–593. Springer, 2016. 1
2016
-
[29]
Colorization as a proxy task for visual understanding
Gustav Larsson, Michael Maire, and Gregory Shakhnarovich. Colorization as a proxy task for visual understanding. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 6874–6883,
-
[30]
Predicting what you already know helps: Provable self- supervised learning
Jason D Lee, Qi Lei, Nikunj Saunshi, and Jiacheng Zhuo. Predicting what you already know helps: Provable self- supervised learning. Advances in Neural Information Pro- cessing Systems, 34:309–323, 2021. 1
2021
-
[31]
Pytorch distributed: Expe- riences on accelerating data parallel training
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, et al. Pytorch distributed: Expe- riences on accelerating data parallel training. arXiv preprint arXiv:2006.15704, 2020. 5
2006 arXiv
-
[32]
Audio self-supervised learning: A survey
Shuo Liu, Adria Mallol-Ragolta, Emilia Parada-Cabaleiro, Kun Qian, Xin Jing, Alexander Kathan, Bin Hu, and Bjo- ern W Schuller. Audio self-supervised learning: A survey. Patterns, 3(12), 2022. 1
2022
-
[33]
Temporal contrastive pretrain- ing for video action recognition
Guillaume Lorre, Jaonary Rabarisoa, Astrid Orcesi, Samia Ainouz, and Stephane Canu. Temporal contrastive pretrain- ing for video action recognition. In Proceedings of the IEEE/CVF winter conference on applications of computer vision, pages 662–670, 2020. 1
2020
-
[34]
Learning word em- beddings efficiently with noise-contrastive estimation
Andriy Mnih and Koray Kavukcuoglu. Learning word em- beddings efficiently with noise-contrastive estimation. Ad- vances in neural information processing systems , 26, 2013. 1
2013
-
[35]
Joint self-supervised image-volume representation learning with intra-inter con- trastive clustering
Duy MH Nguyen, Hoang Nguyen, Truong TN Mai, Tri Cao, Binh T Nguyen, Nhat Ho, Paul Swoboda, Shadi Albarqouni, Pengtao Xie, and Daniel Sonntag. Joint self-supervised image-volume representation learning with intra-inter con- trastive clustering. In Proceedings of the AAAI Confer...
-
[36]
Unsupervised learning of visual representations by solving jigsaw puzzles
Mehdi Noroozi and Paolo Favaro. Unsupervised learning of visual representations by solving jigsaw puzzles. In Euro- pean conference on computer vision, pages 69–84. Springer,
-
[37]
Boosting self-supervised learning via knowledge transfer
Mehdi Noroozi, Ananth Vinjimoor, Paolo Favaro, and Hamed Pirsiavash. Boosting self-supervised learning via knowledge transfer. In Proceedings of the IEEE conference on computer vision and pattern recognition , pages 9359– 9367, 2018. 1
2018
-
[38]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 1
2023 arXiv
-
[39]
Spatiotempo- ral contrastive video representation learning
Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, and Yin Cui. Spatiotempo- ral contrastive video representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 6964–6974, 2021. 1
2021
-
[40]
Self-taught learning: transfer learning from unlabeled data
Rajat Raina, Alexis Battle, Honglak Lee, Benjamin Packer, and Andrew Y Ng. Self-taught learning: transfer learning from unlabeled data. InProceedings of the 24th international conference on Machine learning, pages 759–766, 2007. 1
2007
-
[41]
Masked jigsaw puzzle: A versatile po- sition embedding for vision transformers
Bin Ren, Yahui Liu, Yue Song, Wei Bi, Rita Cucchiara, Nicu Sebe, and Wei Wang. Masked jigsaw puzzle: A versatile po- sition embedding for vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20382–20391, 2023. 1
2023
-
[42]
Berg, and Li Fei-Fei
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, San- jeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Chal- lenge. International Journal of Computer Vision ...
2015
-
[43]
Puzzle- ae: Novelty detection in images through solving puzzles
Mohammadreza Salehi, Ainaz Eftekhar, Niousha Sadjadi, Mohammad Hossein Rohban, and Hamid R Rabiee. Puzzle- ae: Novelty detection in images through solving puzzles. arXiv preprint arXiv:2008.12959, 2020. 1
2008 arXiv
-
[44]
Self-supervised learning for videos: A survey
Madeline C Schiappa, Yogesh S Rawat, and Mubarak Shah. Self-supervised learning for videos: A survey. ACM Com- puting Surveys, 55(13s):1–37, 2023. 1
2023
-
[45]
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudi- nov. Unsupervised learning of video representations using lstms. In International conference on machine learning , pages 843–852. PMLR, 2015. 1
2015
-
[46]
Self- supervised video representation learning using inter-intra contrastive framework
Li Tao, Xueting Wang, and Toshihiko Yamasaki. Self- supervised video representation learning using inter-intra contrastive framework. In Proceedings of the 28th ACM International Conference on Multimedia, pages 2193–2201,
-
[47]
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Advances in neural information processing systems, 35:10078–10093, 2022. 1
2022
-
[48]
Mocogan: Decomposing motion and content for 10 video generation
Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz. Mocogan: Decomposing motion and content for 10 video generation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 1526–1535,
-
[49]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. 2
2017
-
[50]
Generating videos with scene dynamics
Carl V ondrick, Hamed Pirsiavash, and Antonio Torralba. Generating videos with scene dynamics. Advances in neu- ral information processing systems, 29, 2016. 1
2016
-
[51]
Video anomaly detection by solving decoupled spatio-temporal jigsaw puzzles
Guodong Wang, Yunhong Wang, Jie Qin, Dongming Zhang, Xiuguo Bao, and Di Huang. Video anomaly detection by solving decoupled spatio-temporal jigsaw puzzles. In Eu- ropean Conference on Computer Vision , pages 494–511. Springer, 2022. 1
2022
-
[52]
Enhancing unsupervised video representation learning by decoupling the scene and the motion
Jinpeng Wang, Yuting Gao, Ke Li, Jianguo Hu, Xinyang Jiang, Xiaowei Guo, Rongrong Ji, and Xing Sun. Enhancing unsupervised video representation learning by decoupling the scene and the motion. In Proceedings of the AAAI Con- ference on Artificial Intelligence , volume 35, page...
2021
-
[53]
Videomae v2: Scaling video masked autoencoders with dual masking
Limin Wang, Bingkun Huang, Zhiyu Zhao, Zhan Tong, Yi- nan He, Yi Wang, Yali Wang, and Yu Qiao. Videomae v2: Scaling video masked autoencoders with dual masking. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 14549–14560, 2023. 1
2023
-
[54]
Knowledge distillation and student-teacher learning for visual intelligence: A review and new outlooks
Lin Wang and Kuk-Jin Yoon. Knowledge distillation and student-teacher learning for visual intelligence: A review and new outlooks. IEEE transactions on pattern analysis and machine intelligence, 44(6):3048–3068, 2021. 1
2021
-
[55]
Bevt: Bert pretraining of video transformers
Rui Wang, Dongdong Chen, Zuxuan Wu, Yinpeng Chen, Xiyang Dai, Mengchen Liu, Yu-Gang Jiang, Luowei Zhou, and Lu Yuan. Bevt: Bert pretraining of video transformers. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 14733–14743, 2022. 1
2022
-
[56]
Masked feature predic- tion for self-supervised visual pre-training
Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan Yuille, and Christoph Feichtenhofer. Masked feature predic- tion for self-supervised visual pre-training. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14668–14678, 2022. 1
2022
-
[57]
Iterative reorganiza- tion with weak spatial constraints: Solving arbitrary jigsaw puzzles for unsupervised representation learning
Chen Wei, Lingxi Xie, Xutong Ren, Yingda Xia, Chi Su, Jiaying Liu, Qi Tian, and Alan L Yuille. Iterative reorganiza- tion with weak spatial constraints: Solving arbitrary jigsaw puzzles for unsupervised representation learning. In Pro- ceedings of the IEEE/CVF Conference on Co...
1910
-
[58]
Structured sparsity learning for efficient video super-resolution
Bin Xia, Jingwen He, Yulun Zhang, Yitong Wang, Yapeng Tian, Wenming Yang, and Luc Van Gool. Structured sparsity learning for efficient video super-resolution. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 22638–22647, 2023. 1
2023
-
[59]
Self- supervised video representation learning with motion-aware masked autoencoders
Haosen Yang, Deng Huang, Bin Wen, Jiannan Wu, Hongxun Yao, Yi Jiang, Xiatian Zhu, and Zehuan Yuan. Self- supervised video representation learning with motion-aware masked autoencoders. arXiv preprint arXiv:2210.04154 ,
-
[60]
Seco: Exploring sequence supervision for un- supervised representation learning
Ting Yao, Yiheng Zhang, Zhaofan Qiu, Yingwei Pan, and Tao Mei. Seco: Exploring sequence supervision for un- supervised representation learning. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 10656–10664, 2021. 1
2021
-
[61]
Contextualized spatio-temporal contrastive learning with self-supervision
Liangzhe Yuan, Rui Qian, Yin Cui, Boqing Gong, Flo- rian Schroff, Ming-Hsuan Yang, Hartwig Adam, and Ting Liu. Contextualized spatio-temporal contrastive learning with self-supervision. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pag...
2022
-
[62]
A survey on masked au- toencoder for visual self-supervised learning
Chaoning Zhang, Chenshuang Zhang, Junha Song, John Seon Keun Yi, and In So Kweon. A survey on masked au- toencoder for visual self-supervised learning. InProceedings of the Thirty-Second International Joint Conference on Arti- ficial Intelligence, pages 6805–6813, 2023. 1
2023
-
[63]
A survey on masked autoencoder for self-supervised learning in vision and beyond
Chaoning Zhang, Chenshuang Zhang, Junha Song, John Seon Keun Yi, Kang Zhang, and In So Kweon. A survey on masked autoencoder for self-supervised learning in vision and beyond. arXiv preprint arXiv:2208.00173, 2022. 1
2022 arXiv
-
[64]
How mask mat- ters: Towards theoretical understandings of masked autoen- coders
Qi Zhang, Yifei Wang, and Yisen Wang. How mask mat- ters: Towards theoretical understandings of masked autoen- coders. Advances in Neural Information Processing Systems, 35:27127–27139, 2022. 1
2022
-
[65]
Deep mutual learning
Ying Zhang, Tao Xiang, Timothy M Hospedales, and Huchuan Lu. Deep mutual learning. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 4320–4328, 2018. 1, 2, 4, 5
2018
-
[66]
Masked autoencoders in computer vision: A comprehensive survey
Zexian Zhou and Xiaojing Liu. Masked autoencoders in computer vision: A comprehensive survey. IEEE Access ,
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.