REVIEW 4 major objections 4 minor 66 references
Human Gaze Boosts Object-Centered Representation Learning
T0 review · 4 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Self-supervised models learn more object-centered representations when trained on gaze-centered crops of egocentric video than on the full visual field.
desk verdict Gaze-centered crops are a simple, plausible win for egocentric SSL, and the paper's central comparison is consistent across datasets; the main unresolved issue is statistical selection and reporting, not the core idea. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a gaze-centered cropping pipeline feeding a time-augmented momentum-contrast learner. For each frame, a spatio-temporal gaze-estimation model supplies a predicted fixation point, and a square crop of side N (usually 224–336 pixels) is shifted to stay inside the frame, simulating central vision. The learner is a variant of MoCoV3, a momentum-contrast self-supervised model, adapted to sample an indirect temporal neighbor within a window ΔT from the same video and minimize an InfoNCE loss that aligns close-in-time views, enforcing slowly changing representations. The two dials are crop size N and window ΔT: intermediate crops give object-centered representations, and a nonzero ΔT is required for the gain.
What would settle it
Retrain the same model with a graded fovea (sharp center, smoothly blurred periphery) or with ground-truth eye-tracker gaze instead of predicted gaze; if the accuracy gains on ImageNet-1k and Core50 vanish or reverse, the claim that central vision itself drives the improvement is falsified.
Extended reading notes
Core claim
The paper claims, on its own terms, that feeding an SSL learner the visual sequence captured by central vision—a square crop around a predicted gaze point—rather than the full egocentric frame produces better object-centered representations. In a single-epoch pre-training run on Ego4D with a time-augmented MoCoV3 model, gaze-centered training raises ImageNet-1k linear-probe accuracy from 48.982 to 50.572, average fine-grained recognition from 33.761 to 37.854, and average instance recognition from 64.691 to 67.556, while easy-category and scene benchmarks do not improve. An ImageNet-9 analysis attributes this to reduced background sensitivity: intermediate crop sizes make the model rely more on the foreground object and less on the background. A further comparison shows that gaze-based crops beat fixed center crops, especially for instance recognition under changing backgrounds, and that temporal slowness is critical—with no temporal window the improvements largely disappear.
Load-bearing premise
The argument treats central vision as a hard square crop around a predicted gaze point, so if real foveal acuity is graded or the gaze estimates are systematically wrong, the reported benefits could change or disappear.
Editorial extensions
If this is right
- If the claim holds, egocentric SSL pipelines should crop training inputs around gaze rather than use full frames, since the change improves hard category, fine-grained, and instance recognition on nearly every benchmark tested.
- Central-vision training should be especially useful where background is misleading, because the ImageNet-9 results show it reduces background sensitivity at intermediate crop sizes.
- Temporal slowness is a necessary ingredient: with no temporal window (ΔT = 0), recognition scores drop across all semantic groups, so gaze-centered crops alone are not enough.
- The effect is not just about removing peripheral content: gaze-based crops outperform fixed center crops, with the largest margin on Core50 instance recognition under changing backgrounds.
- Crop size has a sweet spot around N = 224–336; very small crops (N = 112) hurt recognition, while full frames favor scene over object representations.
Reading between the lines
- A natural testable extension is replacing the hard square crop with a graded foveal filter, which might preserve the object-centered boost while giving the model peripheral context; the paper's own discussion flags this as future work.
- Because gaze is predicted rather than measured on most of the data, the comparison against the small eye-tracked subset could show whether the reported gains are dampened by gaze-estimation error.
- The same central-vision bias could transfer to embodied agents: an agent that actively looks at objects and learns from gaze-centered views may build object representations from fewer interactions than one trained on uniform frames.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether an SSL model pre-trained on egocentric video should be trained on the whole field of view or on a gaze-centered crop, motivated by the anatomical magnification of central vision in humans. Using Ego4D, the authors predict gaze with GLC, crop N×N patches around the predicted fixation, and train a temporal variant of MoCoV3 (InfoNCE between temporally nearby frames) for one epoch. They evaluate with linear probes on seventeen datasets grouped into hard/easy category, fine-grained, instance, and scene recognition. The headline result is that central-vision pre-training improves hard-category, fine-grained, and instance recognition relative to full-field pre-training, with an ImageNet-9 analysis attributing the gain to reduced background sensitivity. Additionally, temporal slowness is shown to be important, and gaze-based crops outperform center crops on most groups.
Significance. If the central comparison is robust, the paper provides a simple and biologically motivated training intervention that yields consistent gains across diverse object-centric tasks, with a plausible mechanism (reduced background reliance). The study's breadth—many downstream datasets, a large real-world egocentric corpus, and a foreground/background sensitivity analysis—is a strength, as is the explicit gaze-vs-center-crop control. The main unaddressed threats are statistical: single-run comparisons, no measure of run-to-run variability, and hyperparameters (crop size and temporal window) chosen with knowledge of the same evaluation sets. Because several headline gains are small, the result is not yet established to the standard the central claim requires.
major comments (4)
- [Section 4.1, Table 1] All reported accuracies are single runs; the paper gives no standard deviations, seed count, or significance statements. Several differences that support the central claim are small: ImageNet-1k 100% +1.59, ImageNet-1k 10% +1.01, ImageNet-100 +0.56, CIFAR100 +0.77. Without repeated pre-training runs with different seeds, or at least a variance estimate over linear-probe seeds, the observed improvement could be within run-to-run noise, especially for these small effects. This is load-bearing because Table 1 is the evidence for the paper's central claim.
- [Sections 4.2 and 4.3, Figures 2 and 4] The crop size N and temporal window ΔT are selected after observing performance on the same evaluation datasets: Figure 2 identifies N=336 for hard/instance and N=224 for easy/fine-grained, and Figure 4 identifies group-specific best ΔT values, and the paper then reports results using those choices. Since the downstream datasets are also used to compute the reported gains, this selection can inflate the central-vision advantage. Please provide a protocol in which hyperparameters are chosen on a development set or held-out datasets, or report all configurations side by side; at minimum, state explicitly which N and ΔT were used for each table entry.
- [Section 4.1, Table 1 and Figure 2] Table 1 does not state the crop size N used for the 'Central vision' column, although Figure 2 shows the result depends strongly on N. Moreover, for easy category recognition, Table 1 shows central vision is worse on both STL10 and CIFAR10 (71.514 vs 71.689 and 78.654 vs 79.574), so the text's blanket statement that 'focusing on central vision leads to better object-centered representations' is only supported for the hard/fine-grained/instance groups. Please state N and temper the summary accordingly.
- [Section 3.1] Most training data uses GLC-predicted gaze; only 45 hours have ground-truth gaze. The paper does not report GLC accuracy on this data or compare models trained with ground-truth versus predicted gaze. If predicted gaze is biased (e.g., toward image center, as Appendix B suggests), the gaze-centered crop may behave differently from human central vision. A comparison on the ground-truth subset or a corruption analysis would strengthen the interpretation that 'human gaze' rather than a generic center-biased crop drives the effect.
minor comments (4)
- [Section 3.1] The displayed formula for (x_cor, y_cor) has an unbalanced parenthesis and appears to mix the right-boundary and left-boundary corrections incorrectly; please write it with explicit max/min terms for each boundary.
- [Appendix, Table 4 and Table 5] The appendix table captions and cross-references are inconsistent: Table 4 is captioned as detailed results of Figure 2 but actually reports the temporal-window sweep, the text refers to 'Appendix Table C' when it should refer to Table 4, and Table 5's header says 'Places375' instead of 'Places365'.
- [Figures 2 and 4] The y-axis label 'Accuracy recognition improvement' is awkward; consider 'Improvement in recognition accuracy.'
- [Section 3.1] The sentence 'Our final preprocessed dataset contains 64,380,024 images' does not specify whether this count is before or after cropping, and it would be helpful to report the number of unique clips and frames used for training.
Circularity Check
No significant circularity: the paper's central claim is an empirical comparison benchmarked on external datasets, and no derivation reduces to its own inputs.
full rationale
The paper's central claim, that focusing on central vision boosts object-centered representation learning, is supported by linear-probe accuracies on external benchmark datasets (ImageNet-1k, CIFAR, fine-grained and instance recognition sets) after pre-training on Ego4D with gaze-centered crops. There is no equation or construction by which the reported outcome is defined in terms of the input manipulation: the gaze-crop preprocessing is a data transformation, and the evaluation uses held-out test splits of standard benchmarks. The choice of crop size N in Figure 2 and temporal window ΔT in Figure 4 is a hyperparameter sweep on evaluation sets, which raises a statistical robustness concern about selection-induced inflation, but it is not circularity because the final accuracies are measured quantities, not algebraic consequences of the sweep. The self-citations to [2,42] for temporal augmentation are methodological borrowings from the authors' prior work; they are not invoked to justify the central empirical finding, and the temporal-slowness component is itself tested against the ΔT = 0 baseline rather than assumed. No uniqueness theorem, fitted parameter renamed as a prediction, or ansatz smuggled through citation appears in the paper. Accordingly, the derivation chain is self-contained and the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (2)
- crop size N =
224 (default); sweeps up to 448
- temporal window Delta-T =
15 seconds (default); sweeps 0-5 seconds in Figure 4
assumptions (4)
- domain assumption Square crop centered on the predicted gaze location is an adequate model of human central vision
- domain assumption GLC-predicted gaze locations are accurate proxies for human gaze over the full Ego4D corpus
- domain assumption Ego4D head-mounted video approximates a human's daily egocentric visual experience (about 5 months)
- standard math MoCoV3 with temporal slow-feature alignment is a valid SSL model for egocentric video
Cite this review
Pith. "Pith review of Human Gaze Boosts Object-Centered Representation Learning." pith.science (2026). https://pith.science/paper/HXDJDB2N
@misc{pith2026250102966,
author = {Pith},
title = {Pith review of: Human Gaze Boosts Object-Centered Representation Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/HXDJDB2N}},
note = {Machine review of arXiv:2501.02966}
}
read the original abstract
Recent self-supervised learning (SSL) models trained on human-like egocentric visual inputs substantially underperform on image recognition tasks compared to humans. These models train on raw, uniform visual inputs collected from head-mounted cameras. This is different from humans, as the anatomical structure of the retina and visual cortex relatively amplifies the central visual information, i.e. around humans' gaze location. This selective amplification in humans likely aids in forming object-centered visual representations. Here, we investigate whether focusing on central visual information boosts egocentric visual object learning. We simulate 5-months of egocentric visual experience using the large-scale Ego4D dataset and generate gaze locations with a human gaze prediction model. To account for the importance of central vision in humans, we crop the visual area around the gaze location. Finally, we train a time-based SSL model on these modified inputs. Our experiments demonstrate that focusing on central vision leads to better object-centered representations. Our analysis shows that the SSL model leverages the temporal dynamics of the gaze movements to build stronger visual representations. Overall, our work marks a significant step toward bio-inspired learning of visual representations.
Figures
Reference graph
Works this paper leans on
-
[1]
Chart demonstrating variations in acuity with retinal position
Stuart M Anstis. Chart demonstrating variations in acuity with retinal position. Vision research, 14(7):589–592, 1974. 1
work page 1974
-
[2]
Time to augment self-supervised visual representation learning
Arthur Aubret, Markus Ernst, C ´eline Teuli`ere, and Jochen Triesch. Time to augment self-supervised visual represen- tation learning. arXiv preprint arXiv:2207.13492, 2022. 2, 3
work page Pith review arXiv 2022
-
[3]
Toddler- inspired embodied vision for learning object representations
Arthur Aubret, C ´eline Teuli`er, and Jochen Triesch. Toddler- inspired embodied vision for learning object representations. In 2022 IEEE International Conference on Development and Learning (ICDL), pages 81–87. IEEE, 2022. 1
work page 2022
-
[4]
Learning Object Semantic Similarity with Self-Supervision
Arthur Aubret, Timothy Schauml ¨offel, Gemma Roig, and Jochen Triesch. Learning object semantic similarity with self-supervision. arXiv preprint arXiv:2405.05143, 2024. 2
work page Pith review arXiv 2024
-
[5]
Self-supervised visual learning from interactions with objects
Arthur Aubret, C ´eline Teuli`ere, and Jochen Triesch. Self- supervised visual learning from interactions with objects. arXiv preprint arXiv:2407.06704, 2024. 2
work page Pith review arXiv 2024
-
[6]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021. 2
2021
-
[7]
Big self-supervised mod- els are strong semi-supervised learners
Ting Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi, and Geoffrey E Hinton. Big self-supervised mod- els are strong semi-supervised learners. Advances in neural information processing systems, 33:22243–22255, 2020. 4, 11
work page 2020
-
[8]
An empirical study of training self-supervised vision transformers
Xinlei Chen, Saining Xie, and Kaiming He. An empirical study of training self-supervised vision transformers. InPro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9640–9649, 2021. 3
work page 2021
Show all 66 references
-
[9]
Describing textures in the wild
Mircea Cimpoi, Subhransu Maji, Iasonas Kokkinos, Sammy Mohamed, and Andrea Vedaldi. Describing textures in the wild. In 2014 IEEE Conference on Computer Vision and Pattern Recognition, pages 3606–3613, 2014. 4, 11
2014
-
[10]
An analysis of single-layer networks in unsupervised feature learning
Adam Coates, Andrew Ng, and Honglak Lee. An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics , pages 215–223. JMLR Workshop and Conference Proceedings, 2011. 4, 11
2011
-
[11]
solo-learn: A library of self- supervised methods for visual representation learning
Victor Guilherme Turrisi da Costa, Enrico Fini, Moin Nabi, Nicu Sebe, and Elisa Ricci. solo-learn: A library of self- supervised methods for visual representation learning. Jour- nal of Machine Learning Research, 23(56):1–6, 2022. 4
2022
-
[12]
Hvm-1: Large-scale video models pretrained with nearly 5000 hours of human-like video data
A Emin Orhan. Hvm-1: Large-scale video models pretrained with nearly 5000 hours of human-like video data. arXiv e- prints, pages arXiv–2407, 2024. 2
2024
-
[13]
Learning invariance from transformation se- quences
Peter F ¨oldi´ak. Learning invariance from transformation se- quences. Neural computation, 3(2):194–200, 1991. 2
1991
-
[14]
Partial success in closing the gap between human and machine vision
Robert Geirhos, Kantharaju Narayanappa, Benjamin Mitzkus, Tizian Thieringer, Matthias Bethge, Felix A Wichmann, and Wieland Brendel. Partial success in closing the gap between human and machine vision. Advances in Neural Information Processing Systems , 34:23885–23899,
-
[15]
Watching the world go by: Representation learning from un- labeled videos
Daniel Gordon, Kiana Ehsani, Dieter Fox, and Ali Farhadi. Watching the world go by: Representation learning from un- labeled videos. arXiv preprint arXiv:2003.07990, 2020. 2
2003 arXiv
-
[16]
Girshick, Pieter No- ordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tul- 7 loch, Yangqing Jia, and Kaiming He
Priya Goyal, Piotr Doll ´ar, Ross B. Girshick, Pieter No- ordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tul- 7 loch, Yangqing Jia, and Kaiming He. Accurate, large mini- batch sgd: Training imagenet in 1 hour. ArXiv preprint , abs/1706.02677, 2017. 4
2017 arXiv
-
[17]
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Mar- tin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Meng...
2022
-
[18]
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. InPro- ceedings of the IEEE/CVF Conference on Computer Vision...
2022
-
[19]
The visual experience dataset: Over 200 recorded hours of integrated eye movement, odometry, and egocentric video
Michelle R Greene, Benjamin J Balas, Mark D Lescroart, Paul R MacNeilage, Jennifer A Hart, Kamran Binaee, Pe- ter A Hausamann, Ronald Mezile, Bharath Shankar, Chris- tian B Sinnott, et al. The visual experience dataset: Over 200 recorded hours of integrated eye movement, odome...
2024 arXiv
-
[20]
Bootstrap your own latent-a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altch ´e, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Doersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Ghesh- laghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. Advances in neur...
2020
-
[21]
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. 4
2016
-
[22]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 16000– 16009, 2022. 2
2022
-
[23]
Space-time correspondence as a contrastive random walk
Allan Jabri, Andrew Owens, and Alexei Efros. Space-time correspondence as a contrastive random walk. Advances in neural information processing systems, 33:19545–19560,
-
[24]
Learning image representations tied to ego-motion
Dinesh Jayaraman and Kristen Grauman. Learning image representations tied to ego-motion. In Proceedings of the IEEE International Conference on Computer Vision , pages 1413–1421, 2015. 2
2015
-
[25]
Slow and steady feature analysis: higher order temporal coherence in video
Dinesh Jayaraman and Kristen Grauman. Slow and steady feature analysis: higher order temporal coherence in video. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 3852–3861, 2016. 2
2016
-
[26]
3d object representations for fine-grained categorization
Jonathan Krause, Michael Stark, Jia Deng, and Li Fei-Fei. 3d object representations for fine-grained categorization. In 2013 IEEE International Conference on Computer Vision Workshops, pages 554–561, 2013. 4, 11
2013
-
[27]
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 4, 11
2009
-
[28]
In the eye of transformer: Global-local correlation for egocentric gaze estimation
Bolin Lai, Miao Liu, Fiona Ryan, and James Rehg. In the eye of transformer: Global-local correlation for egocentric gaze estimation. In 33rd British Machine Vision Conference 2022, BMVC 2022, London, UK, November 21-24, 2022 . BMV A Press, 2022. 2, 3
2022
-
[29]
Core50: a new dataset and benchmark for continuous object recognition
Vincenzo Lomonaco and Davide Maltoni. Core50: a new dataset and benchmark for continuous object recognition. In Conference on robot learning, pages 17–26. PMLR, 2017. 4, 11
2017
-
[30]
The babyview dataset: High-resolution egocentric videos of in- fants’ and young children’s everyday experiences
Bria Long, Violet Xiang, Stefan Stojanov, Robert Z Sparks, Zi Yin, Grace E Keene, Alvin WM Tan, Steven Y Feng, Chengxu Zhuang, Virginia A Marchman, et al. The babyview dataset: High-resolution egocentric videos of in- fants’ and young children’s everyday experiences. arXiv pre...
2024 arXiv
-
[31]
Nymeria: A massive collection of multimodal egocentric daily motion in the wild
Lingni Ma, Yuting Ye, Fangzhou Hong, Vladimir Guzov, Yifeng Jiang, Rowan Postyeni, Luis Pesqueira, Alexander Gamino, Vijay Baiyya, Hyo Jin Kim, et al. Nymeria: A massive collection of multimodal egocentric daily motion in the wild. arXiv preprint arXiv:2406.09905, 2024. 2
2024 arXiv
-
[32]
Vip: Towards universal visual reward and representation via value-implicit pre-training
Yecheng Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Os- bert Bastani, Vikash Kumar, and Amy Zhang. Vip: Towards universal visual reward and representation via value-implicit pre-training. In The Eleventh International Conference on Learning Representations. 2
-
[33]
Blaschko, and Andrea Vedaldi
Subhransu Maji, Esa Rahtu, Juho Kannala, Matthew B. Blaschko, and Andrea Vedaldi. Fine-grained visual classi- fication of aircraft. Technical Report, abs/1306.5151, 2013. 4, 11
2013 arXiv
-
[34]
Where are we in the search for an artificial visual cortex for embodied intelli- gence? Advances in Neural Information Processing Systems, 36:655–677, 2023
Arjun Majumdar, Karmesh Yadav, Sergio Arnaud, Jason Ma, Claire Chen, Sneha Silwal, Aryan Jain, Vincent-Pierre Berges, Tingfan Wu, Jay Vakil, et al. Where are we in the search for an artificial visual cortex for embodied intelli- gence? Advances in Neural Information Processing...
2023
-
[35]
R3m: A universal visual repre- sentation for robot manipulation
Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta. R3m: A universal visual repre- sentation for robot manipulation. In 6th Annual Conference on Robot Learning. 2
-
[36]
Columbia object image library (coil-20)
Sameer A Nene, Shree K Nayar, Hiroshi Murase, et al. Columbia object image library (coil-20). 1996. 4, 11 8
1996
-
[37]
Automated flower classification over a large number of classes
Maria-Elena Nilsback and Andrew Zisserman. Automated flower classification over a large number of classes. In 2008 Sixth Indian Conference on Computer Vision, Graphics & Image Processing, pages 722–729, 2008. 4, 11
2008
-
[38]
Scaling may be all you need for achieving human-level object recognition capacity with human-like vi- sual experience
A Emin Orhan. Scaling may be all you need for achieving human-level object recognition capacity with human-like vi- sual experience. arXiv preprint arXiv:2308.03712, 2023. 1, 7
2023 arXiv
-
[39]
Learning high-level vi- sual representations from a child’s perspective without strong inductive biases
A Emin Orhan and Brenden M Lake. Learning high-level vi- sual representations from a child’s perspective without strong inductive biases. Nature Machine Intelligence , 6(3):271– 283, 2024. 1, 5, 10
2024
-
[40]
Self-supervised learning of video representations from a child’s perspective
A Emin Orhan, Wentao Wang, Alex N Wang, Mengye Ren, and Brenden M Lake. Self-supervised learning of video representations from a child’s perspective. arXiv preprint arXiv:2402.00300, 2024. 2
2024 arXiv
-
[41]
Self- supervised learning through the eyes of a child
Emin Orhan, Vaibhav Gupta, and Brenden M Lake. Self- supervised learning through the eyes of a child. Advances in Neural Information Processing Systems , 33:9960–9971,
-
[42]
Are vi- sion transformers more data hungry than newborn visual sys- tems? Advances in Neural Information Processing Systems, 36, 2024
Lalit Pandey, Samantha Wood, and Justin Wood. Are vi- sion transformers more data hungry than newborn visual sys- tems? Advances in Neural Information Processing Systems, 36, 2024. 2, 3
2024
-
[43]
Omkar M Parkhi, Andrea Vedaldi, Andrew Zisserman, and C. V . Jawahar. Cats and dogs. In 2012 IEEE Conference on Computer Vision and Pattern Recognition , pages 3498– 3505, 2012. 4, 11
2012
-
[44]
Self-supervised video pretraining yields robust and more human-aligned visual representations
Nikhil Parthasarathy, SM Eslami, Joao Carreira, and Olivier Henaff. Self-supervised video pretraining yields robust and more human-aligned visual representations. Advances in Neural Information Processing Systems , 36:65743–65765,
-
[45]
Object recognition in primates: What can early visual areas contribute? Fron- tiers in Behavioral Neuroscience, 18:1425496, 2024
Christian Quaia and Richard J Krauzlis. Object recognition in primates: What can early visual areas contribute? Fron- tiers in Behavioral Neuroscience, 18:1425496, 2024. 1
2024
-
[46]
Bottom-up saliency mod- els for still images: A practical review
Nicolas Riche and Matei Mancas. Bottom-up saliency mod- els for still images: A practical review. From Human At- tention to Computational Attention: A Multidisciplinary Ap- proach, pages 141–175, 2016. 3
2016
-
[47]
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, San- jeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al. Imagenet large scale visual recognition challenge. International journal of computer vision, 115:211–252, 2015. 1, 4, 7, 11
2015
-
[48]
Time does tell: Self-supervised time- tuning of dense image representations
Mohammadreza Salehi, Efstratios Gavves, Cees GM Snoek, and Yuki M Asano. Time does tell: Self-supervised time- tuning of dense image representations. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 16536–16547, 2023. 2
2023
-
[49]
A computational account of self-supervised visual learning from egocentric object play
Deepayan Sanyal, Joel Michelson, Yuan Yang, James Ain- ooson, and Maithilee Kunda. A computational account of self-supervised visual learning from egocentric object play. In Proceedings of the Annual Meeting of the Cognitive Sci- ence Society, 2023. 2
2023
-
[50]
Caregiver talk shapes toddler vision: A computational study of dyadic play
Timothy Schauml ¨offel, Arthur Aubret, Gemma Roig, and Jochen Triesch. Caregiver talk shapes toddler vision: A computational study of dyadic play. In 2023 IEEE Inter- national Conference on Development and Learning (ICDL), pages 67–72. IEEE, 2023. 2
2023
-
[51]
Time-contrastive networks: Self-supervised learn- ing from video
Pierre Sermanet, Corey Lynch, Yevgen Chebotar, Jasmine Hsu, Eric Jang, Stefan Schaal, Sergey Levine, and Google Brain. Time-contrastive networks: Self-supervised learn- ing from video. In 2018 IEEE international conference on robotics and automation (ICRA) , pages 1134–1141. IEEE,
2018
-
[52]
Curriculum learning with infant egocentric videos
Saber Sheybani, Himanshu Hansaria, Justin Wood, Linda Smith, and Zoran Tiganj. Curriculum learning with infant egocentric videos. Advances in Neural Information Process- ing Systems, 36, 2024. 2
2024
-
[53]
Saycam: A large, longitudinal audiovisual dataset recorded from the infant’s perspective
Jessica Sullivan, Michelle Mei, Andrew Perfors, Erica Wo- jcik, and Michael C Frank. Saycam: A large, longitudinal audiovisual dataset recorded from the infant’s perspective. Open mind, 5:20–29, 2021. 2
2021
-
[54]
Con- trastive multiview coding
Yonglong Tian, Dilip Krishnan, and Phillip Isola. Con- trastive multiview coding. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI 16 , pages 776–794. Springer,
2020
-
[55]
Repre- sentation learning with contrastive predictive coding, 2019
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. Repre- sentation learning with contrastive predictive coding, 2019. 4
2019
-
[56]
Is imagenet worth 1 video? learning strong image encoders from 1 long unlabelled video
Shashanka Venkataramanan, Mamshad Nayeem Rizve, Jo ˜ao Carreira, Yuki Asano, and Yannis Avrithis. Is imagenet worth 1 video? learning strong image encoders from 1 long unlabelled video. In ICLR 2024-Twelfth International Con- ference on Learning Representations, pages 1–21, 2024. 2
2024
-
[57]
On the use of cortical magnification and saccades as biological proxies for data augmentation
Binxu Wang, David Mayo, Arturo Deza, Andrei Barbu, and Colin Conwell. On the use of cortical magnification and saccades as biological proxies for data augmentation. In SVRHM 2021 Workshop@ NeurIPS, 2021. 7
2021
-
[58]
Eliott, James Ainooson, Joshua H
Xiaohan Wang, Fernanda M. Eliott, James Ainooson, Joshua H. Palmer, and Maithilee Kunda. An object is worth six thousand pictures: The egocentric, manual, multi-image (emmi) dataset. In The IEEE International Conference on Computer Vision (ICCV) Workshops, 2017. 4, 11
2017
-
[59]
Cortical magnification factor and the gan- glion cell density of the primate retina
Heinz W ¨assle, Ulrike Gr ¨unert, J ¨urgen R ¨ohrenbeck, and Brian B Boycott. Cortical magnification factor and the gan- glion cell density of the primate retina. Nature, 341(6243): 643–646, 1989. 1
1989
-
[60]
Slow feature analysis: Unsupervised learning of invariances
Laurenz Wiskott and Terrence J Sejnowski. Slow feature analysis: Unsupervised learning of invariances. Neural com- putation, 14(4):715–770, 2002. 2
2002
-
[61]
Noise or signal: The role of image backgrounds in object recognition
Kai Xiao, Logan Engstrom, Andrew Ilyas, and Aleksander Madry. Noise or signal: The role of image backgrounds in object recognition. In International Conference on Learning Representations, 2021. 5
2021
-
[62]
Rethinking self-supervised correspondence learning: A video frame-level similarity per- spective
Jiarui Xu and Xiaolong Wang. Rethinking self-supervised correspondence learning: A video frame-level similarity per- spective. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10075–10085, 2021. 2, 6
2021
-
[63]
Scaling SGD batch size to 32k for imagenet training
Yang You, Igor Gitman, and Boris Ginsburg. Scaling SGD batch size to 32k for imagenet training. ArXiv preprint , abs/1708.03888, 2017. 4 9
2017 arXiv
-
[64]
Representation of central and peripheral vision in the primate cerebral cortex: Insights from studies of the marmoset brain
H-H Yu, TA Chaplin, and MGP Rosa. Representation of central and peripheral vision in the primate cerebral cortex: Insights from studies of the marmoset brain. Neuroscience Research, 93:47–61, 2015. 1
2015
-
[65]
Places: A 10 million image database for scene recognition
Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba. Places: A 10 million image database for scene recognition. IEEE Transactions on Pattern Analy- sis and Machine Intelligence, 2017. 4, 11
2017
-
[66]
Mugs: A multi- granular self-supervised learning framework
Pan Zhou, Yichen Zhou, Chenyang Si, Weihao Yu, Teck Khim Ng, and Shuicheng Yan. Mugs: A multi- granular self-supervised learning framework. arXiv preprint arXiv:2203.14415, 2022. 2 A. Datasets In Table A, we present the datasets and benchmarks used in our experiments. The prov...
2022 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.