REVIEW 3 major objections 2 minor 7 cited by
Disentangling the Factors of Convergence between Brains and Computer Vision Models
T0 review · 3 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Larger self-supervised vision models trained on human-centric images align with the human visual cortex in a fixed order.
desk verdict A clean-sounding parametric sweep of DINOv3 against fMRI/MEG that I can't actually verify because the supplied full text is an unrelated acoustics paper; the abstract alone suggests an interesting chronology result. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is DINOv3, a self-supervised vision transformer that learns image representations from data without explicit labels, trained here in a systematically varied family (model size, training amount, image type). It is compared against human brain recordings from fMRI and MEG, which provide spatial and temporal resolution respectively, using three complementary metrics: overall representational similarity, topographical organization, and temporal dynamics. The argument is carried by the interaction of these measurement axes: factors that alter a model's alignment with early sensory areas are distinguishable from factors that shift alignment with late and prefrontal areas, and th
What would settle it
Train DINOv3 variants where the human-centric image set is replaced by images matched on low-level statistics (color histograms, contrast energy, spatial-frequency spectra) but with semantic content scrambled; if brain-similarity scores remain just as high, the convergence effects could come from low-level features, not semantic representational alignment.
Extended reading notes
Core claim
The central claim is that brain-model convergence is not a single monolithic phenomenon but the product of separable, interacting factors: architecture scale, training quantity, and image content. Each factor independently shapes, and jointly determines, how closely DINOv3 representations match human fMRI and MEG responses on three complementary metrics (overall representational similarity, topographic organization, and temporal dynamics). The same convergence has an ordered developmental trajectory: models first align with early sensory cortices and only align with late and prefrontal areas after much more training. The late-acquired alignments are not arbitrary: they coincide specifically
Load-bearing premise
The three brain-similarity metrics must isolate shared representational content rather than trivial low-level image statistics; if color, contrast, or spatial frequency dominate, the reported factor effects would be an artifact.
Editorial extensions
If this is right
- Scaling alone is not the whole story: image content interacts with model size, so human-centric training data is a necessary ingredient for maximum brain similarity.
- Training longer does not improve all brain areas uniformly; it progressively adds alignment to late and prefrontal cortices.
- Cortical anatomy and physiology (expansion, thickness, myelination, timescale) can serve as predictors of when in training a given area will become model-aligned.
- fMRI- and MEG-based metrics capture distinct facets of brain-model similarity—representational, topographical, and temporal—so evaluating models on only one of them can miss important differences.
- Self-supervised training on natural, human-relevant images is a viable route to more brain-like visual representations, with implications for designing models that better predict or interface with human vision.
Reading between the lines
- If the chronology is a general property of self-supervised visual learning, then the common practice of measuring brain similarity only after final training may hide the fact that late-frontal alignment requires far more training than early sensory alignment; evaluation schedules should sample across training time.
- The cortical-property indexing suggests a testable prediction: models trained on different data distributions should still acquire areas in the same cortical-anatomy order if the ordering is set by intrinsic geometry, or in different orders if data matters more.
- The specific finding that human-centric images improve brain similarity, while objective image statistics are controlled, would strengthen the case that ecological relevance—not just image diversity—drives convergence.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The abstract (arXiv:2508.18226) claims a systematic disentanglement of factors driving brain-like representations in self-supervised vision transformers (DINOv3). It reports that model size, training amount, and image type independently and interactively affect brain similarity measured with fMRI and MEG, via three metrics: overall representational similarity, topographical organization, and temporal dynamics. It further claims that brain-like representations emerge in a specific chronology during training—early sensory areas first, then late and prefrontal areas—and that this developmental trajectory is indexed by cortical expansion, thickness, myelination, and timescales. However, the supplied full text is not the actual paper: it is an unrelated acoustics manuscript on dual corona discharge transducers (arXiv:2508.18232v1). Thus, only the abstract is available for review, and none of the methods, model specifications, statistical controls, or supporting analyses described in the claims can be inspected.
Significance. If the reported findings held, they would be a notable contribution: a factorial manipulation of architecture, training scale, and data composition in modern self-supervised ViTs, evaluated against both fMRI and MEG with spatial and temporal resolution, plus a proposed developmental alignment chronology. The combination of multiple brain-similarity metrics and a parametric model family is a promising design. However, the significance cannot currently be assessed: the manuscript body does not contain the study, no code or data are provided, and the central claims are made at a level of abstraction that prevents verification. The potential is real, but the submission as it stands is not a reviewable scientific paper.
major comments (3)
- [Full text (arXiv:2508.18232v1)] The submitted full text is not the manuscript described by the abstract. It is a physics/applied-photonics paper on controlling loudspeaker directivity with corona discharge transducers. None of the methods needed to evaluate the abstract's claims—DINOv3 training details, factor definitions, fMRI/MEG preprocessing, similarity metric construction, or statistical inference—are present. This is not a local presentation problem; it renders the central claims unverifiable in the submitted manuscript.
- [Abstract, 'three complementary metrics'] The causal claim that model size, training amount, and image type independently and interactively impact brain similarity presupposes that the three metrics isolate representational content rather than low-level image statistics. The abstract gives no evidence that the overall representational similarity, topographical organization, and temporal dynamics metrics are not dominated by color, contrast, spatial frequency, or texture differences between human-centric and other images. Without control analyses (e.g., image-statistic matching, permutation controls, or metric-level ablations), the headline conclusion that larger models and human-centric images are more brain-like could be an artifact of low-level stimulus properties.
- [Abstract, 'follows a specific chronology'] The claim that models align first with early sensory cortices and only later with late/prefrontal areas requires a precise operationalization of alignment, a threshold for when an area is considered aligned, and appropriate multiple-comparison correction across regions and time windows. The abstract reports no such criteria and no evidence that the ordering is stable across training runs or metric choices. This is load-bearing for the developmental-trajectory conclusion and cannot be checked in the absence of the methods.
minor comments (2)
- [Abstract] The abstract does not define the levels of each factor (e.g., what counts as 'largest' model or 'most human-centric' images), nor does it cite the DINOv3 architecture or the metric definitions. A reader cannot assess the scope of the factorial manipulation.
- [Abstract] No statistical details—sample sizes, number of participants, number of images, significance thresholds—are reported. These should appear in an extended summary or full text.
Circularity Check
No circularity found in the abstract; supplied full text is an unrelated paper and cannot support any circularity claim.
full rationale
The abstract claims that model size, training amount, and image type affect brain-similarity as measured by fMRI and MEG. These benchmarks are external to the model training and are not defined in terms of the fitted model parameters, so the comparison is not circular by construction. The training chronology is described as an observed property of training trajectories rather than an input to the fitting procedure. There are no quotations, equations, or self-citations in the abstract to reduce any prediction to an input. The full text supplied is actually a different physics paper (arXiv:2508.18232) on corona discharge transducers, so the methods, metrics, controls, and statistical guardrails of the brain-similarity study cannot be inspected. This creates an unverifiability/correctness-risk concern, but not a demonstrated circularity. No specific step can be exhibited where a result equals its input by construction, so per the hard rules the score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Self-supervised vision transformers trained on natural images can meaningfully model human visual representations
- domain assumption fMRI and MEG provide valid, complementary measurements of brain representations with sufficient spatial and temporal resolution
- domain assumption The three metrics (representational similarity, topographical organization, temporal dynamics) are independent and collectively capture brain-model convergence
Cite this review
Pith. "Pith review of Disentangling the Factors of Convergence between Brains and Computer Vision Models." pith.science (2026). https://pith.science/paper/KM3GPU4M
@misc{pith2026250818226,
author = {Pith},
title = {Pith review of: Disentangling the Factors of Convergence between Brains and Computer Vision Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/KM3GPU4M}},
note = {Machine review of arXiv:2508.18226}
}
read the original abstract
Many AI models trained on natural images develop representations that resemble those of the human brain. However, the factors that drive this brain-model similarity remain poorly understood. To disentangle how the model, training and data independently lead a neural network to develop brain-like representations, we trained a family of self-supervised vision transformers (DINOv3) that systematically varied these different factors. We compare their representations of images to those of the human brain recorded with both fMRI and MEG, providing high resolution in spatial and temporal analyses. We assess the brain-model similarity with three complementary metrics focusing on overall representational similarity, topographical organization, and temporal dynamics. We show that all three factors - model size, training amount, and image type - independently and interactively impact each of these brain similarity metrics. In particular, the largest DINOv3 models trained with the most human-centric images reach the highest brain-similarity. This emergence of brain-like representations in AI models follows a specific chronology during training: models first align with the early representations of the sensory cortices, and only align with the late and prefrontal representations of the brain with considerably more training. Finally, this developmental trajectory is indexed by both structural and functional properties of the human cortex: the representations that are acquired last by the models specifically align with the cortical areas with the largest developmental expansion, thickness, least myelination, and slowest timescales. Overall, these findings disentangle the interplay between architecture and experience in shaping how artificial neural networks come to see the world as humans do, thus offering a promising framework to understand how the human brain comes to represent its visual world.
Forward citations
Cited by 7 Pith papers
-
CanViT: Toward Active-Vision Foundation Models
CanViT is the first task- and policy-agnostic AVFM pretrained via passive-to-active dense latent distillation on 13.2M scenes and 1B random glimpses, achieving 38.5% ADE20K mIoU in one glimpse and 84.5% ImageNet-1k to...
-
IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers
IRIS measures orientation selectivity in vision transformers and shows that a representational similarity score's peak predicts the best layer depth for fine-tuning.
-
Misalignment Between Backpropagation and the Hierarchy of Brain Responses to Images
Backpropagated gradients from vision models predict higher visual cortex signals but diverge from brain hierarchies in spatial and temporal organization.
-
What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?
Text-embedder representations of machine-generated image captions rival vision-model features for predicting high-level visual brain responses and human similarity judgments.
-
Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
Predicting the outputs of several hidden layers of an EMA teacher, rather than only the final layer or pixels, substantially improves self-supervised ViT representations on ImageNet and downstream tasks.
-
Revisiting the Platonic Representation Hypothesis: An Aristotelian View
After permutation-based null-calibration, cross-modal convergence in global spectral similarity disappears, while local neighborhood overlap remains aligned.
-
Toward Aristotelian Medical Representations: Backpropagation-Free Layer-wise Analysis for Interpretable Generalized Metric Learning on MedMNIST
A-ROM delivers competitive MedMNIST performance via pretrained ViT metric spaces, a concept dictionary, and kNN without backpropagation or fine-tuning, framed as interpretable few-shot learning under the Platonic Repr...
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
Allen, Ghislain St-Yves, Yihan Wu, Jesse L
Emily J. Allen, Ghislain St-Yves, Yihan Wu, Jesse L. Breedlove, Jacob S. Prince, Logan T. Dowdle, Matthias Nau, Brad Caron, Franco Pestilli, Ian Charest, J. Benjamin Hutchinson, Thomas Naselaris, and Kendrick Kay. A massive 7t fmri dataset to bridge cognitive neuroscience and artificial intelligence. Nature Neuroscience, 25 0 (1): 0 116–126, January 2022....
-
[3]
Scaling laws for decoding images from brain activity
Hubert Banville, Yohann Benchetrit, St \'e phane d'Ascoli, J \'e r \'e my Rapin, and Jean-R \'e mi King. Scaling laws for decoding images from brain activity. arXiv preprint arXiv:2501.15322, 2025
arXiv 2025
-
[4]
Tomasini, Alessandro Favero, and Matthieu Wyart
Francesco Cagnetta, Leonardo Petrini, Umberto M. Tomasini, Alessandro Favero, and Matthieu Wyart. How deep neural networks learn compositional data: The random hierarchy model. Physical Review X, 14 0 (3), July 2024. ISSN 2160-3308. doi:10.1103/physrevx.14.031001. http://dx.doi.org/10.1103/PhysRevX.14.031001
-
[5]
Brains and algorithms partially converge in natural language processing
Charlotte Caucheteux and Jean-R \'e mi King. Brains and algorithms partially converge in natural language processing. Communications biology, 5 0 (1): 0 134, 2022
work page 2022
-
[6]
Schwing, Alexander Kirillov, and Rohit Girdhar
Bowen Cheng, Ishan Misra, Alexander G. Schwing, Alexander Kirillov, and Rohit Girdhar. Masked-attention mask transformer for universal image segmentation. 2022
2022
-
[7]
Radoslaw Martin Cichy, Aditya Khosla, Dimitrios Pantazis, Antonio Torralba, and Aude Oliva. Comparison of deep neural networks to spatio-temporal cortical dynamics of human visual object recognition reveals hierarchical correspondence. Scientific Reports, 6 0 (1), June 2016. ISSN 2045-2322. doi:10.1038/srep27755. http://dx.doi.org/10.1038/srep27755
-
[8]
Colin Conwell, Jacob S. Prince, George A. Alvarez, and Talia Konkle. What can 5.17 billion regression fits tell us about artificial models of the human visual system? In SVRHM 2021 Workshop @ NeurIPS, 2021. https://openreview.net/forum?id=i_xiyGq6FNT
work page 2021
Show all 62 references
-
[9]
What can 1.8 billion regressions tell us about the pressures shaping high-level visual representation in brains and machines? BioRxiv, pages 2022--03, 2022
Colin Conwell, Jacob S Prince, Kendrick N Kay, George A Alvarez, and Talia Konkle. What can 1.8 billion regressions tell us about the pressures shaping high-level visual representation in brains and machines? BioRxiv, pages 2022--03, 2022
2022
-
[10]
How we learn: Why brains learn better than any machine
Stanislas Dehaene. How we learn: Why brains learn better than any machine... for now. Penguin, 2021
2021
-
[11]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE Conference on Computer Vision and Pattern Recognition. IEEE, June 2009. doi:10.1109/cvpr.2009.5206848. http://dx.doi.org/10.1109/CVPR.2009.5206848
2009
-
[12]
How does the brain solve visual object recognition? Neuron, 73 0 (3): 0 415--434, 2012
James J DiCarlo, Davide Zoccolan, and Nicole C Rust. How does the brain solve visual object recognition? Neuron, 73 0 (3): 0 415--434, 2012
2012
-
[13]
Kietzmann, Emily Allen, Yihan Wu, Thomas Naselaris, Kendrick Kay, and Ian Charest
Adrien Doerig, Tim C. Kietzmann, Emily Allen, Yihan Wu, Thomas Naselaris, Kendrick Kay, and Ian Charest. High-level visual representations in the human brain are aligned with large language models. Nature Machine Intelligence, August 2025. ISSN 2522-5839. doi:10.1038/s42256-02...
2025 doi
-
[14]
Seeing it all: Convolutional network layers map the function of the human visual system
Michael Eickenberg, Alexandre Gramfort, Ga \"e l Varoquaux, and Bertrand Thirion. Seeing it all: Convolutional network layers map the function of the human visual system. NeuroImage, 152: 0 184--194, 2017
2017
-
[15]
Dermatologist-level classification of skin cancer with deep neural networks
Andre Esteva, Brett Kuprel, Roberto A Novoa, Justin Ko, Susan M Swetter, Helen M Blau, and Sebastian Thrun. Dermatologist-level classification of skin cancer with deep neural networks. nature, 542 0 (7639): 0 115--118, 2017
2017
-
[16]
Emergence of language in the developing brain
Linnea Evanson, Christine Bulteau, Mathilde Chipaux, Georg Dorfmüller, Sarah Ferrand-Sorbets, Emmanuel Raffo, Sarah Rosenberg, Pierre Bourdillon, and Jean-Rémi King. Emergence of language in the developing brain. Manuscript, May 2025. mailto:jeanremi@meta.com. Equal contributi...
2025
-
[17]
Distributed hierarchical processing in the primate cerebral cortex
Daniel J Felleman and David C Van Essen. Distributed hierarchical processing in the primate cerebral cortex. Cerebral cortex (New York, NY: 1991), 1 0 (1): 0 1--47, 1991
1991
-
[18]
Freesurfer
Bruce Fischl. Freesurfer. NeuroImage, 62 0 (2): 0 774–781, August 2012. ISSN 1053-8119. doi:10.1016/j.neuroimage.2012.01.021. http://dx.doi.org/10.1016/j.neuroimage.2012.01.021
2012 doi
-
[19]
Gifford, Maya A
Alessandro T. Gifford, Maya A. Jastrzębowska, Johannes J. D. Singer, and Radoslaw M. Cichy. In silico discovery of representational relationships across visual cortex. Nature Human Behaviour, June 2025. ISSN 2397-3374. doi:10.1038/s41562-025-02252-z. http://dx.doi.org/10.1038/...
2025 doi
-
[20]
Meg and eeg data analysis with mne-python
Alexandre Gramfort, Martin Luessi, Eric Larson, Denis A Engemann, Daniel Strohmeier, Christian Brodbeck, Roman Goj, Mainak Jas, Teon Brooks, Lauri Parkkonen, et al. Meg and eeg data analysis with mne-python. Frontiers in Neuroinformatics, 7: 0 267, 2013
2013
-
[21]
Monika Graumann, Caterina Ciuffi, Kshitij Dwivedi, Gemma Roig, and Radoslaw M. Cichy. The spatiotemporal neural dynamics of object location representations in the human brain. Nature Human Behaviour, 6 0 (6): 0 796–811, February 2022. ISSN 2397-3374. doi:10.1038/s41562-022-013...
2022 doi
-
[22]
Nastase, and Ariel Goldstein
Uri Hasson, Samuel A. Nastase, and Ariel Goldstein. Direct fit to nature: An evolutionary perspective on biological and artificial neural networks. Neuron, 105 0 (3): 0 416–434, February 2020. ISSN 0896-6273. doi:10.1016/j.neuron.2019.12.002. http://dx.doi.org/10.1016/j.neuron...
2020 doi
-
[23]
things-meg
Martin N. Hebart, Oliver Contier, Lina Teichmann, Adam H. Rockter, Charles Zheng, Alexis Kidder, Anna Corriveau, Maryam Vaziri-Pashkam, and Chris I. Baker. "things-meg", 2023 a
2023
-
[24]
THINGS -data, a multimodal collection of large-scale datasets for investigating object representations in human brain and behavior
Martin N Hebart, Oliver Contier, Lina Teichmann, Adam H Rockter, Charles Y Zheng, Alexis Kidder, Anna Corriveau, Maryam Vaziri-Pashkam, and Chris I Baker. THINGS -data, a multimodal collection of large-scale datasets for investigating object representations in human brain and ...
2023 doi
-
[25]
Similar patterns of cortical expansion during human development and evolution
Jason Hill, Terrie Inder, Jeffrey Neil, Donna Dierker, John Harwell, and David Van Essen. Similar patterns of cortical expansion during human development and evolution. Proceedings of the National Academy of Sciences, 107 0 (29): 0 13135--13140, 2010
2010
-
[26]
The platonic representation hypothesis
Minyoung Huh, Brian Cheung, Tongzhou Wang, and Phillip Isola. The platonic representation hypothesis. arXiv preprint arXiv:2405.07987, 2024
2024 arXiv
-
[27]
Characterizing the dynamics of mental representations: the temporal generalization method
Jean-R \'e mi King and Stanislas Dehaene. Characterizing the dynamics of mental representations: the temporal generalization method. Trends in cognitive sciences, 18 0 (4): 0 203--210, 2014
2014
-
[28]
Deep neural networks: A new framework for modeling biological vision and brain information processing
Nikolaus Kriegeskorte. Deep neural networks: A new framework for modeling biological vision and brain information processing. Annual Review of Vision Science, 1 0 (1): 0 417–446, November 2015. ISSN 2374-4650. doi:10.1146/annurev-vision-082114-035447. http://dx.doi.org/10.1146...
2015 doi
-
[29]
Representational similarity analysis-connecting the branches of systems neuroscience
Nikolaus Kriegeskorte, Marieke Mur, and Peter A Bandettini. Representational similarity analysis-connecting the branches of systems neuroscience. Frontiers in systems neuroscience, 2: 0 249, 2008
2008
-
[30]
Feature-space selection with banded ridge regression
Tom Dupr \'e La Tour, Michael Eickenberg, Anwar O Nunez-Elizalde, and Jack L Gallant. Feature-space selection with banded ridge regression. NeuroImage, 264: 0 119728, 2022
2022
-
[31]
De Lorenci, Seung Eun Yi, Th \'e o Moutakanni, Piotr Bojanowski, Camille Couprie, Juan C
Alice V. De Lorenci, Seung Eun Yi, Th \'e o Moutakanni, Piotr Bojanowski, Camille Couprie, Juan C. Caicedo, and Wolfgang Maximilian Anton Pernice. Scaling channel-adaptive self-supervised learning. Transactions on Machine Learning Research, 2025. ISSN 2835-8856. https://openre...
2025
-
[32]
Mahner, Lukas Muttenthaler, Umut G\" u c l\" u , and Martin N
Florian P. Mahner, Lukas Muttenthaler, Umut G\" u c l\" u , and Martin N. Hebart. Dimensions underlying the representational alignment of deep neural networks with humans. Nature Machine Intelligence, 7 0 (6): 0 848–859, June 2025. ISSN 2522-5839. doi:10.1038/s42256-025-01041-...
2025 doi
-
[33]
Neuromaps: structural and functional interpretation of brain maps
Ross D Markello, Justine Y Hansen, Zhen-Qi Liu, Vincent Bazinet, Golia Shafiei, Laura E Su \'a rez, Nadia Blostein, Jakob Seidlitz, Sylvain Baillet, Theodore D Satterthwaite, et al. Neuromaps: structural and functional interpretation of brain maps. Nature Methods, 19 0 (11): 0...
2022
-
[34]
Spoerer, Nikolaus Kriegeskorte, and Tim C
Johannes Mehrer, Courtney J. Spoerer, Nikolaus Kriegeskorte, and Tim C. Kietzmann. Individual differences among deep neural network models. Nature Communications, 11 0 (1), November 2020. ISSN 2041-1723. doi:10.1038/s41467-020-19632-w. http://dx.doi.org/10.1038/s41467-020-19632-w
2020 doi
-
[35]
Toward a realistic model of speech processing in the brain with self-supervised learning, 2023
Juliette Millet, Charlotte Caucheteux, Pierre Orhan, Yves Boubenec, Alexandre Gramfort, Ewan Dunbar, Christophe Pallier, and Jean-Remi King. Toward a realistic model of speech processing in the brain with self-supervised learning, 2023. https://arxiv.org/abs/2206.01685
2023 arXiv
-
[36]
Encoding and decoding in fMRI
Thomas Naselaris, Kendrick N Kay, Shinji Nishimoto, and Jack L Gallant. Encoding and decoding in fMRI . Neuroimage, 56 0 (2): 0 400--410, 2011
2011
-
[37]
Modality-agnostic fmri decoding of vision and language, 2024
Mitja Nikolaus, Milad Mozafari, Nicholas Asher, Leila Reddy, and Rufin VanRullen. Modality-agnostic fmri decoding of vision and language, 2024. https://arxiv.org/abs/2403.11771
2024 arXiv
-
[38]
Pedregosa, G
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. Scikit-learn: Machine learning in P ython. Journal of Machine Learnin...
2011
-
[39]
Brain-like emergent properties in deep networks: impact of network architecture, datasets and training, 2024
Niranjan Rajesh, Georgin Jacob, and SP Arun. Brain-like emergent properties in deep networks: impact of network architecture, datasets and training, 2024. https://arxiv.org/abs/2411.16326
2024
-
[40]
You only look once: Unified, real-time object detection, 2016
Joseph Redmon, Santosh Divvala, Ross Girshick, and Ali Farhadi. You only look once: Unified, real-time object detection, 2016. https://arxiv.org/abs/1506.02640
2016 arXiv
-
[41]
Majaj, Rishi Rajalingham, Elias B
Martin Schrimpf, Jonas Kubilius, Ha Hong, Najib J. Majaj, Rishi Rajalingham, Elias B. Issa, Kohitij Kar, Pouya Bashivan, Jonathan Prescott-Roy, Franziska Geiger, Kailyn Schmidt, Daniel L. K. Yamins, and James J. DiCarlo. Brain-score: Which artificial neural network for object ...
2018 doi
-
[42]
Seeliger, M
K. Seeliger, M. Fritsche, U. G\" u c l\" u , S. Schoenmakers, J.-M. Schoffelen, S.E. Bosch, and M.A.J. van Gerven. Convolutional neural network-based encoding and decoding of visual object recognition in space and time. NeuroImage, 180: 0 253–266, October 2018. ISSN 1053-8119....
2018 doi
-
[43]
Human electromagnetic and haemodynamic networks systematically converge in unimodal cortex and diverge in transmodal cortex
Golia Shafiei, Sylvain Baillet, and Bratislav Misic. Human electromagnetic and haemodynamic networks systematically converge in unimodal cortex and diverge in transmodal cortex. September 2021. doi:10.1101/2021.09.07.458941. http://dx.doi.org/10.1101/2021.09.07.458941
2021 doi
-
[44]
Alignment between brains and ai: Evidence for convergent evolution across modalities, scales and training trajectories, 2025
Guobin Shen, Dongcheng Zhao, Yiting Dong, Qian Zhang, and Yi Zeng. Alignment between brains and ai: Evidence for convergent evolution across modalities, scales and training trajectories, 2025. https://arxiv.org/abs/2507.01966
2025 arXiv
-
[45]
Oriane Sim \'e oni, Huy V. Vo, Maximilian Seitzer, Federico Baldassarre, Maxime Oquab, Cijo Jose, Vasil Khalidov, Marc Szafraniec, Seungeun Yi, Michaël Ramamonjisoa, Massa Francisco, Daniel Haziza, Luca Wehrstedt, Jianyuan Wang, Timothée Darcet, Théo Moutakanni, Leonel Sentana...
2025
-
[46]
Representations in vision and language converge in a shared, multidimensional space of perceived similarities, 2025
Katerina Marie Simkova, Adrien Doerig, Clayton Hickey, and Ian Charest. Representations in vision and language converge in a shared, multidimensional space of perceived similarities, 2025. https://arxiv.org/abs/2507.21871
2025 arXiv
-
[47]
Solomon, K
S.H. Solomon, K. Kay, and A.C. Schapiro. Semantic plasticity across timescales in the human brain. bioRxiv, 2024. doi:10.1101/2024.02.07.579310. https://www.biorxiv.org/content/early/2024/05/24/2024.02.07.579310. Publisher: Cold Spring Harbor Laboratory
2024 doi
-
[48]
Vo, Vasudev Lal, and Alexander G
Jerry Tang, Meng Du, Vy A. Vo, Vasudev Lal, and Alexander G. Huth. Brain encoding models based on multimodal transformers can transfer across language and vision, 2023. https://arxiv.org/abs/2305.12248
2023 arXiv
-
[49]
Many-two-one: Diverse representations across visual pathways emerge from a single objective
Yingtian Tang, Abdulkadir Gokce, Khaled Jedoui Al-Karkari, Daniel Yamins, and Martin Schrimpf. Many-two-one: Diverse representations across visual pathways emerge from a single objective. July 2025. doi:10.1101/2025.07.22.664908
2025 doi
-
[50]
Prince, Rosa Cao, and Daniel LK Yamins
Imran Thobani, Javier Sagastuy-Brena, Aran Nayebi, Jacob S. Prince, Rosa Cao, and Daniel LK Yamins. Model-brain comparison using inter-animal transforms. In 8th Annual Conference on Cognitive Computational Neuroscience, 2025. https://openreview.net/forum?id=bra729zCMm
2025
-
[51]
Siglip 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense features, 2025
Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muhammad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, Olivier Hénaff, Jeremiah Harmsen, Andreas Steiner, and Xiaohua Zhai. Siglip 2: Multilingual vision-language enc...
2025 arXiv
-
[52]
The wu-minn human connectome project: an overview
David C Van Essen, Stephen M Smith, Deanna M Barch, Timothy EJ Behrens, Essa Yacoub, Kamil Ugurbil, Wu-Minn HCP Consortium, et al. The wu-minn human connectome project: an overview. Neuroimage, 80: 0 62--79, 2013
2013
-
[53]
Loek van Rossem and Andrew M. Saxe. When representations align: Universality in representation learning dynamics, 2024. https://arxiv.org/abs/2402.09142
2024 arXiv
-
[54]
Scipy 1.0: fundamental algorithms for scientific computing in python
Pauli Virtanen, Ralf Gommers, Travis E Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, et al. Scipy 1.0: fundamental algorithms for scientific computing in python. Nature methods, 17 0 (3): 0 261--272, 2020
2020
-
[55]
Butterfly effects in perceptual development: A review of the ‘adaptive initial degradation’ hypothesis
Lukas Vogelsang, Marin Vogelsang, Gordon Pipa, Sidney Diamond, and Pawan Sinha. Butterfly effects in perceptual development: A review of the ‘adaptive initial degradation’ hypothesis. Developmental Review, 71: 0 101117, March 2024. ISSN 0273-2297. doi:10.1016/j.dr.2024.101117
2024
-
[56]
Using goal-driven deep learning models to understand sensory cortex
Daniel L K Yamins and James J DiCarlo. Using goal-driven deep learning models to understand sensory cortex. Nature Neuroscience, 19 0 (3): 0 356–365, February 2016. ISSN 1546-1726. doi:10.1038/nn.4244. http://dx.doi.org/10.1038/nn.4244
2016 doi
-
[57]
Daniel L. K. Yamins, Ha Hong, Charles F. Cadieu, Ethan A. Solomon, Darren Seibert, and James J. DiCarlo. Performance-optimized hierarchical models predict neural responses in higher visual cortex. Proceedings of the National Academy of Sciences, 111 0 (23): 0 8619–8624, May 20...
2014 doi
-
[58]
Performance-optimized hierarchical models predict neural responses in higher visual cortex
Daniel LK Yamins, Ha Hong, Charles F Cadieu, Ethan A Solomon, Darren Seibert, and James J DiCarlo. Performance-optimized hierarchical models predict neural responses in higher visual cortex. Proceedings of the national academy of sciences, 111 0 (23): 0 8619--8624, 2014 b
2014
-
[59]
Frank, James J
Chengxu Zhuang, Siming Yan, Aran Nayebi, Martin Schrimpf, Michael C. Frank, James J. DiCarlo, and Daniel L. K. Yamins. Unsupervised neural network models of the ventral visual stream. Proceedings of the National Academy of Sciences, 118 0 (3), January 2021. ISSN 1091-6490. doi...
2021 doi
-
[60]
@esa (Ref
\@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...
-
[61]
\@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...
-
[62]
@open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.