REVIEW 2 major objections 4 minor 97 references
The paper claims that seen images can be decoded from fMRI within about ten seconds of stimulus onset, using only one hour of a new participant's scan data.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 07:32 UTC pith:BI4TK4VC
load-bearing objection A real engineering proof-of-concept for putting MindEye2 inside a real-time fMRI loop; the live-session evidence is latency-only, so the paper's central demonstration is still simulation. the 2 major comments →
Real-time Reconstruction of Human Visual Perception from fMRI
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is that single-trial visual decoding survives the switch from offline to real-time processing. In the fastest condition — waiting about 7.9 seconds for the BOLD response to peak, then running motion correction, a per-trial GLM, and decoder inference — the model retrieves the seen image from a 50-image pool 36–40% of the time (chance 2%) and produces reconstructions that score above chance on multiple metrics. This works with a condensed version of a 700M+ parameter architecture, on 3T rather than 7T data, and after only one training session of roughly one hour from the new participant.
What carries the argument
A large pretrained fMRI-to-image backbone that maps single-trial beta estimates into a vision-language embedding space. In real time, each new brain volume is aligned to the training session, a general linear model converts the overlapping BOLD response into one beta vector per trial, and that vector is projected into the shared embedding space; a frozen diffusion model turns the embedding into an image, or nearest-neighbor search retrieves the closest candidate. The key is keeping inference under about five seconds so the only real latency is the unavoidable hemodynamic delay.
Load-bearing premise
The load-bearing assumption is that the temporary overlap between pretraining and test images did not meaningfully inflate the reported real-time accuracy; if that leakage is large, the claim that held-out perception is being decoded weakens.
What would settle it
Fine-tune or pretrain the model with a strict continuous-block train/test split, so no test image ever shares a block with a training image, then measure fast real-time retrieval accuracy on the same 50-image test set. If top-1 accuracy falls to the 2% chance level, the real-time decoding claim is falsified; if it stays above roughly 30%, the central result holds.
If this is right
- Real-time neurofeedback can now target fine-grained visual content, not just coarse category or arousal levels.
- A new participant can get a working decoder after a single one-hour training session on standard hospital 3T scanners.
- Researchers can trade delay against accuracy: waiting roughly 30 seconds gives most of the offline benefit, suggesting an operating point for closed-loop experiments.
- Retrieval at a 2-item pool reached about 90% in the fastest condition, so relative comparisons between two mental images are already usable for latent-space feedback.
- The same real-time streaming architecture can potentially host other computationally heavy decoding models beyond image reconstruction.
Where Pith is reading between the lines
- If the one-hour fine-tuning result generalizes beyond the single participant tested, the main cost of adopting fMRI-based brain-computer interfaces shifts from data collection to access to a scanner.
- The leakage caveat in the appendix suggests a clean test: re-train with a strictly block-wise split; until that is done, the exact size of the real-time decoding advantage is uncertain.
- The same pipeline could be pointed at imagined rather than seen images, since the latent-space mapping may transfer; the paper does not test this.
- Real-time decoding could expose private cognitive content, so ethical safeguards may need to be built into the interface itself, not just added as consent language.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a real-time-compatible adaptation of the MindEye2 fMRI-to-image decoding pipeline, integrated with the RT-Cloud platform. Using a 3T scanner and approximately one hour of fine-tuning data from a new participant, the authors report above-chance single-trial image retrieval and reconstruction in 'fast' (14.5 s), 'slow' (36 s), and 'end-of-run' (2.7 min) real-time-compatible settings, with retrieval accuracy up to 36–40% (chance 2%) in the fast condition. They replicate the qualitative pattern on held-out NSD subj01 and include a no-pretraining control. The actual live RT-Cloud session is described as a proof-of-concept and is used to report processing latencies; all decoding accuracy metrics come from simulated real-time replay of previously acquired data. The paper also documents preprocessing, model architecture, training, and evaluation details, and makes code and data publicly available.
Significance. If the central claim were fully supported, this would be a notable practical advance: it would show that a state-of-the-art generative fMRI decoder can operate inside the real-time fMRI envelope with only ~1 hour of new-participant data, opening doors to closed-loop neurofeedback and BCI applications. The paper's strengths include a clear comparison of pipeline variants, a no-pretraining control, replication on an independent NSD subject, detailed latency measurements, and open release of code and data. These are real contributions. However, the headline claim that the authors 'demonstrate for the first time' real-time single-trial decoding is not directly supported by the evidence, because the live session was not quantitatively scored. The simulated real-time analyses are valuable and well designed, but they do not by themselves prove that the end-to-end live system produces above-chance outputs.
major comments (2)
- [Abstract; §3.3; §6] The headline claim, stated in the abstract as 'we demonstrate for the first time that it is possible to decode seen images from fMRI at single-trial resolution in real-time' and repeated in the contributions list, is not supported by the evidence reported for the live session. Every decoding accuracy result in Tables 1, 3, and 4 and Figures 4–8 comes from 'simulated real-time analyses' (explicitly stated in §3.3 and §6). Table 2 reports only latencies from the live RT-Cloud session; no retrieval or reconstruction accuracy is reported for live trials. A simulation can establish that an algorithm is compatible with a real-time latency budget, but it does not verify the end-to-end live system — DICOM streaming, online registration and motion correction, time-pressured GLM fitting, and inference within a fixed wall-clock window — actually produces above-chance outputs. The abstract's phrasin
- [§2.6.1; Appendix A.1] The pretraining/test interleaving leakage admitted in Appendix A.1 is load-bearing for the claimed generalization to held-out perception and should be quantified. Because the test images (the 50 special515 images) were temporally interleaved with pretraining images in the NSD acquisition, BOLD responses to test trials can contaminate the beta estimates of adjacent pretraining trials. The authors argue this is minor because the model saw fMRI data but not CLIP labels for the test images; however, leakage through feature representations does not require access to labels. If contamination materially inflated the reported single-trial retrieval/reconstruction numbers, the central claim that held-out perception is being decoded would be weakened. The fact that the test set is fixed across all evaluations makes this concern concrete. Please add a quantitative control, for example re-running pr
minor comments (4)
- [§3.3; Table 2] The advertised '9.5 seconds post image onset for retrieval' is not directly derivable from Table 2. Summing the fast-condition latencies for retrieval (stimulus delay 7.85 s + motion correction 0.39 s + registration 0.18 s + GLM fit 1.09 s + inference 0.19 s + retrieval 0.50 s) gives about 10.2 s. Section 3.3 itself says 'fast' retrieval takes ~10s. Please correct the abstract and contributions to use a consistent, well-defined latency.
- [Table 2] The 'Total Latency' row appears to sum reconstruction and retrieval times as if they were serial components. If reconstruction and retrieval are alternative inference branches that can be run in parallel or selectively, the total latency should be defined and labeled accordingly to avoid ambiguity.
- [Figure 2] Figure 2 is labeled 'Hand-picked example reconstructions.' Since the paper also provides randomly selected reconstructions in Figure 10, the text could briefly note that Figure 2 is intended to show favorable examples, to prevent over-interpretation.
- [§2.7] For the two-way reconstruction metrics, the description says chance is 50%, but the exact averaging procedure over pairwise comparisons could be clarified a bit more, especially regarding whether all mismatched pairs are used or a sampled subset.
Circularity Check
No derivation-level circularity; one admitted pretraining/test leakage compromises full independence of the held-out evaluation.
specific steps
-
other
[Appendix A.1 (Limitations); §2.6.1 Train and Test Split]
"One limitation of our 7T pretraining procedure is that the images used for pretraining were interleaved with some of the images that were later (in a separate session) used for testing. Due to the temporal lag of the BOLD response, this could have led to a minor form of data leakage, whereby the neural response to the test images affects the beta maps for images presented after them during pretraining"
This is not classic derivation circularity (no equation reduces a prediction to a fit), but it is an evaluation-independence leak: the 'held-out' test images' BOLD activity may have entered the construction of pretraining beta maps via temporally overlapping responses. The headline claim of decoding 'seen images from fMRI at single-trial resolution in real-time' is supported by retrieval/reconstruction scores on these same test images, so part of the support is self-referential: the model's training inputs contained neural information from the test trials. The authors argue the inflation is minor because CLIP labels for the test images were withheld, but the mechanism is real, admitted, and load-bearing for the 'held-out' framing.
full rationale
The paper is an empirical engineering/decoding study rather than a derivation, so the classic circularity patterns (self-definitional identities, fitted parameters renamed as predictions, ansatz smuggled via citation) do not apply. The training/test logic is: pretrain MindEye2 on 7 NSD subjects, fine-tune on ~1 hour of new 3T data, then evaluate on a separate session with a fixed 50-image test set. This chain is externally checkable: MindEye2 and RT-Cloud are public code/models, and the authors replicate the main pattern on the held-out NSD subj01 and include a no-pretraining control. Self-citations to MindEye2, RT-Cloud, and GLMsingle are therefore real evidence rather than circularity. The one flagged item is the authors' own admission in Appendix A.1 that 7T pretraining images were interleaved with later test images, so test-image BOLD responses may have bled into pretraining beta maps. This is an independence leak, not an equation-level reduction, but because the central claim rests on above-chance scores from this test set, it is load-bearing enough to raise the score to 3. Separately, the live RT session contributes latency while accuracy numbers come from simulated replay; that is a support gap, not circularity. Overall: no self-definitional or fitted-input-as-prediction circularity.
Axiom & Free-Parameter Ledger
free parameters (4)
- Reliability threshold r =
0.2
- Fast stimulus delay =
~7.9 s
- Slow stimulus delay =
~29 s
- Shared-subject latent dimensionality =
1024
axioms (5)
- domain assumption BOLD response can be modeled as a linear time-invariant system convolved with an HRF, and single-trial betas from a GLM capture stimulus-specific information.
- domain assumption CLIP embedding space is a meaningful target space: visual-semantic similarity in CLIP corresponds to perceptual similarity, and fMRI-to-CLIP mapping generalizes across participants.
- domain assumption A model pretrained on 7 NSD subjects can be adapted to a new participant with about one hour of fine-tuning data.
- domain assumption The reliability mask (r>0.2) plus the nsdgeneral ROI defines the set of informative voxels.
- ad hoc to paper The pretraining/test interleaving leak is minor and does not materially inflate results.
read the original abstract
Real-time closed-loop neurofeedback based on functional magnetic resonance imaging (fMRI) has led to important scientific and clinical advances. However, the sophistication of the analysis methods used in real-time fMRI lags behind the state-of-the-art in fMRI decoding, largely due to computational factors: Most advanced decoding pipelines do not fit within the envelope of real-time processing, where the analysis needs to be conducted in a matter of seconds and without leveraging data acquired later in the session. Here, we present a real-time compatible adaptation of a computationally intensive state-of-the-art pipeline for reconstructing perceived natural images (MindEye2), and we demonstrate that reliable fine-grained decoding is still achievable in this setting. Using RT-Cloud, an open-source, scalable cloud-based platform, we performed a real-time scan where we decoded single-trial visual perception within seconds after an image was shown to the participant. Finally, we use simulated analyses to document the factors driving changes in performance from offline to real-time analysis. This work serves as a proof-of-concept that it is feasible to deploy these powerful fMRI decoding pipelines in real-time analysis, paving the way for their use in brain-computer interfaces for scientific discovery and clinical treatment.
Figures
Reference graph
Works this paper leans on
-
[1]
Learning Transferable Visual Models From Natural Language Supervision.arXiv, February 2021
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agar- wal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision.arXiv, February 2021. doi: 10.48550/arXiv.2103.00020
-
[2]
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language Models are Unsupervised Multitask Learners. https://cdn.openai.com/better-language- models/language_models_are_unsupervised_multitask_learners.pdf, 2019
2019
-
[3]
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.arXiv, May 2019
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.arXiv, May 2019. doi: 10.48550/arXiv.1810.04805
-
[4]
LLaMA: Open and Efficient Foundation Language Models.arXiv, February 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timo- thée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. LLaMA: Open and Efficient Foundation Language Models.arXiv, February 2023. doi: 10.48550/arXiv.2302.13971
-
[5]
High-resolution image reconstruction with latent diffusion models from human brain activity
Yu Takagi and Shinji Nishimoto. High-resolution image reconstruction with latent diffusion models from human brain activity. In2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14453–14463, June 2023. doi: 10.1109/CVPR52729.2023. 01389
arXiv 2023
-
[6]
Matteo Ferrante, Tommaso Boccato, Furkan Ozcelik, Rufin VanRullen, and Nicola Toschi. Through their eyes: Multi-subject brain decoding with simple alignment techniques.Imaging Neuroscience, 2:imag–2–00170, May 2024. ISSN 2837-6056. doi: 10.1162/imag_a_00170
-
[7]
Furkan Ozcelik, Bhavin Choksi, Milad Mozafari, Leila Reddy, and Rufin VanRullen. Re- construction of Perceived Images from fMRI Patterns and Semantic Brain Exploration using Instance-Conditioned GANs.arXiv, February 2022. doi: 10.48550/arXiv.2202.12692
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2202.12692 2022
-
[8]
Furkan Ozcelik and Rufin VanRullen. Natural scene reconstruction from fMRI signals using generative latent diffusion.Scientific Reports, 13(1):15666, September 2023. ISSN 2045-2322. doi: 10.1038/s41598-023-42891-8
-
[9]
Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion Priors
Paul S. Scotti, Atmadeep Banerjee, Jimmie Goode, Stepan Shabalin, Alex Nguyen, Ethan Cohen, Aidan J. Dempster, Nathalie Verlinde, Elad Yundler, David Weisberg, Kenneth A. Norman, and Tanishq Mathew Abraham. Reconstructing the Mind’s Eye: fMRI-to-Image with Contrastive Learning and Diffusion Priors.arXiv, October 2023. doi: 10.48550/arXiv.2305.18274
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2305.18274 2023
-
[10]
Paul S. Scotti, Mihir Tripathy, Cesar Kadir Torrico Villanueva, Reese Kneeland, Tong Chen, Ashutosh Narang, Charan Santhirasegaran, Jonathan Xu, Thomas Naselaris, Kenneth A. Nor- man, and Tanishq Mathew Abraham. MindEye2: Shared-Subject Models Enable fMRI-To- Image With 1 Hour of Data.arXiv, June 2024. doi: 10.48550/arXiv.2403.11207
-
[11]
MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction
Zixuan Gong, Qi Zhang, Guangyin Bao, Lei Zhu, Ke Liu, Liang Hu, and Duoqian Miao. MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction. arXiv, December 2024. doi: 10.48550/arXiv.2404.12630
-
[12]
NSD-Imagery: A benchmark dataset for extending fMRI vision decoding methods to mental imagery
Reese Kneeland, Paul S. Scotti, Ghislain St-Yves, Jesse Breedlove, Kendrick Kay, and Thomas Naselaris. NSD-Imagery: A benchmark dataset for extending fMRI vision decoding methods to mental imagery.arXiv, June 2025. doi: 10.48550/arXiv.2506.06898
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2506.06898 2025
-
[13]
Jerry Tang, Amanda LeBel, Shailee Jain, and Alexander G. Huth. Semantic reconstruction of continuous language from non-invasive brain recordings.Nature Neuroscience, 26(5):858–866, May 2023. ISSN 1546-1726. doi: 10.1038/s41593-023-01304-9
-
[14]
Ranganatha Sitaram, Tomas Ros, Luke Stoeckel, Sven Haller, Frank Scharnowski, Jarrod Lewis- Peacock, Nikolaus Weiskopf, Maria Laura Blefari, Mohit Rana, Ethan Oblak, Niels Birbaumer, and James Sulzer. Closed-loop brain training: The science of neurofeedback.Nature Reviews Neuroscience, 18(2):86–100, February 2017. ISSN 1471-0048. doi: 10.1038/nrn.2016.164
-
[15]
Asma Motiwala, Joana Soldado-Magraner, Aaron P. Batista, Matthew A. Smith, and Byron M. Yu. Brain–computer interfaces as a causal probe for scientific inquiry.Trends in Cognitive Sciences, 30(1):40–53, January 2026. ISSN 13646613. doi: 10.1016/j.tics.2025.06.017. 12
-
[16]
Kymberly D. Young, Greg J. Siegle, Vadim Zotev, Raquel Phillips, Masaya Misaki, Han Yuan, Wayne C. Drevets, and Jerzy Bodurka. Randomized Clinical Trial of Real-Time fMRI Amygdala Neurofeedback for Major Depressive Disorder: Effects on Symptoms and Autobiographical Memory Recall.The American Journal of Psychiatry, 174(8):748–755, August 2017. ISSN 1535-72...
arXiv 2017
-
[17]
Anne C. Mennen, Nicholas B. Turk-Browne, Grant Wallace, Darsol Seok, Adna Jaganjac, Janet Stock, Megan T. deBettencourt, Jonathan D. Cohen, Kenneth A. Norman, and Yvette I. Sheline. Cloud-Based Functional Magnetic Resonance Imaging Neurofeedback to Reduce the Negative Attentional Bias in Depression: A Proof-of-Concept Study.Biological Psychiatry: Cognitiv...
-
[18]
Wammes, Alex Nguyen, Coraline Rinn Iordan, Kenneth A
Kailong Peng, Jeffrey D. Wammes, Alex Nguyen, Coraline Rinn Iordan, Kenneth A. Norman, and Nicholas B. Turk-Browne. Inducing representational change in the hippocampus through real-time neurofeedback.Philosophical Transactions of the Royal Society B: Biological Sciences, 379(1915):20230091, October 2024. ISSN 0962-8436. doi: 10.1098/rstb.2023.0091
arXiv 1915
-
[19]
Grant Wallace, Stephen Polcyn, Paula P. Brooks, Anne C. Mennen, Ke Zhao, Paul S. Scotti, Sebastian Michelmann, Kai Li, Nicholas B. Turk-Browne, Jonathan D. Cohen, and Kenneth A. Norman. RT-Cloud: A cloud-based software framework to simplify and standardize real-time fMRI.NeuroImage, 257:119295, August 2022. ISSN 1053-8119. doi: 10.1016/j.neuroimage. 2022.119295
arXiv 2022
-
[20]
Allen, Ghislain St-Yves, Yihan Wu, Jesse L
Emily J. Allen, Ghislain St-Yves, Yihan Wu, Jesse L. Breedlove, Jacob S. Prince, Logan T. Dow- dle, Matthias Nau, Brad Caron, Franco Pestilli, Ian Charest, J. Benjamin Hutchinson, Thomas Naselaris, and Kendrick Kay. A massive 7T fMRI dataset to bridge cognitive neuroscience and artificial intelligence.Nature Neuroscience, 25(1):116–126, January 2022. ISSN...
-
[21]
Lawrence Zitnick, and Piotr Dollár
Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár. Microsoft COCO: Common Objects in Context.arXiv, February 2015. doi: 10.48550/arXiv.1405.0312
-
[22]
Guo Wanjia, Serra E. Favila, Ghootae Kim, Robert J. Molitor, and Brice A. Kuhl. Abrupt hippocampal remapping signals resolution of memory interference.Nature Communications, 12(1):4816, August 2021. ISSN 2041-1723. doi: 10.1038/s41467-021-25126-0
-
[23]
A Style-Based Generator Architecture for Generative Adversarial Networks.arXiv, March 2019
Tero Karras, Samuli Laine, and Timo Aila. A Style-Based Generator Architecture for Generative Adversarial Networks.arXiv, March 2019. doi: 10.48550/arXiv.1812.04948
-
[24]
Jacob S Prince, Ian Charest, Jan W Kurzawski, John A Pyles, Michael J Tarr, and Kendrick N Kay. Improving the accuracy of single-trial fMRI response estimates using GLMsingle.eLife, 11:e77599, November 2022. ISSN 2050-084X. doi: 10.7554/eLife.77599
-
[25]
Mark Jenkinson, Christian F. Beckmann, Timothy E. J. Behrens, Mark W. Woolrich, and Stephen M. Smith. FSL.NeuroImage, 62(2):782–790, August 2012. ISSN 1053-8119. doi: 10.1016/j.neuroimage.2011.09.015
-
[26]
Nilearn contributors, Ahmad Chamma, Aina Frau-Pascual, Alex Rothberg, Alexandre Abadie, Alexandre Abraham, Alexandre Gramfort, Alexandre Savio, Alexandre Cionca, Alexandre Sayal, Alexis Thual, Alisha Kodibagkar, Amadeus Kanaan, Ana Luisa Pinho, Anand Joshi, Andrés Hoyos Idrobo, Anne-Sophie Kieslinger, Anupriya Kumari, Ariel Rokem, Arthur Mensch, Aswin Vij...
2024
-
[28]
Scikit-learn: Machine Learning in Python.Journal of Machine Learning Research, 12(85):2825–2830, 2011
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake Vander- plas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay. Scikit-learn: Machine Learning in Python.Journal of Machine Learning Resear...
2011
-
[29]
Jeanette A. Mumford, Benjamin O. Turner, F. Gregory Ashby, and Russell A. Poldrack. Deconvolving BOLD activation in event-related designs for multivoxel pattern classifica- tion analyses.Neuroimage, 59(3):2636–2643, February 2012. ISSN 1053-8119. doi: 10.1016/j.neuroimage.2011.08.076
-
[30]
Trevor Hastie, Robert Tibshirani, and Jerome Friedman. Linear Methods for Regression. In Trevor Hastie, Robert Tibshirani, and Jerome Friedman, editors,The Elements of Statistical Learning: Data Mining, Inference, and Prediction, pages 43–99. Springer, New York, NY , 2009. ISBN 978-0-387-84858-7. doi: 10.1007/978-0-387-84858-7_3
-
[31]
SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.arXiv, July 2023
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.arXiv, July 2023. doi: 10.48550/arXiv.2307.01952
-
[32]
Zhou Wang, A.C. Bovik, H.R. Sheikh, and E.P. Simoncelli. Image quality assessment: From error visibility to structural similarity.IEEE Transactions on Image Processing, 13(4):600–612, April 2004. ISSN 1057-7149, 1941-0042. doi: 10.1109/TIP.2003.819861
arXiv 2004
-
[33]
Mingxing Tan and Quoc V . Le. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks.arXiv, September 2020. doi: 10.48550/arXiv.1905.11946
-
[34]
Unsupervised Learning of Visual Features by Contrasting Cluster Assignments.arXiv, January
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, and Armand Joulin. Unsupervised Learning of Visual Features by Contrasting Cluster Assignments.arXiv, January
-
[35]
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. ImageNet Classification with Deep Convolutional Neural Networks. InAdvances in Neural Information Processing Systems, volume 25. Curran Associates, Inc., 2012
2012
-
[36]
Rethink- ing the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethink- ing the Inception Architecture for Computer Vision. In2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 2818–2826, Las Vegas, NV , USA, June 2016. IEEE. ISBN 978-1-4673-8851-1. doi: 10.1109/CVPR.2016.308
-
[37]
Brain decoding: Toward real-time reconstruction of visual perception.arXiv, March 2024
Yohann Benchetrit, Hubert Banville, and Jean-Rémi King. Brain decoding: Toward real-time reconstruction of visual perception.arXiv, March 2024. doi: 10.48550/arXiv.2310.19812
-
[38]
ENIGMA: A Unified Lightweight EEG-to-Image Model for Multi- 14 Subject Visual Decoding
Reese Kneeland, Wangshu Jiang, Ugo Bruzadin Nunes, Si Kai Lee, Paul Steven Scotti, Arnaud Delorme, and Jonathan Xu. ENIGMA: A Unified Lightweight EEG-to-Image Model for Multi- 14 Subject Visual Decoding. InNeurIPS 2025 Workshop on Foundation Models for the Brain and Body, October 2025
2025
-
[39]
Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer.arXiv, October 2025
Roman Beliy, Amit Zalcher, Jonathan Kogman, Navve Wasserman, and Michal Irani. Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer.arXiv, October 2025. doi: 10.48550/arXiv.2510.25976
-
[40]
Nitzan Lubianiker, Christian Paret, Peter Dayan, and Talma Hendler. Neurofeedback through the lens of reinforcement learning.Trends in Neurosciences, 45(8):579–593, August 2022. ISSN 0166-2236. doi: 10.1016/j.tins.2022.03.008
-
[41]
Saampras Ganesan, Nicholas T. Van Dam, Sunjeev K. Kamboj, Aki Tsuchiyagaito, Matthew D. Sacchet, Masaya Misaki, Bradford A. Moffat, Valentina Lorenzetti, and Andrew Zalesky. Neurofeedback Training Facilitates Awareness and Enhances Emotional Well-being Associated with Real-World Meditation Practice: A 7-T MRI Study.Mindfulness, 16(10):2787–2807, 2025. ISS...
-
[42]
Nitzan Lubianiker, Tamar Koren, Meshi Djerasi, Margarita Sirotkin, Neomi Singer, Itamar Jalon, Avigail Lerner, Roi Sar-el, Haggai Sharon, Moni Shahar, Hilla Azulay-Debby, Asya Rolls, and Talma Hendler. Upregulation of reward mesolimbic activity and immune response to vaccination: A randomized controlled trial.Nature Medicine, 32(2):572–581, February 2026....
-
[43]
Ethan F. Oblak, Jarrod A. Lewis-Peacock, and James S. Sulzer. Self-regulation strategy, feedback timing and hemodynamic properties modulate learning in a simulated fMRI neurofeedback environment.PLOS Computational Biology, 13(7):e1005681, July 2017. ISSN 1553-7358. doi: 10.1371/journal.pcbi.1005681
-
[44]
Busch, E
Erica L. Busch, E. Chandra Fincke, Guillaume Lajoie, Smita Krishnaswamy, and Nicholas B. Turk-Browne. Human learning of noninvasive brain–computer interfaces via manifold ge- ometry.Nature Neuroscience, pages 1–9, June 2026. ISSN 1546-1726. doi: 10.1038/ s41593-026-02311-2
2026
-
[45]
Scaling laws for decoding images from brain activity.arXiv, January 2025
Hubert Banville, Yohann Benchetrit, Stéphane d’Ascoli, Jérémy Rapin, and Jean-Rémi King. Scaling laws for decoding images from brain activity.arXiv, January 2025. doi: 10.48550/ arXiv.2501.15322
-
[46]
Christopher DiMattina and Kechen Zhang. Adaptive stimulus optimization for sensory systems neuroscience.Frontiers in Neural Circuits, 7:101, June 2013. ISSN 1662-5110. doi: 10.3389/ fncir.2013.00101
arXiv 2013
-
[47]
Adaptive stimulus selection for optimizing neural population responses
Benjamin Cowley, Ryan Williamson, Katerina Clemens, Matthew Smith, and Byron M Yu. Adaptive stimulus selection for optimizing neural population responses. InAdvances in Neural Information Processing Systems, volume 30. Curran Associates, Inc., 2017
2017
-
[48]
Aditi Jha, Zoe C. Ashwood, and Jonathan W. Pillow. Active Learning for Discrete Latent Variable Models.Neural computation, 36(3):437–474, February 2024. ISSN 0899-7667. doi: 10.1162/neco_a_01646
-
[49]
Dynadiff: Single-stage Decoding of Images from Continuously Evolving fMRI.arXiv, May 2025
Marlène Careil, Yohann Benchetrit, and Jean-Rémi King. Dynadiff: Single-stage Decoding of Images from Continuously Evolving fMRI.arXiv, May 2025. doi: 10.48550/arXiv.2505.14556
-
[50]
Rizvi, Matteo Rosati, Christopher Averill, James L
Josue Ortega Caro, Antonio Henrique de Oliveira Fonseca, Syed A. Rizvi, Matteo Rosati, Christopher Averill, James L. Cross, Prateek Mittal, Emanuele Zappala, Rahul Madhav Dho- dapkar, Chadi Abdallah, and David van Dijk. BrainLM: A foundation model for brain activity recordings. InThe Twelfth International Conference on Learning Representations, October 2023
2023
-
[51]
SwiFT: Swin 4D fMRI Transformer.arXiv, October 2023
Peter Yongho Kim, Junbeom Kwon, Sunghwan Joo, Sangyoon Bae, Donggyu Lee, Yoonho Jung, Shinjae Yoo, Jiook Cha, and Taesup Moon. SwiFT: Swin 4D fMRI Transformer.arXiv, October 2023. doi: 10.48550/arXiv.2307.05916
-
[52]
Zijian Dong, Ruilin Li, Yilei Wu, Thuan Tinh Nguyen, Joanna Su Xian Chong, Fang Ji, Nathanael Ren Jie Tong, Christopher Li Hsian Chen, and Juan Helen Zhou. Brain-JEPA: Brain Dynamics Foundation Model with Gradient Positioning and Spatiotemporal Masking.arXiv, September 2024. doi: 10.48550/arXiv.2409.19407
-
[53]
Zijian Dong, Ruilin Li, Joanna Su Xian Chong, Niousha Dehestani, Yinghui Teng, Yi Lin, Zhizhou Li, Yichi Zhang, Yapei Xie, Leon Qi Rong Ooi, B. T. Thomas Yeo, and Juan Helen 15 Zhou. Brain Harmony: A Multimodal Foundation Model Unifying Morphology and Function into 1D Tokens.arXiv, September 2025. doi: 10.48550/arXiv.2509.24693
-
[54]
Sam Gijsen, Marc-Andre Schulz, and Kerstin Ritter. Brain-Semantoks: Learning Semantic Tokens of Brain Dynamics with a Self-Distilled Foundation Model.arXiv, March 2026. doi: 10.48550/arXiv.2512.11582
-
[55]
Kaplan, Benjamin Warner, Tanishq Mathew Abraham, and Paul S
Connor Lane, Mihir Tripathy, Leema Krishna Murali, Ratna Sagari Grandhi, Shamus Sim Zi Yang, Sam Gijsen, Debojyoti Das, Manish Ram, Utkarsh Kumar Singh, Cesar Kadir Torrico Villanueva, Yuxiang Wei, Will Beddow, Gianfranco Cortés, Suin Cho, Daniel Z. Kaplan, Benjamin Warner, Tanishq Mathew Abraham, and Paul S. Scotti. Scaling Vision Transformers for Functi...
-
[56]
Nachimuthu, Michael Jacob Mendelson, Blake Aaron Richards, Matthew G
Mehdi Azabou, Vinam Arora, Venkataramana Ganesh, Ximeng Mao, Santosh B. Nachimuthu, Michael Jacob Mendelson, Blake Aaron Richards, Matthew G. Perich, Guillaume Lajoie, and Eva L. Dyer. A Unified, Scalable Framework for Neural Population Decoding. InThirty-Seventh Conference on Neural Information Processing Systems, November 2023
2023
-
[57]
Neural Encoding and Decoding at Scale.arXiv, May 2025
Yizi Zhang, Yanchen Wang, Mehdi Azabou, Alexandre Andre, Zixuan Wang, Hanrui Lyu, The International Brain Laboratory, Eva Dyer, Liam Paninski, and Cole Hurwitz. Neural Encoding and Decoding at Scale.arXiv, May 2025. doi: 10.48550/arXiv.2504.08201
-
[58]
Krishna, Ximeng Mao, Mehdi Azabou, Eva L
Avery Hee-Woon Ryoo, Nanda H. Krishna, Ximeng Mao, Mehdi Azabou, Eva L. Dyer, Matthew G. Perich, and Guillaume Lajoie. Generalizable, real-time neural decoding with hybrid state-space models.arXiv, November 2025. doi: 10.48550/arXiv.2506.05320
-
[59]
Harris, K
Charles R. Harris, K. Jarrod Millman, Stéfan J. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fer- nández del Río, Mark Wiebe, Pearu Peterson, Pierre Gérard-Marchant, Kevin She...
2020
-
[60]
Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J
Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J. van der Walt, Matthew Brett, Joshua Wilson, K. Jarrod Millman, Nikolay Mayorov, Andrew R. J. Nelson, Eric Jones, Robert Kern, Eric Larson, C. J. Carey, ˙Ilhan Polat, Yu Feng, Eric W....
2020
-
[61]
John D. Hunter. Matplotlib: A 2D Graphics Environment.Computing in Science & Engineering, 9(3):90–95, May 2007. ISSN 1558-366X. doi: 10.1109/MCSE.2007.55
-
[62]
Pandas-dev/pandas: Pandas
The pandas development team. Pandas-dev/pandas: Pandas. Zenodo, February 2020
2020
-
[63]
Transformers: State-of-the- Art Natural Language Processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. Transformers: State-of-the- Art N...
2020
-
[64]
PyTorch: An Imperative Style, High- Performance Deep Learning Library.arXiv, December 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. PyTorch: An Imperative Style, High- Perform...
-
[65]
Markiewicz, Michael Hanke, Marc-Alexandre Côté, Ben Cipollini, Dimitri Papadopoulos Orfanos, Paul McCarthy, Dorota Jarecka, Christopher P
Matthew Brett, Christopher J. Markiewicz, Michael Hanke, Marc-Alexandre Côté, Ben Cipollini, Dimitri Papadopoulos Orfanos, Paul McCarthy, Dorota Jarecka, Christopher P. Cheng, Eric Larson, Yaroslav O. Halchenko, Michiel Cottaar, Satrajit Ghosh, Demian Wassermann, Stephan Gerhard, Gregory R. Lee, Zvi Baratz, Brendan Moloney, Hao-Ting Wang, Erik Kastman, 16...
2025
-
[66]
Pels, Erik J
Elmar G.M. Pels, Erik J. Aarnoutse, Sacha Leinders, Zac V . Freudenburg, Mariana P. Branco, Benny H. van der Vijgh, Tom J. Snijders, Timothy Denison, Mariska J. Vansteensel, and Nick F. Ramsey. Stability of a chronic implanted brain-computer interface in late-stage amyotrophic lateral sclerosis.Clinical Neurophysiology, 130(10):1798–1803, October 2019. ISSN 1388-
2019
-
[67]
Kevin C. Davis, Kimberley R. Wyse-Sookoo, Fouzia Raza, Benyamin Meschede-Krasa, Noe- line W. Prins, Letitia Fisher, Emery N. Brown, Iahn Cajigas, Michael E. Ivan, Jonathan R. Jagid, and Abhishek Prasad. 5-Year Follow-up of a Fully Implanted Brain-Computer Interface in A Spinal Cord Injury Patient.Journal of neural engineering, 22(2):10.1088/1741–2552/adc4...
doi:10.1088/1741 2025
-
[68]
Peter Mitchell, Sarah C. M. Lee, Peter E. Yoo, Andrew Morokoff, Rahul P. Sharma, Daryl L. Williams, Christopher MacIsaac, Mark E. Howard, Lou Irving, Ivan Vrljic, Cameron Williams, Steven Bush, Anna H. Balabanski, Katharine J. Drummond, Patricia Desmond, Douglas Weber, Timothy Denison, Susan Mathers, Terence J. O’Brien, J. Mocco, David B. Grayden, David S...
arXiv 2023
-
[69]
Aoki, Kei Majima, Yusuke Muraki, and Yukiyasu Kamitani
Ken Shirakawa, Yoshihiro Nagano, Misato Tanaka, Shuntaro C. Aoki, Kei Majima, Yusuke Muraki, and Yukiyasu Kamitani. Spurious reconstruction from brain activity.arXiv, May 2025. doi: 10.48550/arXiv.2405.10078
-
[70]
The Belmont report: Ethical principles and guidelines for the protection of human subjects of research
National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research. The Belmont report: Ethical principles and guidelines for the protection of human subjects of research. https://www.hhs.gov/ohrp/regulations-and-policy/belmont-report/read-the- belmont-report/index.html, 1979
1979
-
[71]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.arXiv, June 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.arXiv, June 2021. doi: 10.48550/arXiv.2010.11929
-
[72]
Adversarial Diffusion Distillation
Axel Sauer, Dominik Lorenz, Andreas Blattmann, and Robin Rombach. Adversarial Diffusion Distillation. In Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sattler, and Gül Varol, editors,Computer Vision – ECCV 2024, volume 15144, pages 87–103. Springer Nature Switzerland, Cham, 2025. ISBN 978-3-031-73015-3 978-3-031-73016-0. doi: 10.1007...
2024
-
[73]
Huettel, Allen W
Scott A. Huettel, Allen W. Song, and Gregory McCarthy.Functional Magnetic Resonance Imaging. Sinauer Associates, Sunderland, Mass, 2004. ISBN 978-0-87893-288-7 978-0-87893- 289-4
2004
-
[74]
Johnson, Alexandre Sayal, jstaph, JohannesWiesner, Jon Clucas, Tinashe Michael Tapera, and justbennet
Andrew Jahn, Dan Levitas, Eric Holscher, John T. Johnson, Alexandre Sayal, jstaph, JohannesWiesner, Jon Clucas, Tinashe Michael Tapera, and justbennet. Andrew- jahn/AndysBrainBook:. Zenodo, January 2022. 17
2022
-
[75]
fMRIPrep: a robust preprocessing pipeline for functional MRI.Nature Methods, 16:111–116, 2019
Oscar Esteban, Christopher Markiewicz, Ross W Blair, Craig Moodie, Ayse Ilkay Isik, Asier Erramuzpe Aliaga, James Kent, Mathias Goncalves, Elizabeth DuPre, Madeleine Snyder, Hi- royuki Oya, Satrajit Ghosh, Jessey Wright, Joke Durnez, Russell Poldrack, and Krzysztof Jacek Gorgolewski. fMRIPrep: a robust preprocessing pipeline for functional MRI.Nature Meth...
-
[76]
Oscar Esteban, Ross Blair, Christopher J. Markiewicz, Shoshana L. Berleant, Craig Moodie, Feilong Ma, Ayse Ilkay Isik, Asier Erramuzpe, James D. Kent, Mathias Goncalves, Elizabeth DuPre, Kevin R. Sitek, Daniel E. P. Gomez, Daniel J. Lurie, Zhifang Ye, Russell A. Poldrack, and Krzysztof J. Gorgolewski. fmriprep.Software, 2018. doi: 10.5281/zenodo.852659
-
[77]
K. Gorgolewski, C. D. Burns, C. Madison, D. Clark, Y . O. Halchenko, M. L. Waskom, and S. Ghosh. Nipype: a flexible, lightweight and extensible neuroimaging data processing frame- work in python.Frontiers in Neuroinformatics, 5:13, 2011. doi: 10.3389/fninf.2011.00013
Pith/arXiv arXiv 2011
-
[78]
Gorgolewski, Oscar Esteban, Christopher J
Krzysztof J. Gorgolewski, Oscar Esteban, Christopher J. Markiewicz, Erik Ziegler, David Gage Ellis, Michael Philipp Notter, Dorota Jarecka, Hans Johnson, Christopher Burns, Alexandre Manhães-Savio, Carlo Hamalainen, Benjamin Yvernault, Taylor Salo, Kesshi Jordan, Mathias Goncalves, Michael Waskom, Daniel Clark, Jason Wong, Fred Loney, Marc Modat, Blake E ...
2018
-
[79]
Andersson, Stefan Skare, and John Ashburner
Jesper L.R. Andersson, Stefan Skare, and John Ashburner. How to correct susceptibility distortions in spin-echo echo-planar images: application to diffusion tensor imaging.NeuroIm- age, 20(2):870–888, 2003. ISSN 1053-8119. doi: 10.1016/S1053-8119(03)00336-7. URL https://www.sciencedirect.com/science/article/pii/S1053811903003367
-
[80]
N. J. Tustison, B. B. Avants, P. A. Cook, Y . Zheng, A. Egan, P. A. Yushkevich, and J. C. Gee. N4itk: Improved n3 bias correction.IEEE Transactions on Medical Imaging, 29(6):1310–1320,
-
[81]
B.B. Avants, C.L. Epstein, M. Grossman, and J.C. Gee. Symmetric diffeomorphic im- age registration with cross-correlation: Evaluating automated labeling of elderly and neu- rodegenerative brain.Medical Image Analysis, 12(1):26–41, 2008. ISSN 1361-8415. doi: 10.1016/j.media.2007.06.004. URL https://www.sciencedirect.com/science/ article/pii/S1361841507000606
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.