Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Leveraging unlabelled data for generalizable neural population decoding

T0 review · 3 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read MOJO shows that adding a masked-autoencoder objective to spike-tokenizing neural decoders improves decoding, especially when labeled data are scarce.

desk verdict MOJO's joint SSL–SL objective consistently beats supervised baselines, but the paper's headline claim about leveraging unlabelled data is contradicted by its own Appendix D.6. read the letter →

arxiv 2607.14086 v1 pith:HYZHHHK5 submitted 2026-07-15 cs.LG q-bio.NC

classification cs.LGq-bio.NC
keywords neuraldecodingself-supervisedlearningmaskedautoencoderspiketokenizationfew-shotfinetuningunitembeddingsbrain-computerinterfaceneuro-foundationmodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper sets out to show that spike-tokenizing neural decoders, which normally require paired behavioral labels, can be pretrained with unlabeled neural data by adding a masked-autoencoder reconstruction objective alongside supervised decoding. The proposed MOJO framework shares almost all parameters between the two pathways, adding only a small overhead (about 231K parameters on a 9.88M-parameter model). Across monkey reaching, mouse vision, mouse decision, and human speech electrocorticography, MOJO variants outperform their purely supervised counterparts under the same finetuning strategy, with the largest gains in few-shot finetuning. The learned per-neuron embeddings also become more interpretable: linear probes predict brain region and single-neuron spike statistics better than supervised-only embeddings. The paper argues this opens a path to using vast unlabeled neural recordings to train scalable neuro-foundation models.

What carries the argument

The central object is the MOJO training objective L = α_SSL L_SSL + α_SL L_SL, a weighted sum of a Poisson negative log-likelihood for reconstructing spike rates from masked latents and a supervised behavioral loss. The self-supervised pathway applies temporal masking to latent tokens and reconstructs spike rates for sampled units; the supervised pathway uses output cross-attention queries to predict behavior. The key design is pathway integration: the two pathways share the input cross-attention output and backbone parameters, so the self-supervised pathway adds only one output cross-attention module. This lets unlabeled neural data improve representations without doubling the model.

What would settle it

Run a pre-registered, code-released head-to-head on several held-out sessions from new animals, finetuning each model with only 2–4 labeled trials (and unlabeled data for MOJO); if MOJO's few-shot advantage over its purely supervised counterpart disappears or reverses across sessions, the central claim collapses.

Watch

Extended reading notes

Core claim

The central claim is that augmenting spike-tokenizing models with the MOJO joint self-supervised-supervised objective improves decoding performance over purely supervised training, especially in label-impoverished few-shot finetuning, and yields more interpretable unit embeddings. Operationally, MOJO trains a masked autoencoder on latent tokens produced by input cross-attention while simultaneously training a supervised behavioral decoder, sharing the input cross-attention and backbone parameters. Evaluated on monkey reaching, mouse vision, mouse decision, and human ECoG speech, MOJO consistently outperforms purely supervised baselines under identical finetuning strategies, achieves the larg

Load-bearing premise

The load-bearing premise is that the public datasets, with the paper's preprocessing and re-splitting, fairly measure cross-session and cross-animal generalization—in particular, that the modified IBL data preparation and the re-run baselines do not systematically favor MOJO.

Editorial extensions

If this is right

  • MOJO-trained decoders can be finetuned with very few labeled trials from a new session or animal, and can use unlabeled trials during finetuning to recover 60–75% of fully supervised performance with only 2–4 labeled calibration trials.
  • Unlabeled data can be used during pretraining: even when up to 90% of pretraining data are unlabeled, MOJO maintains strong decoding and unit-embedding quality.
  • Unit embeddings learned by MOJO support linear-probe classification of brain regions and regression of spike statistics with substantially higher accuracy than purely supervised embeddings.
  • Joint pretraining across monkey reaching and mouse vision transfers positively and scales with model size and data, improving convergence and finetuning on mouse tasks while preserving monkey performance.
  • MOJO generalizes beyond spiking data to human ECoG speech decoding, outperforming supervised POYO and matching a specialized continuous-signal foundation model.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If this result generalizes, brain-computer interface calibration could shift from collecting hours of labeled data per session to a short labeled calibration plus passive unlabeled recording, which would make clinical deployment faster and less burdensome.
  • The finding that self-supervision makes unit embeddings more interpretable suggests masked autoencoding could serve as a generic pretraining scheme for neural recordings where behavior labels are absent or noisy, such as sleep or freely moving behavior.
  • A testable extension would add spatial masking (masking neurons or brain regions) in addition to temporal masking; the paper notes this is absent and could further improve embeddings for unseen neurons or regions.
  • The joint loss coefficients were set to 1 for both pathways across experiments; on more heterogeneous data or with imbalanced label availability, adaptive weighting might be needed, and this could be tested directly.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces MOJO, a joint self-supervised (masked autoencoder) and supervised training objective for spike-tokenizing neural decoders of the POYO family. MOJO is evaluated on monkey reaching, mouse vision and decision-making, and human ECoG speech, with comparisons against supervised spike-tokenizing models and binned baselines (NDT-2/3, NEDS, Du-IN, EEGNet). The authors report consistent gains from MOJO when finetuning to held-out sessions, especially in few-shot regimes, as well as more interpretable unit embeddings. However, Appendix D.6 reports that adding unlabelled data during pretraining does not improve finetuning transfer, which directly conflicts with the title, abstract, and Discussion claim of leveraging unlabelled data for generalizable decoding.

Significance. The empirical scope is substantial: multiple species, modalities, tasks, and backbones, with ablations of mask ratio, pathway integration, and alternative SSL objectives, plus a running-time comparison. The interpretability analyses (region classification, spike-statistic prediction, probe-distance geometry) are thoughtful and go beyond standard decoding metrics. If the central claims were fully supported, MOJO would be a useful step toward using unannotated neural recordings in foundation-model pretraining. However, the manuscript's own D.6 undermines the main unlabelled-data claim, and the lack of released code weakens reproducibility of the baseline comparisons. The results still suggest that joint SSL-SL training improves over purely supervised training on labelled data, but the title and discussion overstate the evidence for leveraging unlabelled data.

major comments (3)
  1. [§3.1 'Unlabelled Data During Pretraining', Appendix D.6, §5] The paper's central claim that MOJO leverages unlabelled data for generalizable decoding is contradicted by its own Appendix D.6 (Fig. 4a): 'having additional unlabelled data during pretraining does not lead to better finetuning performance, especially in UI.' Since finetuning to new sessions is the operational objective of pretraining and the source of the strongest reported gains, the §5 statement that 'We see significant downstream improvement even if up to 90% of the pretraining data is unlabelled' is not supported. Fig. 2b shows only pretraining R2, not downstream finetuning. The authors must reconcile this, e.g., by reporting finetuning results for the unlabelled-pretraining sweep or by reframing the unlabelled-data contribution to the few-shot finetuning setting (Fig. 2a) and interpretability analyses.
  2. [§3.3, §3.4, Appendix F] The reported superiority over external baselines rests on modifications to the baselines and data splits, but no code is released. The paper acknowledges IBL split and re-sorting changes and reruns NEDS, but for NDT-2/3, Du-IN, and EEGNet it states only that original code was used 'with some modifications' (C.1, C.2, C.4). Given the causal evaluation and non-trial-aligned preprocessing choices, the gains could in part be an artifact of evaluation setup. I recommend releasing code and exact preprocessing scripts, and providing a sensitivity analysis (e.g., evaluating NDT-3 under the same causal protocol) to verify that the comparisons are apples-to-apples.
  3. [§3.1 'Few-shot Finetuning', Figure 2a] The few-shot finetuning experiment is the main surviving evidence for unlabelled-data benefit, but the protocol is underspecified. The text says 'incorporating up to 32 trials of unlabelled data' without explaining how the SSL loss is applied during finetuning, which parameters are updated, how unlabelled trials are selected, or how they are balanced with labelled trials. Without these details, the reader cannot determine whether the gain comes from the SSL objective on unlabelled data or from increased training signal/regularization. Please provide the full finetuning protocol and, ideally, an ablation that replaces the unlabelled trials with additional labelled trials or discards them.
minor comments (5)
  1. [Table 2 and Table 21] The notation 'MOJO-POYO(J)' and 'MOJO-POYO-L(J)' is used in the main table but only defined in the text; similarly, 'MOJO-POYO(A)' appears in Table 21. Please add definitions to the table captions for readability.
  2. [§3.4] The phrase '47% improvement for classifying syllables' should report absolute accuracies (e.g., from X% to Y%) to avoid ambiguity between relative and absolute improvement.
  3. [Appendix D.5 and §5] There are typos: 'certain extend' should be 'certain extent' in D.5, and 'underling neural dynamics' should be 'underlying neural dynamics' in §5.
  4. [Appendix F/Table 24] The 'NEDS+bugfix' entry is not defined in the main text; please state what the bug fix is and why it is reported.
  5. [§3.3] The preprocessing changes relative to Zhang et al. [16] (removing trial alignment and firing-rate exclusion, different whisker normalization) are important for interpreting Table 3; they should be stated prominently before the results rather than only in Appendix A.3.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: MOJO's gains are measured on held-out sessions; self-citations are ordinary prior-work dependencies, not load-bearing reductions.

full rationale

MOJO's objective (Section 2.5) is a weighted sum L = α_SSL L_SSL + α_SL L_SL, with L_SSL a Poisson spike-rate NLL and L_SL the behavioural losses. Nothing in this definition encodes the held-out decoding R2 or accuracy that the paper claims to predict; the SSL and SL pathways share the input cross-attention and backbone (Section 2.6), but that is an architectural choice, not an equation-level identity. The central comparisons (Tables 1–3, Figure 3c, few-shot curves in Figure 2a) are empirical evaluations on held-out sessions/trials, and the baselines include externally published methods (NDT-2, NDT-3, NEDS, Du-IN, EEGNet, MLP/GRU) rerun on the same splits where needed, so the reported advantages are not equivalences to the training objective. Citations to the authors' own POYO/POSSM papers [13,14] supply the tokenization scheme, backbone, and evaluation pipeline, but they do not by themselves entail the MOJO result; the MOJO-vs-supervised gap is computed here on unseen data. Appendix D.6 is a genuine internal-evidence problem for the headline 'leverage unlabelled data': it states that adding unlabelled data during pretraining 'does not lead to better finetuning performance, especially in UI', which the Discussion does not reconcile. That is an overclaim or inconsistency about what the data show, not a case of a prediction being identical to its input by construction. No self-definitional step, no fitted parameter renamed as a prediction, and no uniqueness theorem imported from the authors' prior work was found.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

MOJO is empirical and builds directly on prior POYO/POSSM architectures. The central claim rests on dataset validity, the masking premise, and several hand-chosen hyperparameters. No new physical entities are introduced; the mask token is a standard learned parameter. The most consequential choices are the loss coefficients and the temporal-masking assumption, both of which are only partially ablated.

free parameters (6)
  • Temporal mask ratio = 0.5
    Chosen by hand in Section 2.3; ablation D.3.2 shows SL scores are robust across 0.02-0.9, so this parameter is not load-bearing.
  • Loss coefficients αSSL, αSL = 1 (monkey/mouse); SSL coefficient 0.5 (ECoG)
    Set to 1 in Section 2.5 as an empirical default; tuned on ECoG (Appendix C.4). Affects the joint objective balance.
  • SSL unit-query sampling count = up to 10 units
    Appendix B.4 randomly draws up to 10 neural units at masked timesteps; a hand-chosen detail of the SSL reconstruction head.
  • Whisker normalization constant = 2.5
    Appendix A.3: chosen so whisker signal amplitude is similar to wheel; directly affects reported whisker R2 in Table 3.
  • Task loss weights for choice/block = 0.2
    Appendix C.3: empirically weighted to avoid overfitting on mouse decision tasks.
  • Architecture hyperparameters (latents per chunk, layers, hidden dims) = varies per dataset (Tables 5-10)
    Model sizes and latent counts were tuned per dataset; not derived from first principles.
assumptions (5)
  • domain assumption Input cross-attention computed per contiguous time chunk and Bernoulli temporal masking remove all spike information from masked intervals.
    Section 2.3: the SSL reconstruction is only meaningful if masking truly deletes information; leakage through shared position or unit embeddings would contaminate the pretext task.
  • domain assumption Public datasets (Perich, O'Doherty, Churchland, NLB, Allen, IBL, Bouchard ECoG) provide accurate spikes, behaviour, and ECoG labels, and held-out sessions are representative of transfer.
    Section 3 and Appendix A: all conclusions depend on external data quality and split fairness.
  • domain assumption POYO-style unit embeddings can be shared across sessions, animals, and species with only unit/session embedding re-learning (UI).
    Sections 2.1, B.6: cross-session and cross-species generalization claims inherit this representation assumption from prior POYO/POSSM work.
  • domain assumption Poisson negative log-likelihood is an appropriate reconstruction loss for spike rates in the SSL pathway.
    Section 2.4-2.5: spike-rate prediction is scored with Poisson NLL; if neural noise is not Poisson, the SSL objective is misspecified.
  • standard math Scaled dot-product attention, RoPE, and SSM transitions behave as standard implementations.
    Sections 2.1-2.2 and Appendix B rely on standard attention and state-space machinery without new mathematical derivation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Leveraging unlabelled data for generalizable neural population decoding." pith.science (2026). https://pith.science/paper/HYZHHHK5

@misc{pith2026260714086,
  author       = {Pith},
  title        = {Pith review of: Leveraging unlabelled data for generalizable neural population decoding},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HYZHHHK5}},
  note         = {Machine review of arXiv:2607.14086}
}
read the original abstract

Robust and accurate neural decoders are integral to neurotechnologies such as brain-computer interfaces and closed-loop experiments. Recent work has shown that tokenizing neural data at the spike level facilitates multi-session pretraining and delivers state-of-the-art decoding performance. However, current spike-based models are restricted to supervised learning (SL), limiting training to datasets with paired behavioural labels. To address this limitation, we introduce MOJO (Masked autOencoder-based JOint training), a training framework for spike-tokenizing models that jointly leverages self-supervised learning (SSL) via masked autoencoding and SL objectives. We evaluate MOJO on three spiking datasets spanning monkey motor cortex during reaching tasks and multi-regional mouse recordings during vision and decision making tasks, demonstrating superior performance over purely SL-trained models. This improvement is especially pronounced when training with limited labelled data, particularly in few-shot finetuning, where only a small amount of labelled data from a new session is available. Incorporating SSL also yields more interpretable neuronal representations, improving performance on brain region classification and spike-statistics prediction without explicit optimization for these tasks. We further show that MOJO generalizes beyond spiking data to human electrocorticography during speech, where it continues to outperform purely SL-trained models and achieves performance comparable to neuro-foundation models (NFMs) designed specifically for continuous signals. Overall, augmenting spike-tokenizing models with SSL improves performance in label-impoverished settings and enables the use of unlabelled data across various tasks and species, while generalizing to other neural modalities. These results suggest a path towards more flexible and scalable data usage when training NFMs.

Figures

Figures reproduced from arXiv: 2607.14086 by the authors.

Figure 1
Figure 1. Model and task schematics. (a) Schematic showing a POYO-style model augmented with MOJO. Latent representations extracted from tokenized neural data are simultaneously used for supervised learning (SL) and self-supervised learning (SSL). The former is carried out by minimizing error in predicted behaviour while the latter is carried out by reconstructing spike counts from masked latents. (b-e) Schematics describing … view at source ↗
Figure 2
Figure 2. Leveraging unlabelled data for fine￾tuning and pretraining. (a) MOJO improves few￾shot finetuning performance over standard SL and leverages additional unlabelled data to improve per￾formance further. (b) MOJO can improve decoding performance by exploiting all available unlabelled data even when little labelled data is available. Few-shot Finetuning. MOJO also enables ro￾bust and efficient few-shot adaptation. Figur… view at source ↗
Figure 3
Figure 3. Mouse brain region classification and human speech decoding results. (a-b) Confusion matrices for (a) 3-class and (b) 7-class brain region classification of neurons in mice from the Allen visual coding dataset. (c) Classification accuracy of syllables, consonants, and vowels for speech decoding from human electrocorticography. For (a-c), we report the mean accuracy over sessions averaged across 5 seeds. (**,***): p … view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Additional results on pretraining with unlabelled data. (a) Finetuning performance with additional unlabelled data during pretraining. (b) Brain region classification performance with additional unlabelled data during pretraining [PITH_FULL_IMAGE:figures/full_fig_p023…
Figure 5
Figure 5. Figure 5: Confusion matrix for mouse brain region classifications across different models. Left: from model with paired data from Allen datasets; Middle: Left + additional unlabelled data from Allen datasets; Right: Middle + monkey dataset. latent sequence (MOJO-POYO and MOJO-PO…
Figure 6
Figure 6. Figure 6: Unit-level metadata is encoded better by MOJO than by pure-SL POYO. (a) Linear-probe regression of 18 single-neuron spike-statistic features (ISI moments, firing rate, variability indices, gamma-distribution fit parameters, and band-limited PSDs of the binned spike tra…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NeuroPB: Scaling Neural Decoding with Pretrained Behavioral Representations

    cs.LG 2026-08 conditional novelty 7.0 of 10

    Behavioral pretraining on macaque and robotic trajectories improves neural trajectory decoding and reduces calibration data needs across sessions, subjects, and tasks.

Reference graph

Works this paper leans on

59 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [1]

    Machine Learning for Neural Decoding

    J. I. Glaser, A. S. Benjamin, R. H. Chowdhury, M. G. Perich, L. E. Miller, and K. P. Kording. “Machine Learning for Neural Decoding”.eNeuro7.4 (2020)

  2. [2]

    Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation

    K. Cho, B. van Merriënboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y . Bengio. “Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation”. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Doha, Qatar, 2014, pp. 1724–1734

  3. [3]

    BRAND: a platform for closed-loop experiments with deep network models

    Y . H. Ali, K. Bodkin, M. Rigotti-Thompson, K. Patel, N. S. Card, B. Bhaduri, S. R. Nason-Tomaszewski, D. M. Mifsud, X. Hou, C. Nicolas, et al. “BRAND: a platform for closed-loop experiments with deep network models”.Journal of Neural Engineering21.2 (2024), p. 026046

  4. [4]

    Making brain–machine interfaces robust to future neural variability

    D. Sussillo, S. D. Stavisky, J. C. Kao, S. I. Ryu, and K. V . Shenoy. “Making brain–machine interfaces robust to future neural variability”.Nature Communications7.1 (2016)

  5. [5]

    Attention is All you Need

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. “Attention is All you Need”.Advances in Neural Information Processing Systems. V ol. 30. 2017

  6. [6]

    Mamba: Linear-Time Sequence Modeling with Selective State Spaces

    A. Gu and T. Dao. “Mamba: Linear-Time Sequence Modeling with Selective State Spaces”. 2024. arXiv: 2312.00752

  7. [7]

    Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

    T. Dao and A. Gu. “Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality”.Proceedings of the 41st International Conference on Machine Learning. V ol. 235. 2024, pp. 10041–10071

  8. [8]

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”.International Conference on Learning Representations. 2021

Show all 59 references
  1. [9]

    Large Brain Model for Learning Generic Representations with Tremen- dous EEG Data in BCI

    W. Jiang, L. Zhao, and B.-l. Lu. “Large Brain Model for Learning Generic Representations with Tremen- dous EEG Data in BCI”.The Twelfth International Conference on Learning Representations. 2024

  2. [10]

    Jamba: Hybrid Transformer-Mamba Language Models

    B. Lenz, O. Lieber, A. Arazi, A. Bergman, A. Manevich, B. Peleg, B. Aviram, C. Almagor, C. Fridman, D. Padnos, et al. “Jamba: Hybrid Transformer-Mamba Language Models”.The Thirteenth International Conference on Learning Representations. 2025

  3. [11]

    Brant: Foundation Model for Intracranial Neural Signal

    D. Zhang, Z. Yuan, Y . Yang, J. Chen, J. Wang, and Y . Li. “Brant: Foundation Model for Intracranial Neural Signal”.Advances in Neural Information Processing Systems. 2023

  4. [12]

    BrainBERT: Self- supervised representation learning for intracranial recordings

    C. Wang, V . Subramaniam, A. U. Yaari, G. Kreiman, B. Katz, I. Cases, and A. Barbu. “BrainBERT: Self- supervised representation learning for intracranial recordings”.The Eleventh International Conference on Learning Representations. 2023

  5. [13]

    A Unified, Scalable Framework for Neural Population Decoding

    M. Azabou, V . Arora, V . Ganesh, X. Mao, S. Nachimuthu, M. Mendelson, B. Richards, M. Perich, G. Lajoie, and E. Dyer. “A Unified, Scalable Framework for Neural Population Decoding”.Advances in Neural Information Processing Systems. V ol. 36. 2023, pp. 44937–44956

  6. [14]

    Gener- alizable, real-time neural decoding with hybrid state-space models

    A. H.-W. Ryoo, N. H. Krishna, X. Mao, M. Azabou, E. L. Dyer, M. G. Perich, and G. Lajoie. “Gener- alizable, real-time neural decoding with hybrid state-space models”.Advances in Neural Information Processing Systems. 2025

  7. [15]

    Representation learning for neural population activity with Neural Data Transformers

    J. Ye and C. Pandarinath. “Representation learning for neural population activity with Neural Data Transformers”.Neurons, Behavior, Data analysis, and Theory5.3 (2021), pp. 1–18

  8. [16]

    Neural Encoding and Decoding at Scale

    Y . Zhang, Y . Wang, M. Azabou, A. Andre, Z. Wang, H. Lyu, I. B. Laboratory, E. L. Dyer, L. Paninski, and C. L. Hurwitz. “Neural Encoding and Decoding at Scale”.Proceedings of the 42nd International Conference on Machine Learning. V ol. 267. Proceedings of Machine Learning Res...

  9. [17]

    Deep Learning

    I. Goodfellow, Y . Bengio, and A. Courville. “Deep Learning”.http://www.deeplearningbook.org. MIT Press, 2016

  10. [18]

    Language Models are Few-Shot Learners

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. “Language Models are Few-Shot Learners”.Advances in Neural Information Processing Systems. V ol. 33. 2020, pp. 1877–1901

  11. [19]

    BERT: Pre-training of Deep Bidirectional Transform- ers for Language Understanding

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. “BERT: Pre-training of Deep Bidirectional Transform- ers for Language Understanding”.Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies...

  12. [20]

    A simple framework for contrastive learning of visual representations

    T. Chen, S. Kornblith, M. Norouzi, and G. Hinton. “A simple framework for contrastive learning of visual representations”.Proceedings of the 37th International Conference on Machine Learning. 2020

  13. [21]

    Self- Supervised Learning from Images with a Joint-Embedding Predictive Architecture

    M. Assran, Q. Duval, I. Misra, P. Bojanowski, P. Vincent, M. G. Rabbat, Y . LeCun, and N. Ballas. “Self- Supervised Learning from Images with a Joint-Embedding Predictive Architecture”. 2023, pp. 15619– 15629

  14. [22]

    POCO: Scalable Neural Forecasting through Population Conditioning

    Y . Duan, H. T. Chaudhry, M. B. Ahrens, C. D. Harvey, M. G. Perich, K. Deisseroth, and K. Rajan. “POCO: Scalable Neural Forecasting through Population Conditioning”.The Thirty-ninth Annual Conference on Neural Information Processing Systems. 2026

  15. [23]

    OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens

    K. F. Willeke, P. Turishcheva, A. Gilbert, G. Chakrabarty, H. A. Bedel, P. G. Fahey, Y . Qiu, M. A. Weis, M. Vystrˇcilová, T. Muhammad, et al. “OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens”.The Fourteenth International Conference ...

  16. [24]

    Masked Autoencoders Are Scalable Vision Learners

    K. He, X. Chen, S. Xie, Y . Li, P. Dollár, and R. Girshick. “Masked Autoencoders Are Scalable Vision Learners”.2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 15979–15988

  17. [25]

    wav2vec 2.0: a framework for self-supervised learning of speech representations

    A. Baevski, H. Zhou, A. Mohamed, and M. Auli. “wav2vec 2.0: a framework for self-supervised learning of speech representations”.Proceedings of the 34th International Conference on Neural Information Processing Systems. Vancouver, BC, Canada, 2020

  18. [26]

    Multi-session, multi-task neural decoding from distinct cell-types and brain regions

    M. Azabou, K. X. Pan, V . Arora, I. J. Knight, E. L. Dyer, and B. A. Richards. “Multi-session, multi-task neural decoding from distinct cell-types and brain regions”.The Thirteenth International Conference on Learning Representations. 2025

  19. [27]

    Neural Data Transformer 2: Multi-context Pretraining for Neural Spiking Activity

    J. Ye, J. Collinger, L. Wehbe, and R. Gaunt. “Neural Data Transformer 2: Multi-context Pretraining for Neural Spiking Activity”.Advances in Neural Information Processing Systems. V ol. 36. 2023, pp. 80352– 80374

  20. [28]

    A Generalist Intracortical Motor Decoder

    J. Ye, F. Rizzoglio, X. Ma, A. Smoulder, H. Mao, G. H. Blumenthal, W. Hockeimer, N. G. Kunigk, D. D. Moore, P. J. Marino, et al. “A Generalist Intracortical Motor Decoder”.Advances in Neural Information Processing Systems. 2025

  21. [29]

    Sharing neurophysiology data from the Allen Brain Observa- tory

    S. E. de Vries, J. H. Siegle, and C. Koch. “Sharing neurophysiology data from the Allen Brain Observa- tory”.eLife12 (July 2023). Ed. by M. Meister, J. I. Gold, M. Meister, and J. L. Chen, e85550

  22. [30]

    Survey of spiking in the mouse visual system reveals functional hierarchy

    J. H. Siegle, X. Jia, S. Durand, S. Gale, C. Bennett, N. Graddis, G. Heller, T. K. Ramirez, H. Choi, J. A. Luviano, et al. “Survey of spiking in the mouse visual system reveals functional hierarchy”.Nature 592 (2021), pp. 86–92

  23. [31]

    Transcriptomic cell type structures in vivo neuronal activity across multiple timescales

    A. Schneider, M. Azabou, L. McDougall-Vigier, D. F. Parks, S. Ensley, K. Bhaskaran-Nair, T. Nowakowski, E. L. Dyer, and K. B. Hengen. “Transcriptomic cell type structures in vivo neuronal activity across multiple timescales”.Cell Reports42.4 (2023), p. 112318

  24. [32]

    Reproducibility of in vivo electrophysiological measurements in mice

    I. B. L. IBL, K. Banga, J. Benson, J. Bhagat, D. Biderman, D. Birman, N. Bonacchi, S. A. Bruijns, K. Buchanan, R. A. Campbell, et al. “Reproducibility of in vivo electrophysiological measurements in mice”.eLife(Mar. 2025)

  25. [33]

    Functional organization of human sensori- motor cortex for speech articulation

    K. E. Bouchard, N. Mesgarani, K. Johnson, and E. F. Chang. “Functional organization of human sensori- motor cortex for speech articulation”.Nature. 7441. 2013, pp. 327–332

  26. [34]

    Human ECoG speaking consonant-vowel syllables

    K. E. Bouchard and E. F. Chang. “Human ECoG speaking consonant-vowel syllables”. 2019

  27. [35]

    Du-IN: Discrete units-guided mask modeling for decoding speech from Intracranial Neural signals

    H. Zheng, H. Wang, W. Jiang, Z. Chen, L. He, P. Lin, P. Wei, G. Zhao, and Y . Liu. “Du-IN: Discrete units-guided mask modeling for decoding speech from Intracranial Neural signals”.The Thirty-eighth Annual Conference on Neural Information Processing Systems. 2024

  28. [36]

    A New Approach to Linear Filtering and Prediction Problems

    R. E. Kalman. “A New Approach to Linear Filtering and Prediction Problems”.Journal of Basic Engineering82.1 (1960), pp. 35–45

  29. [37]

    Neural Decoding of Cursor Motion Using a Kalman Filter

    W. Wu, M. Black, Y . Gao, M. Serruya, A. Shaikhouni, J. Donoghue, and E. Bienenstock. “Neural Decoding of Cursor Motion Using a Kalman Filter”.Advances in Neural Information Processing Systems. V ol. 15. 2002

  30. [38]

    Comparison of brain–computer interface decoding algorithms in open-loop and closed-loop control

    S. Koyama, S. M. Chase, A. S. Whitford, M. Velliste, A. B. Schwartz, and R. E. Kass. “Comparison of brain–computer interface decoding algorithms in open-loop and closed-loop control”.Journal of Computational Neuroscience29.1–2 (2010), pp. 73–87. 11

  31. [39]

    Principled BCI Decoder Design and Parameter Selection Using a Feedback Control Model

    F. R. Willett, D. R. Young, B. A. Murphy, W. D. Memberg, C. H. Blabe, C. Pandarinath, S. D. Stavisky, P. Rezaii, J. Saab, B. L. Walter, et al. “Principled BCI Decoder Design and Parameter Selection Using a Feedback Control Model”.Scientific Reports9.1 (2019)

  32. [40]

    A recurrent neural network for closed-loop intracortical brain–machine interface decoders

    D. Sussillo, P. Nuyujukian, J. M. Fan, J. C. Kao, S. D. Stavisky, S. Ryu, and K. Shenoy. “A recurrent neural network for closed-loop intracortical brain–machine interface decoders”.Journal of Neural Engineering 9.2 (Mar. 2012), p. 026027

  33. [41]

    A Recurrent Latent Variable Model for Sequential Data

    J. Chung, K. Kastner, L. Dinh, K. Goel, A. Courville, and Y . Bengio. “A Recurrent Latent Variable Model for Sequential Data”. 2016. arXiv:1506.02216

  34. [42]

    Inferring single-trial neural population dynamics using sequential auto-encoders

    C. Pandarinath, D. J. O’Shea, J. Collins, R. Jozefowicz, S. D. Stavisky, J. C. Kao, E. M. Trautmann, M. T. Kaufman, S. I. Ryu, L. R. Hochberg, et al. “Inferring single-trial neural population dynamics using sequential auto-encoders”.Nature Methods15.10 (2018), pp. 805–815

  35. [43]

    Efficiently Modeling Long Sequences with Structured State Spaces

    A. Gu, K. Goel, and C. Re. “Efficiently Modeling Long Sequences with Structured State Spaces”. International Conference on Learning Representations. 2022

  36. [44]

    Towards a

    Y . Zhang, Y . Wang, D. M. Jiménez-Benetó, Z. Wang, M. Azabou, B. Richards, R. Tung, O. Winter, T. I. B. Laboratory, E. Dyer, et al. “Towards a "Universal Translator" for Neural Dynamics at Single-Cell, Single-Spike Resolution”.Advances in Neural Information Processing Systems...

  37. [45]

    Know Thyself by Knowing Others: Learning Neuron Identity from Population Context

    V . Arora, D. Lachi, I. J. Knight, M. Azabou, B. A. Richards, C. L. Hurwitz, J. Siegle, and E. L. Dyer. “Know Thyself by Knowing Others: Learning Neuron Identity from Population Context”.Advances in Neural Information Processing Systems. 2025

  38. [46]

    Long-term recordings of motor and premotor cortical spiking activity during reaching in monkeys

    M. G. Perich, L. E. Miller, M. Azabou, and E. L. Dyer. “Long-term recordings of motor and premotor cortical spiking activity during reaching in monkeys”. 2025

  39. [47]

    Accurate decoding of reaching movements from field potentials in the absence of spikes

    R. D. Flint, E. W. Lindberg, L. R. Jordan, L. E. Miller, and M. W. Slutzky. “Accurate decoding of reaching movements from field potentials in the absence of spikes”.Journal of Neural Engineering9.4 (2012), p. 046006

  40. [48]

    Nonhuman Primate Reaching with Multichannel Sensorimotor Cortex Electrophysiology

    J. E. O’Doherty, M. M. B. Cardoso, J. G. Makin, and P. N. Sabes. “Nonhuman Primate Reaching with Multichannel Sensorimotor Cortex Electrophysiology”. Zenodo:10.5281/zenodo.3854034. 2020

  41. [49]

    Neural population dynamics during reaching

    M. M. Churchland, J. P. Cunningham, M. T. Kaufman, J. D. Foster, P. Nuyujukian, S. I. Ryu, and K. V . Shenoy. “Neural population dynamics during reaching”.Nature487.7405 (2012), pp. 51–56

  42. [50]

    Neural Latents Benchmark ‘21: Evaluating latent variable models of neural population activity

    F. Pei, J. Ye, D. Zoltowski, A. Wu, R. Chowdhury, H. Sohn, J. O’Doherty, K. V . Shenoy, M. Kaufman, M. Churchland, et al. “Neural Latents Benchmark ‘21: Evaluating latent variable models of neural population activity”.Proceedings of the Neural Information Processing Systems Tr...

  43. [51]

    RoFormer: Enhanced transformer with Rotary Position Embedding

    J. Su, M. Ahmed, Y . Lu, S. Pan, W. Bo, and Y . Liu. “RoFormer: Enhanced transformer with Rotary Position Embedding”.Neurocomputing568 (2024), p. 127063

  44. [52]

    Perceiver IO: A General Architecture for Structured Inputs & Outputs

    A. Jaegle, S. Borgeaud, J.-B. Alayrac, C. Doersch, C. Ionescu, D. Ding, S. Koppula, D. Zoran, A. Brock, E. Shelhamer, et al. “Perceiver IO: A General Architecture for Structured Inputs & Outputs”.International Conference on Learning Representations. 2022

  45. [53]

    TorchEEGEMO: A deep learning toolbox towards EEG-based emotion recognition

    Z. Zhang, S. -h. Zhong, and Y . Liu. “TorchEEGEMO: A deep learning toolbox towards EEG-based emotion recognition”.Expert Systems with Applications(2024), p. 123550

  46. [54]

    Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

    Y . You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli, X. Song, J. Demmel, K. Keutzer, and C.-J. Hsieh. “Large Batch Optimization for Deep Learning: Training BERT in 76 minutes”. 2020. arXiv: 1904.00962

  47. [55]

    Decoupled Weight Decay Regularization

    I. Loshchilov and F. Hutter. “Decoupled Weight Decay Regularization”.International Conference on Learning Representations. 2019

  48. [56]

    xLSTM: Extended Long Short-Term Memory

    M. Beck, K. Pöppel, M. Spanring, A. Auer, O. Prudnikova, M. Kopp, G. Klambauer, J. Brandstetter, and S. Hochreiter. “xLSTM: Extended Long Short-Term Memory”.Advances in Neural Information Processing Systems. 2024

  49. [57]

    Representation Learning with Contrastive Predictive Coding

    A. van den Oord, Y . Li, and O. Vinyals. “Representation Learning with Contrastive Predictive Coding”. CoRRabs/1807.03748 (2018). arXiv:1807.03748

  50. [58]

    High-performance brain-to-text communication via handwriting

    F. R. Willett, D. T. Avansino, L. R. Hochberg, J. M. Henderson, and K. V . Shenoy. “High-performance brain-to-text communication via handwriting”.Nature593.7858 (2021), pp. 249–254. 12 Supplementary Material A Additional Details on Datasets A.1 Monkey Reaching Tasks For experi...

  51. [59]

    neural unit

    in the following three aspects: Firstly, we didn’t enforce trial-alignment during training, and the training intervals were extended to the union of 1) the pre-determined intervals and 2) the full trial durations including additional 200ms before and after; Secondly, we do not...

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.