{"id":"313f2eaf-1775-4e07-8e85-553d5848295e","arxiv_id":"2412.19845","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A dissertation reporting that a hierarchical VAE trained on synthetic optic flow predicts macaque MT neuron responses better than prior models, and that overlapping community analysis of simultaneous fMRI and Ca2+ imaging finds about half of mouse cortical regions in multiple networks.","lead":"This dissertation tests whether hierarchical generative models can reproduce motion processing in the primate brain, and whether resting brain networks in mice are overlapping rather than disjoint. It reports that a hierarchical variational autoencoder trained on synthetic optic flow outperforms earlier models in predicting MT neuron responses, and that roughly half of mouse cortical regions belong to multiple networks.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Per-figure β selection (Sec. 4.9.5) threatens the reported 2x predictive gain: the gain may reflect outcome-guided hyperparameter choice rather than robust model superiority.","rationale":"The reader's weakest_assumption focuses on ecological validity of ROFL. While that is a real limitation for the interpretive leap to 'the brain's understanding', it does not directly undermine the quantitative 2x gain claim, because the comparison could still be fair on the tested stimulus family. In contrast, the per-figure β selection disclosed in Section 4.9.5 threatens the internal validity of the comparison: if β is chosen after seeing the neural scores, the model's advantage could be an artifact of selection. The reader's conditions already require a pre-registered β rule, so this stress-test reinforces that condition rather than changing the verdict. The verdict remains CONDITIONAL; no verdict change is needed, but the existing conditional should be understood as mandatory, not optional.","tokens_in":47182,"tokens_out":4951,"duration_ms":42314,"concrete_test":"Recompute the MT brain-alignment comparison (Tables 4.3–4.6) with a single fixed β that is selected a priori on a validation split using only reconstruction or disentanglement criteria (no neural data), and evaluate against a re-fit Nishimoto-Gallant baseline on identical held-out neurons and stimulus sets. If the 2x gain does not survive a pre-registered β rule, the headline claim should be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 1.8.4 asserts that the hierarchical VAE 'surpasses the previous state-of-the-art model by a gain of over 2x in predictive power' on macaque MT neurons. Section 4.9.5, however, discloses that β values are chosen separately for different figures. If β is selected after inspecting brain-alignment results (e.g., Fig. 4.16), the reported comparison is not a fixed-model evaluation but a search over a hyperparameter that strongly controls the trade-off between reconstruction and disentanglement. The relevant contrast is therefore between an unconstrained family of cNVAE variants and a single fixed Nishimoto-Gallant baseline, so the 2x gain could be inflated by favorable β selection. This is more load-bearing than the ecological validity of ROFL: even if the synthetic training data are simplified, a fair internal comparison could still support the architectural conclusion, whereas per-figure β tuning undermines the internal comparison itself.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript, an arXiv deposit of a PhD dissertation, presents two projects. Project 1 (Chapter 4) develops a hierarchical variational autoencoder, cNVAE, trained on a synthetic optic-flow dataset (ROFL), and claims that its latent representations not only separate self-motion and object-motion causes without extra-retinal signals but also predict macaque MT neuron responses with a gain of over 2x over the previous state of the art (Section 1.8.4). Project 2 (Chapter 5) applies a mixed-membership stochastic blockmodel to simultaneous fMRI and wide-field calcium imaging data in mice and claims that roughly half of cortical regions belong to multiple overlapping communities (abstract and Section 1.13.3). The introduction and background chapters are fully provided, but the two results chapters are truncated in the version under review, so the evidence for these claims is not available in the submitted text.","tokens_in":47198,"tokens_out":5539,"duration_ms":47699,"significance":"If the results hold, Project 1 would provide an important proof-of-concept that unsupervised hierarchical generative models can perform optic-flow parsing and serve as encoding models of primate MT neurons, a domain previously dominated by supervised mechanistic models. The use of external macaque MT recordings to ground the brain-alignment claim is a methodological strength, as is the stated intention to release code and data (Sections 4.8 and 5.5). Project 2 addresses a timely question about overlapping functional organization in the rodent cortex using a rare simultaneous fMRI/Ca2+ dataset. However, because both results chapters are absent from the reviewed text, the actual quantitative support for these claims cannot currently be assessed.","major_comments":[{"comment":"The central quantitative claims—the over-2x predictive gain of the hierarchical VAE over the Nishimoto-Gallant model and the roughly 50% overlap of mouse cortical regions—are asserted in the introduction and abstract, but Chapters 4 and 5 are not present in the submitted text; only their table-of-contents entries and section headings are available. Consequently, the model architecture, training procedure, evaluation metrics, statistical comparisons, and robustness analyses that would substantiate these claims cannot be inspected. This is a load-bearing incompleteness for the manuscript as submitted.","section":"Chapters 4 and 5 (TOC; Sections 1.8.4, 1.13.3)"},{"comment":"The over-2x predictive gain claim is potentially confounded by the per-figure β selection disclosed in Section 4.9.5. If β values were chosen after inspecting brain-alignment results (e.g., Fig. 4.16), then the reported gain compares an outcome-selected member of the cNVAE family against a single fixed Nishimoto-Gallant baseline, rather than evaluating a fixed model. Please report the β selection rule, state whether it was blind to the MT recording data, and include a sensitivity analysis of the predictive gain across β values.","section":"§4.9.5; §1.8.4"},{"comment":"The ROFL synthetic dataset contains one fixating observer, one object of fixed size, and known depth distributions. The transfer of representations learned on this simplified stimulus family to real macaque MT responses assumes that its motion statistics capture ecologically relevant structure. The paper should provide evidence about how the disentanglement and MT-alignment results vary with object size, number of objects, presence of pursuit eye movements, and depth variability, or otherwise justify the ecological validity of ROFL.","section":"§4.3, §4.12"},{"comment":"The Chapter 5 claim that about 50% of cortical regions belong to multiple communities depends on the chosen number of communities K and on the functional connectivity graph threshold. The visible text does not include the robustness analyses that would show this overlap estimate is stable across these choices; the full chapter should report such analyses or qualify the claim accordingly.","section":"§5.4.2, §5.4.3"}],"minor_comments":[{"comment":"The phrase 'hierarchical inference underlines the brain's understanding' should read 'underlies the brain's understanding'; also, 'V AE' is written with an unusual space throughout, and should be standardized to 'VAE'.","section":"Abstract; §1.8.4"},{"comment":"The caption contains a typo: 'Foodforward (ascending) and feedback' should be 'Feedforward (ascending) and feedback'.","section":"Figure 1.11 caption"},{"comment":"The claim of 'over 2x in predictive power' would be clearer if it specified the evaluation metric (e.g., variance explained, correlation coefficient) and the exact comparison protocol used for the baseline model.","section":"§1.8.4"},{"comment":"The chapter title 'Variational Inference & Variational Autoencoders (V AE)' should use the standard abbreviation 'VAE' consistently, both in the title and in the body text.","section":"Chapter 3 title"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is a doctoral dissertation rather than a focused research article, and the two projects are largely independent. If the journal's scope is original research articles, the authors should consider splitting the work into two papers, each with full methods and results. More importantly, the version provided for review lacks the results chapters, which makes it impossible to verify the headline claims; this barrier must be resolved before a definitive editorial decision can be made."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing you should know up front: this is a PhD dissertation, and the two results chapters (Chapters 4 and 5) are truncated in the arXiv posting. What is visible—the abstract, the full introduction, and the historical/mathematical background—is carefully written and honestly hedged. But the load-bearing numbers, the 'over 2x' predictive gain over the Nishimoto-Gallant MT model and the ~50% of mouse cortical regions in multiple communities, are claims we simply cannot check from the text under review.\n\nWhat is genuinely new: the ROFL synthetic optic flow dataset and the cNVAE architecture are real contributions, and the attempt to test an unsupervised hierarchical generative model against macaque MT recordings is a meaningful step beyond the usual feedforward CNN-based encoding models. The second project, applying a mixed-membership stochastic blockmodel to simultaneous fMRI and Ca2+ imaging, is also a sensible way to ask whether overlapping community structure is a BOLD artifact or a general property of cortical organization. The writing is refreshingly direct, and the author explicitly lists limitations in Sections 4.7.1 and 6.1.1.\n\nThe soft spots are real, and the stress-test note points at the most important one. Section 4.9.5 discloses that beta (the VAE loss weight) is chosen separately for different figures. If those choices were made after inspecting brain-alignment results, the 2x gain is not a comparison between a fixed cNVAE and a fixed baseline; it is a comparison between a search over a hyperparameter family and a single Nishimoto-Gallant fit. That undermines the internal validity of the headline claim more than the ecological simplicity of ROFL does, because even a simplified synthetic world could support a fair architectural comparison if the hyperparameters were set in advance. The stress-test note is right on this point.\n\nTwo other fragilities are visible from the front matter. The baseline comparability—whether the 2x gain was computed under identical cross-validation splits and stimulus sets—is not shown in the text we have. And the Chapter 5 overlap percentage depends on the chosen number of communities K and the graph threshold; those choices are reported in the supplementary sections, but the robustness analysis is not available to us. None of this means the results are wrong; it means the posted text does not contain the evidence needed to judge them.\n\nWho is this for? Vision scientists and computational neuroscientists interested in generative models of neural coding, and network neuroscientists working on overlapping brain networks. The paper deserves a serious referee, but only after the full results chapters are made available and the beta selection rule is either fixed a priori or shown to be insensitive across a reasonable range. If it crosses your desk as a journal submission, send it to review with that insistence. As an arXiv posting, it is an interesting draft, not a finished paper.","headline":"A carefully framed dissertation whose headline claims are unverifiable in the posted text; the per-figure beta selection is a real threat to the 2x predictive gain.","tokens_in":47913,"tokens_out":2828,"would_cite":false,"duration_ms":51247,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A hierarchical generative model trained only on synthetic optic flow predicts macaque MT neurons more than twice as well as the prior benchmark.","keywords":["hierarchical variational autoencoder","optic flow parsing","MT neuron encoding models","self-motion versus object-motion","mixed-membership stochastic blockmodel","overlapping cortical networks","wide-field calcium imaging","resting-state fMRI"],"falsifier":"Train the same model on ROFL variants with two independently moving objects, pursuit or saccadic eye movements, occlusions, or natural depth variation, and re-measure the MT alignment gain. If the >2x improvement over the mechanistic baseline disappears or falls below significance, the finding is specific to the simplified stimulus family rather than a general principle of hierarchical inference.","tokens_in":46765,"feed_emoji":"🧠","tokens_out":8627,"duration_ms":95106,"temperature":0.7,"pith_summary":"This dissertation claims that a hierarchical variational autoencoder—a generative neural network trained without labels on synthetic retinal optic flow—learns to decompose the flow into self-motion and object-motion, and that its internal representations predict macaque MT neuron responses with more than double the predictive power of the leading mechanistic model. A second project applies the same inferential framework to spontaneous mouse cortex activity, showing that simultaneously recorded fMRI and calcium-imaging signals are best described by overlapping communities, with roughly half of cortical regions belonging to more than one network. If the motion claim holds, it suggests that hierarchical inference, rather than supervised feature fitting, underlies the visual system's ability to separate the observer's own movement from movement in the world. If the network claim holds, standard disjoint parcellations of the cortex systematically understate how multifunctional many regions are, and the overlapping organization appears in both hemodynamic and more neuron-specific optical signals.","feed_headline":"Generative model beats prior MT-neuron benchmark 2x","feed_subtitle":"Trained on synthetic optic flow alone, it splits self-motion from object motion and predicts macaque MT firing.","key_machinery":"The central object is the compressed hierarchical variational autoencoder (cNVAE), which stacks multiple stochastic latent layers and is trained by maximizing the evidence lower bound (ELBO), a variational free-energy objective that operationalizes Helmholtz's 'perception as unconscious inference.' The hierarchy is what lets the model capture multi-scale causes in the synthetic ROFL optic-flow data: lower latents handle local flow structure and higher latents encode object motion and self-motion. The other load-bearing mechanism is the mixed-membership stochastic blockmodel with variational inference, which assigns each cortical region a vector of membership strengths across K overlapping communities rather than a single community label; the overlap fraction and membership entropy derived from that vector carry the Chapter 5 results.","core_discovery":"In the author's terms, the dissertation demonstrates that optic flow parsing is possible if a neural network is (a) structured hierarchically and (b) trained with an objective function based on Helmholtzian inference. The compressed hierarchical VAE (cNVAE) is trained on the ROFL synthetic world, which contains a fixating observer, translational and rotational self-motion, and one independently moving object; without any supervision, the latents untangle object position and velocity from self-motion components more sharply than a comparable flat VAE. Evaluated as an encoding model against macaque MT neuron recordings, the hierarchical VAE surpasses the previous state-of-the-art mechanistic MT model [27] by a gain of over 2x in predictive power. Chapter 5 applies the same variational-inference strategy to spontaneous activity: a mixed-membership stochastic blockmodel decomposes simultaneous fMRI-BOLD and wide-field calcium fluorescence in mice into overlapping communities, and around half of the cortical regions belong to multiple communities.","pith_inferences":["Taken further than the dissertation states: if the inference objective is the cause of the 2x gain, then training the same model on more ecological optic flow—pursuit eye movements, multiple objects, occlusions—should preserve or increase the gain; if the gain collapses, the result is specific to the simplified stimulus family rather than to hierarchical inference per se.","A testable extension the author leaves implicit: latent units carrying object-motion information in the model should map onto neurons in area MST, the next cortical stage after MT, predicting that MST-like selectivity for combined optic flow and object motion does not require extra-retinal input.","On the network side, the overlap result implies that regions with high membership entropy are the most likely to switch community allegiance when the brain changes state, linking static overlap to dynamic circuit reconfiguration in a way that task or arousal manipulations could test.","Because Ca2+ and fMRI reveal similar principal gradients but different degree and entropy maps, the disparate centrality measures across modalities may index neurovascular rather than purely neural properties; a joint model treating BOLD and calcium as two noisy observations of one shared community latent would test that interpretation."],"forward_implications":["Unsupervised, visual-only learning can separate self-motion from object motion when the network is hierarchical and trained with an inference-based loss; no extra-retinal signals are required.","Hierarchical latent structure is the reason for the improved brain alignment: flat VAEs trained on the same objective and data align less well with MT neurons, so the hierarchy itself carries predictive power.","MT encoding models no longer need hand-designed motion filters to be competitive; the latent representations of a generative model trained on optic flow predict neural responses better than the previous mechanistic standard.","Around half of mouse cortical regions belong to more than one resting-state network, so disjoint network parcellations misdescribe a substantial fraction of cortex.","The overlapping organization is not purely a BOLD artifact: wide-field calcium imaging, a more neuron-specific signal, produces largely concordant overlapping communities, with modality-specific differences in degree and diversity metrics."],"supporting_citations":[{"why":"Supplies the mechanistic three-dimensional spatiotemporal receptive field model of MT neurons and the naturalistic stimulus baseline that the hierarchical VAE claims to surpass by over 2x.","marker":"[27]"},{"why":"Introduces the variational autoencoder and the ELBO objective that the dissertation adapts into a hierarchical, inference-based loss function.","marker":"[413]"},{"why":"Source of the conjecture that the visual cortex performs hierarchical Bayesian inference with feedback carrying prior expectations; motivates the hierarchical architecture.","marker":"[127]"},{"why":"Original formulation of perception as unconscious inference, which supplies the Helmholtzian interpretation of the generative objective.","marker":"[11]"},{"why":"Provides the simultaneous fMRI and wide-field Ca2+ imaging dataset in mice used for the overlapping-community analyses in Chapter 5.","marker":"[276]"},{"why":"Supplies the overlapping-community algorithm based on mixed-membership stochastic blockmodels and variational inference used to decompose the mouse cortical networks.","marker":"[182]"}],"fun_headline_variants":["Hierarchical VAE beats prior MT model by 2x","Unsupervised VAE splits self-motion from object motion like primate MT","Generative model doubles prediction of primate motion neurons","Helmholtz-inspired VAE mirrors monkey visual cortex motion coding"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the simplified simulated retina—one fixating observer, one fixed-size moving object, and prescribed depth statistics—captures the optic-flow structure that real macaque MT neurons encode, so that what the model learns transfers to biological neurons.","fun_headline_variants_meta":{"raw":{"variants":["Hierarchical VAE beats prior MT model by 2x","Unsupervised VAE splits self-motion from object motion like primate MT","Generative model doubles prediction of primate motion neurons","Helmholtz-inspired VAE mirrors monkey visual cortex motion coding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000861,"raw_usage":{"total_tokens":3797,"prompt_tokens":1067,"completion_tokens":2730,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":683,"completion_tokens_details":{"reasoning_tokens":2659}},"tokens_in":683,"tokens_out":2730,"duration_ms":19335,"temperature":1.0,"reasoning_tokens":2659,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:30:08.162059+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same model on ROFL variants with two independently moving objects, pursuit or saccadic eye movements, occlusions, or natural depth variation, and re-measure the MT alignment gain. If the >2x improvement over the mechanistic baseline disappears or falls below significance, the finding is specific to the simplified stimulus family rather than a general principle of hierarchical inference.","supporting_citations":[],"review_version":1}