{"id":"93da9c4b-0b18-4484-8f10-278b587b74ea","arxiv_id":"2501.11566","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 119 ANN-MEG studies shows rapid growth across decoding, BCI, clinical, modeling, and source-localization applications, with recurring reproducibility gaps.","lead":"This review maps 119 studies that use artificial neural networks on magnetoencephalography (MEG) brain recordings, sorting them into classification, brain modeling, and methodological tasks. It gives newcomers a practical map of the field and flags recurring problems like small datasets, missing baselines, and scarce interpretability tools.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Search query omits 'neural network', 'deep neural network', and other common terms, so the 119-study census and growth trend may undercount the field.","rationale":"The review provides a useful first structured map of ANN-MEG work, and its qualitative message—that the area is active and spans classification, modeling, and methodological applications—is likely to survive even a revised census. However, the strongest claim as stated in the abstract is quantitative: 'We identified 119 relevant studies, with 70 focused on Classification, 16 on Modeling, and 33 in the Other category.' The reader's weakest assumption was that the search result is representative; I agree, and the specific mechanism is the query's missing generic ANN terminology. The Methods section (§2.1) reports a query that cannot plausibly retrieve all ANN-MEG work: it omits 'neural network' without the qualifier 'artificial', omits 'deep neural network', and omits common architecture names. Some of the papers actually included in the review, such as a GPT foundation model [163] and a ROCKET-based classifier [159], would not be matched by the stated query if their titles/abstracts were the only evidence screened. This indicates either undocumented query variants or an undocumented broadening of the search; either way, the stated protocol is not reproducible and the 119-study census is not auditable. The DOI requirement adds a temporal bias because older preprints and conference papers often lack DOIs, so the growth curve in Fig. 2a may overstate a recent surge. These issues do not justify rejecting the review, since the qualitative conclusion is well supported by the assembled applications and the recommendations are reasonable. But they do justify the reader's CONDITIONAL verdict: before the review is treated as a definitive field reference, the search should be re-run with a transparent, reproducible protocol and the counts and distributions updated accordingly. The concrete test above would settle whether the concern actually changes the quantitative conclusions.","tokens_in":41237,"tokens_out":7906,"duration_ms":84233,"concrete_test":"Re-run the §2.1 search on PubMed and arXiv with the same MEG block and date/language filters, but execute the stated original query and a broadened query side-by-side. The broadened ANN block should be ('neural network*' OR 'deep neural network' OR 'artificial neural network*' OR 'multilayer perceptron' OR MLP OR autoencoder OR transformer OR LSTM OR RNN OR CNN) AND ('MEG' OR magnetoencephalogra*). Record the number of unique additional studies not in the 119, their assigned categories, and their publication years. If the broadened query adds more than 10% additional studies, or if including older no-DOI papers shifts the year-by-year growth curve materially, the census and trend claims in the abstract and conclusion are biased.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—119 studies, 70/16/33 split, and a 'rapidly growing interest' trend—rests on the literature search in §2.1. The query's ANN term block is ('deep learning' OR 'artificial neural network*' OR 'Convolutional neural net*' OR CNN OR 'Recurrent neural net*' OR RNN). It does not include the plain phrase 'neural network' (only 'artificial neural network*'), nor 'deep neural network', 'multilayer perceptron', 'MLP', 'autoencoder', 'transformer', 'LSTM', or 'RNN' as standalone terms. A MEG study whose title/abstract says 'we trained a deep neural network' or 'a transformer-based model' would therefore fail the stated title/abstract screening even if it is squarely an ANN-MEG study. This is not hypothetical: the included corpus itself contains works such as [163] 'Foundational GPT model for MEG' and [159] 'ROCKET-based models' that would not be caught by the stated query terms if encountered without other qualifying language. The PDF search described in §2.1 does not repair this, because the selection procedure is described as title-first and abstract-second, and the query defines the candidate pool. In addition, the DOI requirement systematically excludes older preprints and conference papers without DOIs, which would bias the year-by-year curve in Fig. 2a toward recent years and thus inflate the 'rapidly growing interest' conclusion. Because the abstract and conclusion make the exact count and growth pattern central, this search gap is the load-bearing weakness; the numeric inconsistencies (71 vs. 70 classification studies; 17 vs. 16 hyperparameter omissions) are secondary symptoms of the same auditability problem.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This review surveys the use of artificial neural networks (ANNs) in magnetoencephalography (MEG) research. The authors searched PubMed, Google Scholar, arXiv, and bioRxiv with a fixed query, screened titles, abstracts, and PDFs, and included 119 primary research articles with a DOI published before November 2024. They categorize the corpus into Classification (70 studies, further split into decoding, BCI, clinical, and event detection), Modeling (16 studies), and Other (33 studies, covering preprocessing, source localization, and methods). For each category, the review summarizes typical pipelines, data characteristics, network architectures, training and validation practices, and current limitations, and closes with recommendations on data augmentation, validation, baselines, interpretability, and reproducibility.","tokens_in":41531,"tokens_out":6686,"duration_ms":63242,"significance":"If the census is accurate, this is a valuable and timely reference: it organizes a rapidly growing literature into a clear taxonomy, provides condensed tables of architectures, datasets, and validation schemes, and identifies recurring methodological weaknesses (missing baselines, limited interpretability, sparse reporting of hyperparameters). The paper's main strengths are its breadth, the explicit categorization scheme, and the practical recommendations grounded in the corpus. The central quantitative claim—119 studies with a 70/16/33 split and a rapid growth trend—is, however, only as reliable as the literature search and the consistency of the reported counts, both of which have issues that need to be addressed.","major_comments":[{"comment":"The search query omits several common terms for neural-network methods, including \"neural network\" (without the qualifier \"artificial\"), \"deep neural network\", \"multilayer perceptron\", \"MLP\", \"autoencoder\", \"transformer\", \"LSTM\", \"GRU\", \"BERT\", and \"GPT\". Because the initial candidate pool is defined by this query and then screened by title and abstract, a MEG study whose abstract says \"we trained a deep neural network\" or \"a transformer-based model\" would be missed. The included corpus itself contains papers such as [163] (a GPT model), [159] (ROCKET-based), and [165] (CLIP-based) that rely on terminology outside the stated query terms, so the concern is concrete. The PDF search described in §2.1 does not repair this, since the query defines the pool before PDF screening. This directly affects the central 119-study count and the growth trajectory in Fig. 2a. I ask the authors to rerun the search with an expanded term set, report the exact search date and screening counts (including how many titles/abstracts were screened at each stage), and discuss how the DOI-only inclusion criterion affects coverage of older conference papers and preprints, which may bias the year-by-year curve toward recent years.","section":"§2.1, Fig. 1, Abstract"},{"comment":"The manuscript contains several internal numeric inconsistencies that affect the quantitative portrait. Specifically: §3.2.5 states \"Among the 71 studies in this category\" though the Classification category contains 70 studies per Tables 2–3 and §3.2.2; §4.3.4 says \"16 of the classification studies did not provide enough information\" about hyperparameters, while §3.2.5 says 17 studies do not mention training parameters; Table 6 lists reference [165] under both \"Methods\" and \"Source localization\", and the Source localization subcategory omits [141] that is listed in Table 5; and §3.3.2 refers to \"the 11 studies using RSA\" while §3.3.3 and the preceding text in §3.3.1 describe 10 RSA-based studies. These discrepancies are individually small but collectively undermine confidence in the reported statistics. The authors should reconcile all counts across the text, tables, and figure captions.","section":"§3.2.5, §4.3.4, Table 6, §3.3.2"},{"comment":"A few citation/count errors also appear in the discussion: §4.2 lists [119] as an example of a classification study, but [119] is a Modeling study; and Table 2 lists the Shu and Fyshe study [77] (a 2013 workshop paper) under publication year 2020. Such errors, while not central to the main thesis, should be corrected to make the review a reliable reference.","section":"§4.2, Table 2"}],"minor_comments":[{"comment":"The section heading contains a typo: \"Litterature research\" should be \"Literature research\".","section":"§2.1"},{"comment":"The sentence \"These techniques have proven to be shown to be useful\" is grammatically awkward and should be rewritten.","section":"§1.6"},{"comment":"The sentence about sampling frequencies reports both a median of 250 Hz in §3.1 and a median of 600 Hz in §3.2.3; please make these consistent and clarify what set each median is computed over.","section":"§3.2.3"},{"comment":"The text says \"Out of the eleven studies investigating the visual cortex (including visual word recognition)\", but the earlier enumeration gives nine visual-cortex studies plus one visual-word-recognition study, i.e., ten. Please correct the number.","section":"§3.3.1"},{"comment":"In the Source localization row, reference [165] appears to be misassigned (it is a multimodal alignment study, not a source-localization study), and [141] is missing from that row.","section":"Table 6"},{"comment":"The sentence \"Roughly half of the 'Methods' studies (9 out of 16) included visualization techniques\" is internally consistent, but it would be clearer to say \"9 of 16\" rather than \"roughly half\".","section":"§4.3.7"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a broad topical review rather than a PRISMA-style systematic review, which is acceptable for this venue, but the authors should be explicit about this scope. The search-strategy gap is the main technical concern; it is fixable by expanding the query and reporting screening statistics. The internal count inconsistencies are numerous enough that the authors should do a careful pass over all numbers before resubmission. The review does not appear to be circular: the census is built from external studies, and the authors' own works are cited only as recommendations. The paper fits the journal's scope as a review of an emerging application area in neuroimaging."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on this review. It's the first systematic survey of artificial neural networks applied to MEG, organized around a sensible three-way taxonomy (classification, modeling, other) with subcategories, and it tables a lot of useful extracted metadata (subjects, trials, architectures, validation, code availability) for 119 studies. That alone makes it a useful entry point for someone entering the field. The practical recommendations—report class imbalance, use baselines, detail hyperparameters, publish code—are sensible and well grounded in what the corpus actually shows.\n\nThe soft spots are real. The search query in §2.1 is narrower than the field: it includes 'artificial neural network*' but not 'neural network', 'deep neural network', 'MLP', 'transformer', 'LSTM', or 'autoencoder'. Combined with a DOI-only inclusion rule, this will miss many relevant preprints and conference papers, and it biases the year-by-year curve toward recent, DOI-bearing work. The stress-test note lands: the census is likely an undercount and the growth trend in Fig. 2a is only indicative. The review also contains an internal inconsistency: §3.2.5 refers to '71 studies' where the category has 70 per the tables, and the count of classification studies missing training parameters is 17 in §3.2.5 but 16 in §4.3.4. These are minor numeric slips, but in a review built on numbers they need fixing.\n\nThe paper does not report PRISMA-style screening counts or a precise search date, so the audit trail is weaker than it should be for a quantitative census. If the authors can supply the screening flow, broaden the query, and reframe the 119/70/16/33 figures as approximate, the review will be much more durable.\n\nWho's it for? Anyone wanting an overview of ANN-MEG applications, especially newcomers; the tables are a quick-reference digest. It deserves serious peer review—the gap is real and the organization is genuinely helpful. I'd recommend revision rather than desk rejection, with the search protocol and internal counts being the priority.","headline":"A useful first map of ANN-MEG, but the census undercounts from a narrow search query; fix the audit trail before treating the numbers as definitive.","tokens_in":42085,"tokens_out":2688,"would_cite":false,"duration_ms":28677,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review identifies 119 studies using artificial neural networks on MEG data and organizes them into classification, modeling, and other methodological applications, arguing the field is growing rapidly.","keywords":["magnetoencephalography","artificial neural networks","deep learning","brain decoding","representational similarity analysis","source localization","artifact removal","systematic review"],"falsifier":"Re-run the literature search from the methods section without requiring a digital object identifier and without the English-only restriction, recording screening counts at each stage; then recompute the total number of studies and the category shares in figures 2a and 2c. If the total grows substantially or the classification share drops below half, the review's portrait of the field depends on its search choices rather than on the underlying research activity.","tokens_in":41031,"feed_emoji":"🧠","tokens_out":6797,"duration_ms":68525,"temperature":0.7,"pith_summary":"This review tries to establish that artificial neural networks have become a recognizable, rapidly growing subfield of MEG research, where MEG is magnetoencephalography, a technique that records magnetic fields from brain activity with millisecond timing. It reports 119 relevant studies, with classification dominant at 70 studies, modeling smallest at 16, and 33 studies in an 'Other' category covering preprocessing, source localization, and methods. If the review's picture is right, ANN-MEG work is not a scattering of one-off demonstrations but a coherent field with distinct pipelines, evaluation practices, and open challenges. The authors also claim the trajectory resembles the earlier expansion of deep learning in EEG analysis, which would predict continued growth, more foundation models, and multimodal decoding.","feed_headline":"MEG meets deep learning: 119 studies, three application families","feed_subtitle":"A systematic review sorts the emerging field into classification, brain modeling, and methodological tooling.","key_machinery":"The organizing device is the three-category taxonomy (Classification, Modeling, Other) with subcategories, applied to a corpus of 119 studies. This taxonomy does the argument's work: it turns a heterogeneous collection of papers into a map of pipelines, enabling the review's quantitative claims about growth, architecture choices, validation practices, baseline reporting, and gaps in interpretability.","core_discovery":"The central claim is that ANNs are being applied to MEG data in three distinct modes: classification, where MEG trials are inputs and outputs are labels for decoding, brain-computer interfaces, clinical diagnosis, or event detection; modeling, where ANN activations are compared with MEG responses to the same stimuli, mostly via representational similarity analysis or neural predictivity; and other methodological uses, including preprocessing, artifact removal, and source localization. The review further claims that classification pipelines are diverse with no standard protocol, that modeling studies are gaining momentum, and that the main bottlenecks are data scarcity, limited interpretability, and reproducibility.","pith_inferences":["The paper does not claim this, but the DOI requirement likely biases the corpus toward published work; including DOI-less preprints could reveal a steeper recent growth curve and a larger share of 'Other' methodological studies.","If the taxonomy were applied to the EEG-ANN literature, tests of whether MEG's temporal resolution yields systematically different modeling insights could be made explicit.","Given that 27 of 70 classification studies lack baseline comparisons, a re-analysis could quantify how often ANN accuracies actually beat classical classifiers; the review leaves this as future work."],"forward_implications":["If the growth curve in figure 2a continues as the review expects, more MEG studies will adopt ANNs, and the field should see more foundation models for MEG, following the trajectory observed in EEG.","Classification research will likely remain CNN-dominated, but without shared benchmarks or standardized preprocessing, cross-study comparisons will stay difficult.","Modeling studies using RSA and neural predictivity may become a standard way to test whether ANN representations track the millisecond-scale dynamics of the human brain.","ANN-based source localization is promising but bounded by forward-model accuracy; gains over classical inverse methods such as MNE, beamforming, and sLORETA will depend on realistic simulations.","Interpretability tools, used in only 16 of 70 classification studies, will need to become routine for decoding claims to be explainable."],"supporting_citations":[{"why":"Establishes MEG's strengths and limitations, motivating why ANNs are suited to its high temporal resolution.","marker":"[12]"},{"why":"Provides the EEG deep-learning review whose growth trajectory the authors use to predict a similar expansion for MEG.","marker":"[5]"},{"why":"Defines representational similarity analysis, the main technique the modeling studies use to compare ANN and brain representations.","marker":"[185]"},{"why":"Exemplifies ANN-based preprocessing via time-contrastive learning as an alternative to ICA.","marker":"[142]"},{"why":"Represents ANN-based source localization benchmarked against classical inverse methods in the Other category.","marker":"[138]"},{"why":"Supports the review's warning that small-sample accuracies can beat chance by chance, motivating permutation baselines.","marker":"[203]"},{"why":"Represents the foundation-model direction the review identifies as a key future prospect.","marker":"[163]"},{"why":"Supplies the MEG-BIDS standard the review recommends for reproducibility and cross-study comparison.","marker":"[205]"}],"fun_headline_variants":["ANNs in MEG: three roles reviewed","MEG and neural networks: a three-way review","Artificial networks for MEG: sorting the field","Neural networks for MEG: from decoding to sources","MEG meets deep learning: three application families"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 119 studies found by the English-only search, which required a digital object identifier, across the four databases described in the methods, accurately represent the whole body of ANN-MEG work; if relevant preprints without DOIs, non-English papers, or studies using different terminology were missed, the reported counts, category shares, and growth trend could shift.","fun_headline_variants_meta":{"raw":{"variants":["ANNs in MEG: three roles reviewed","MEG and neural networks: a three-way review","Artificial networks for MEG: sorting the field","Neural networks for MEG: from decoding to sources","MEG meets deep learning: three application families"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000247,"raw_usage":{"total_tokens":1543,"prompt_tokens":948,"completion_tokens":595,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":521}},"tokens_in":564,"tokens_out":595,"duration_ms":6148,"temperature":1.0,"reasoning_tokens":521,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:06:22.523117+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the literature search from the methods section without requiring a digital object identifier and without the English-only restriction, recording screening counts at each stage; then recompute the total number of studies and the category shares in figures 2a and 2c. If the total grows substantially or the classification share drops below half, the review's portrait of the field depends on its search choices rather than on the underlying research activity.","supporting_citations":[{"cited_title":"Representational similarity analysis- connecting the branches of systems neuroscience","cited_arxiv_id":null,"evidence_quote":"Defines representational similarity analysis, the main technique the modeling studies use to compare ANN and brain representations."},{"cited_title":"Exceeding chance level by chance: The caveat of theoretical chance levels in brain signal classification and statistical assessment of decoding accuracy","cited_arxiv_id":null,"evidence_quote":"Supports the review's warning that small-sample accuracies can beat chance by chance, motivating permutation baselines."},{"cited_title":"Meg-bids, the brain imaging data structure extended to magnetoencephalography","cited_arxiv_id":null,"evidence_quote":"Supplies the MEG-BIDS standard the review recommends for reproducibility and cross-study comparison."}],"review_version":1}