{"id":"ee12d6aa-c7ff-475a-9b05-6e95657ad5d5","arxiv_id":"2607.26666","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 239 studies finds AI for deep brain stimulation is mostly at proof-of-concept readiness (TRL 2-4), with external validation in only 5.9%.","lead":"Across 239 studies, AI for deep brain stimulation in movement disorders is still mostly at early proof-of-concept stages: 86% of systems sit at readiness levels 2-4 and almost none have external validation. The review maps where the field stands and argues the bottleneck is validation, not algorithm quality, giving funders and researchers a concrete agenda.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Per-study TRL ratings may understate system-level maturity because the same AI system is often described across multiple publications; the claim that 'most systems' sit at TRL 2–4 is not directly supported.","rationale":"The reader's weakest assumption concerns missing non-AI-indexed evidence, which is explicitly acknowledged and conservative. My concern is distinct: the TRL assessment unit is the study, not the system, and this mismatch is not acknowledged in the limitations. It directly affects the central claim's wording ('most systems remain at early-to-intermediate translational stages') because multi-study systems can have cumulative validation that no single paper reports. This is a methodological issue with a concrete remediation (system-level clustering and re-scoring). The paper is otherwise well-conducted, data are transparent, and the interpretation is reasonable for study-level evidence. Therefore the verdict should be CONDITIONAL: accept the review as a study-level characterization, but require the system-level reanalysis or a qualified rephrasing before the abstract's system-level conclusion can be taken as established.","tokens_in":22955,"tokens_out":5344,"duration_ms":60372,"concrete_test":"Cluster the 239 included studies into distinct AI systems by matching shared model names, public code repositories, explicit system references, or overlapping author groups and dataset descriptions. For each cluster, pool all evidence from its constituent studies (including any referenced clinical-trial reports) and re-apply the same TRL rubric at the system level. If the system-level distribution remains at least 80% at TRL 2–4, the conclusion holds; if a nontrivial fraction (e.g., >15%) of systems reach TRL 5–6, the per-study analysis understates maturity and the abstract/discussion must be revised to say 'most publications' rather than 'most systems.'","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The review assigns TRL to each included study individually (Methods, 'Technology readiness level assessment'), but the abstract and discussion interpret the distribution as characterizing 'systems' ('most systems remain at early-to-intermediate translational stages'; 'no reviewed system reached routine clinical deployment'). TRL is a property of a technology system, not of a publication. If a single DBS-AI system is developed incrementally, one paper may report a retrospective proof-of-concept (TRL 4), a second a clinician-facing prototype (TRL 5), and a third a prospective evaluation (TRL 6). Rating each paper independently assigns TRL 4 to all three, silencing cumulative evidence across the system's lifecycle. This is not merely hypothetical: the corpus includes multiple publications from the same groups on related decoders, targeting pipelines, and programming algorithms, yet no clustering or aggregation by system is described. The acknowledged limitation about non-AI-indexed clinical trials is a separate and arguably weaker threat because it depends on indexing gaps; the study-versus-system unit mismatch affects every system with more than one publication, even when all AI papers are captured. Consequently, the 86.2% TRL 2–4 figure likely overstates the proportion of immature systems, and the comparative claim that validation—not algorithmic capacity—is the dominant constraint is not established at the level of systems. The paper should either rephrase its central claim to 'most publications/studies' or provide a system-level TRL analysis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is a PRISMA-compliant systematic review of 239 peer-reviewed studies (2000–2025) applying AI to deep brain stimulation (DBS) for movement disorders. It maps AI methods, data modalities, validation practices, and translational maturity using a machine-learning-adapted Technology Readiness Level (TRL) rubric. The authors report a field concentrated on Parkinson's disease and subthalamic nucleus targeting, dominated by internal cross-validation, with rare external (5.86%) and temporal (2.51%) validation. A study-level TRL distribution places 86.2% of reports at TRL 2–4 and none at TRL ≥7. The central conclusion is that most AI-DBS systems remain at early-to-intermediate translational stages and that the bottleneck is validation maturity rather than algorithmic capacity.","tokens_in":23283,"tokens_out":7397,"duration_ms":66090,"significance":"If the findings hold, this review provides a valuable, reproducible evidence map and a structured translational assessment of AI in DBS. Strengths include prospective registration (PROSPERO CRD42023488688), independent dual screening/extraction with adjudication, a publicly available full dataset and code on Figshare, a clearly specified ML-adapted TRL rubric anchored in the external MLTRL framework, and explicit acknowledgement of key limitations. The descriptive distributions (external validation 5.86%, temporal validation 2.51%, TRL 2–4 at 86.2% of studies) are useful and falsifiable. However, the central 'systems' interpretation is partly interpretive and is vulnerable to a unit-of-analysis mismatch between study-level TRL ratings and system-level claims, as well as protocol-expansion transparency concerns.","major_comments":[{"comment":"The TRL rubric is applied at the level of individual studies ('TRL assignment [sic] were based solely on information explicitly reported in each study'), but the abstract and discussion repeatedly generalize to 'systems' ('most systems remain at early-to-intermediate translational stages'; 'no reviewed system reached routine clinical deployment'). TRL is a property of a technology system, not a publication. If the same AI system is described across multiple publications—for example, a retrospective proof-of-concept (TRL 4), then a clinician-facing prototype (TRL 5), then a prospective evaluation (TRL 6)—rating each study independently assigns TRL 4 to all three and undercounts cumulative system maturity. No clustering or aggregation by system is described. This is load-bearing because the 86.2% TRL 2–4 figure and the 'most systems' claim are based on per-study ratings. The acknowledged l","section":"Abstract; Results, 'TRL distribution and translational gaps'; Methods, 'Technology readiness level assessment'"},{"comment":"The explanatory claim that the field is 'constrained more by limited validation than by algorithmic inadequacy' is an interpretation, not a direct result of the reported analyses. The review documents low rates of external/temporal/prospective validation and notes high internal performance, but it does not systematically compare algorithmic adequacy with validation limitations; there is no analysis of whether low-TRL studies have systematically weaker models, no benchmark of achieved performance against known ceilings, and no evidence that algorithmic capacity is uniform across the corpus. As phrased, this overstates an explanatory conclusion. Please present this explicitly as an interpretation/hypothesis and, if possible, support it with a sensitivity analysis (e.g., correlation of TRL with reported performance, or comparison of performance declines under external versus internal valida","section":"Abstract; Discussion, 'With median TRL 3–4...'"},{"comment":"The PROSPERO registration is a strength, but the Methods state that the analytical framework was expanded after initial data extraction to incorporate governance, deployment readiness, and generalisation variables. These added variables are precisely the features that feed the TRL and validation conclusions. The assertion that this expansion 'was driven by the exploratory nature of the evidence base rather than by post-hoc outcome considerations' cannot be verified from the manuscript. Please report the protocol amendments explicitly, with dates, and label the corresponding analyses as exploratory in the Results and Discussion. This is a load-bearing transparency issue for a systematic review and should be weighed in the interpretation of the central claim.","section":"Methods, 'Data Extraction and Synthesis' and 'Technology readiness level assessment'"}],"minor_comments":[{"comment":"Typo: 'TRL assignment were based' should be 'TRL assignments were based'.","section":"Methods, 'Technology readiness level assessment'"},{"comment":"Reference 75 is cited as an arXiv preprint dated 2026. If this is a non-peer-reviewed preprint, it should be labeled as such, and the systematic review should not rely on it as supporting evidence for a clinical-utility claim.","section":"References, ref. 75"},{"comment":"The phrase 'Dedicated segmentation metrics ... appeared in only 12.9% of targeting-related metric mentions' is ambiguous: clarify whether this is 12.9% of targeting studies or 12.9% of metric mentions.","section":"Results, 'Validation landscape'"},{"comment":"The 'External validation rate' column entries such as '2/62' would be clearer as percentages, especially because the text reports percentages (e.g., 14.8%, 2.7%). Ensure all workflow stages are represented consistently.","section":"Table 3"},{"comment":"The legend states that papers using multiple input feature categories appear in multiple columns, but it is not clear how the TRL dot colors are handled for multi-feature papers; please clarify the counting and coloring rules.","section":"Figure 4B"}],"recommendation":"major_revision","confidential_remarks":"The unit-of-analysis issue in the central claim is real and should be addressed before acceptance. The protocol expansion is disclosed, but the manuscript should clearly label the post-hoc variables. I have also flagged reference 75 (a self-citation to a 2026 arXiv preprint) as a possible data-integrity matter; the editor may wish to verify its status."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This review does a genuinely useful thing: it collates 239 studies across the full DBS workflow, applies a ML-adapted TRL rubric, and publicly releases the extracted data and code. The validation numbers alone justify the paper—5.9% external validation and 2.5% temporal validation in a fast-growing literature is a stark finding. The conclusion that validation, not algorithm design, is the main constraint is an interpretation, but a reasonable one given how little independent testing exists. The PRISMA reporting, dual extraction, and explicit limitations sections are credit-worthy.\n\nThe soft spots, in proportion:\n\nFirst, the unit-of-analysis mismatch. TRLs are assigned per study, but the abstract and discussion say \"most systems\" are at TRL 2–4. If the same system appears across multiple publications, aggregating by study understates system-level maturity. The 86.2% figure likely overstates the share of immature systems. This is not fatal—the paper could simply say \"most studies\" or do a system-level grouping—but it is a real gap between the data and the headline.\n\nSecond, there is a citation problem. Reference 75 is a 2026 arXiv preprint by the senior author, cited as if it were a peer-reviewed or included source. The review explicitly covers 2000–2025 and excludes non-peer-reviewed work. The citation is non-load-bearing, but it should be removed or justified.\n\nThird, the \"constrained more by limited validation than by algorithmic inadequacy\" phrasing is more causal than the descriptive data support. A systematic review can show validation is rare; it cannot directly show that is the binding constraint. The paper would be more precise to say \"validation is the clearest bottleneck\" rather than implying it outranks all other factors.\n\nMinor: the TRL adaptation is inherently subjective, and the conservative lower-level assignment when reporting is incomplete could mask maturity. The authors acknowledge this, though not the system-level issue.\n\nOverall, this paper deserves a serious referee. The empirical synthesis is transparent and checkable, and it gives the field a quantified map of where things stand. A referee should push for the system-level reframing and clean up the self-citation, but the core contribution stands.","headline":"Useful, transparent systematic map of AI-DBS with a validation bottleneck worth taking seriously—but the headline 'system' claim overstates the data, and there's a self-citation that slipped past the protocol.","tokens_in":23781,"tokens_out":3146,"would_cite":true,"duration_ms":34986,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Across 239 studies, AI for deep brain stimulation clusters at technology-readiness levels 2–4, with no system reaching routine clinical deployment—evidence that the bottleneck is validation, not algorithmic capability.","keywords":["deep brain stimulation","artificial intelligence","technology readiness level","movement disorders","external validation","clinical translation","Parkinson's disease","model generalization"],"falsifier":"A registry-based audit that identifies more than a handful of AI-assisted DBS systems with regulatory clearance or routine institutional use (TRL 7–9) whose studies were missed because they do not index AI terms would undercut the claim that no system reaches TRL ≥7. Equally, a single large prospective multicentre external-validation study in adaptive DBS for motor-state control in Parkinson's disease would directly falsify the claim that this subfield has not exceeded TRL 5 in any signal modality.","tokens_in":22910,"feed_emoji":"🧠","tokens_out":4299,"duration_ms":40620,"temperature":0.7,"pith_summary":"This review asks whether artificial intelligence for deep brain stimulation in movement disorders is converging on clinical deployment or accumulating proof-of-concept work. Analyzing 239 peer-reviewed studies from 2000 to 2025, it finds that 86.2% sit at technology-readiness levels 2–4, only 13.8% reach levels 5–6, and none reach level 7 or above. External validation appears in 5.86% of studies, temporal validation in 2.51%, and prospective evaluation in 6.27%. The authors conclude that the field is constrained less by algorithmic inadequacy than by scarce validation under real-world variability, and they map which workflow stages—targeting, programming, adaptive stimulation—are closest to near-term clinical value.","feed_headline":"86% of AI brain-stimulation studies stay at early readiness","feed_subtitle":"Review of 239 studies finds external validation in only 5.9% and no system in routine clinical use.","key_machinery":"The central instrument is the machine-learning-adapted Technology Readiness Level (TRL) scale, a nine-level rubric that scores how close a system is to deployment based on evidence reported in the study—from conceptual work (TRL 1–2) through proof-of-concept on retrospective data (TRL 4), clinician-facing prototypes (TRL 5), prospective evaluation in workflow context (TRL 6), to integrated pilot deployment, routine use, and multicentre adoption or regulatory clearance (TRL 7–9). It is complemented by a strict three-way distinction among external validation (testing on an independent cohort from a different institution), temporal validation (chronological data splitting or longitudinal follow","core_discovery":"The central discovery is a quantitative picture of translational immaturity: a machine-learning-adapted Technology Readiness Level assessment of all 239 included studies shows that most AI-DBS systems are proof-of-concept (TRL 4 is the modal level, at 65.7%), none have reached integrated pilot deployment, routine use, or regulatory clearance (TRL ≥7), and the small minority reaching TRL 5–6 concentrate in objective assessment and adaptive DBS. The paper argues that this pattern persists not because the algorithms are too weak but because the field has not generated the external, temporal, and prospective evidence needed to show generalizability across centres, devices, and time.","pith_inferences":["A testable extension would be to run the same TRL rubric against clinical-trial registries and regulatory databases that do not index AI terms; if many systems are already in routine or trial use without AI-related keywords, the reported immaturity is partly an artifact of search scope.","If validation is the bottleneck, then funding or reporting mandates that require external or temporal validation for publication could shift the field faster than any new algorithmic technique; one could compare the TRL distribution of studies published after such mandates with earlier work.","The paper's logic suggests a simple leading indicator for the field: the ratio of externally validated studies to total studies per workflow stage, tracked over time—a metric that future reviews could adopt as a standard reporting item.","The finding that tasks with clear anatomical ground truth transfer across institutions while behavioural-signal tasks degrade suggests a broader principle: AI trustworthiness in surgery may scale with the objectivity of the ground truth, so the first deployable systems will likely be those whose labels are anatomical or electrical rather than clinical-rating-based."],"forward_implications":["If the TRL distribution is accurate, no current AI-DBS system is deployable as a routine clinical tool, and the near-term priority should shift from building new architectures to generating reproducible, externally validated evidence.","Externally validated success is concentrated in tasks with anatomical ground truth (contact selection, STN localization), while tasks generalizing biological or behavioural signals degrade across cohorts—implying near-term clinical value is most likely in targeting and programming.","Adaptive DBS, despite the largest study volume, has the lowest external validation rate (2.7%) and no motor-state control work above TRL 5, so closed-loop systems still need proof of stability across time and centres before routine use.","With only 5.02% of studies referencing regulatory frameworks and 2.51% addressing model lifecycle management, compliance and governance infrastructure will have to be built mostly from scratch as systems approach deployment.","Because 26.4% of studies operate in small-sample, high-dimensional settings and accuracy is the dominant reported metric, reported performances are likely optimistic upper bounds; calibration and decision-curve analysis should become minimum requirements for deployment claims."],"fun_headline_variants":["AI DBS: most systems remain proof-of-concept, not clinic-ready","239 studies show AI brain implants still far from clinical use","AI deep brain stimulation: validation gap blocks deployment","DBS AI: external validation in <6% studies, none in routine use","AI for DBS: modal readiness at TRL 4, none beyond pilot"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The finding rests on treating absence of reported validation evidence as absence of validation: TRL assignments were made only from what each study explicitly reported, with the lower level conservatively assigned when reporting was insufficient, so the 86% figure could understate true readiness if many teams have validated systems in trials or registries that do not use AI keywords.","fun_headline_variants_meta":{"raw":{"variants":["AI DBS: most systems remain proof-of-concept, not clinic-ready","239 studies show AI brain implants still far from clinical use","AI deep brain stimulation: validation gap blocks deployment","DBS AI: external validation in <6% studies, none in routine use","AI for DBS: modal readiness at TRL 4, none beyond pilot"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00076,"raw_usage":{"total_tokens":3195,"prompt_tokens":709,"completion_tokens":2486,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":453,"completion_tokens_details":{"reasoning_tokens":2394}},"tokens_in":453,"tokens_out":2486,"duration_ms":15336,"temperature":1.0,"reasoning_tokens":2394,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T11:10:15.115056+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A registry-based audit that identifies more than a handful of AI-assisted DBS systems with regulatory clearance or routine institutional use (TRL 7–9) whose studies were missed because they do not index AI terms would undercut the claim that no system reaches TRL ≥7. Equally, a single large prospective multicentre external-validation study in adaptive DBS for motor-state control in Parkinson's disease would directly falsify the claim that this subfield has not exceeded TRL 5 in any signal modality.","supporting_citations":[],"review_version":1}