{"id":"e7e19086-640d-45f9-bf18-2b54bf527b17","arxiv_id":"2607.24269","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A Bayesian Thompson-sampling electrode-selection method that reallocates channels over time captured 17.2 percentage points more oracle-available spikes than static selection in 34-hour neuronal culture recordings.","lead":"This paper introduces an algorithm that repeatedly re-picks which electrodes on a dense brain-cell recording chip should be used, so the fixed number of recording channels follows where the activity has moved. In nine 34-hour recordings, the adaptive method captured more of the available spikes than keeping the initial electrode selection fixed.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported 17.2 pp advantage likely reflects hyperparameter tuning on the same offline recordings used for the headline comparison; a held-out validation is required.","rationale":"The reader's weakest_assumption focuses on the downscaled offline replay (100 of 529 electrodes) not representing the full 1,024-of-26,400 problem. That is a real external-validity concern, and the paper itself lists full-array online evaluation as future work, so the central claim is honestly scoped to the offline setting. The hyperparameter-selection issue is more directly load-bearing for the internal validity of the reported statistic: the 17.2 pp advantage and p = 0.004 are presented as if from a fixed parameter configuration, but §3.3 reveals the parameters were chosen as 'best performing' in the same offline replay. Without a separate validation split, the p-value is conditional on a tuning procedure and is optimistically biased. A grid search over four parameters on nine recordings can easily produce a minimum-p-value outcome even when the true effect is zero, especially because the sign-flip test has only 2^9 = 512 possible outcomes. This is a correctness risk, not merely a generality caveat, and it can be resolved by a concrete nested-validation or sensitivity analysis. I therefore make this the primary load-bearing concern and keep the verdict CONDITIONAL pending that evidence.","tokens_in":15389,"tokens_out":6579,"duration_ms":58961,"concrete_test":"Run a nested validation protocol on the nine offline recordings: split into a tuning set (e.g., four recordings) and a held-out test set (remaining five), tune ρ, τ, α0, β0 on the tuning set, then evaluate the final attainment difference between the proposed method and static selection on the held-out recordings only. Report the held-out mean difference and a sign-flip p-value computed solely on test recordings. Additionally, perform a sensitivity sweep with ρ ∈ {0.5, 0.7, 0.9, 1.0} and τ ∈ {0.3, 1.0, 3.0} (other parameters fixed), and check whether the proposed method exceeds static selection in every configuration. If the held-out difference remains positive in the same direction for all nine recordings and across the sweep, the reported 17.2 pp is not an artifact of tuning.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central quantitative claim is that the proposed discounted Poisson–Gamma Thompson sampling exceeds static selection by 17.2 percentage points at the final 34 h time point (two-sided exact paired sign-flip test, p = 0.004; figure 3(D)). However, §3.3 states that the online run used 'the Bayesian hyperparameters that performed best in the offline replay,' and Table 1 fixes these same values (ρ = 0.7, τ = 1.0, α0 = 5.0, β0 = 1.0) for the offline evaluation that produces the headline number. If these parameters were selected by grid search over the same nine recordings used for the p-value, the comparison is not an out-of-sample test: the tuning process can exploit idiosyncrasies of these recordings, and the exact sign-flip p = 0.004 is the minimum attainable for n = 9, meaning the method happened to win in all nine recordings—exactly what a tuned policy would tend to produce even with no true advantage. The paper does not disclose a validation split or a parameter-selection protocol, and the cited supplementary sensitivity analyses (supplementary figures) are not shown in the main text; they would quantify robustness but cannot undo selection bias. This concern is load-bearing because the 17.2 pp gain and its p-value are the paper's primary evidence for adaptive selection; if the advantage disappears under a proper validation protocol, the central claim is unsupported regardless of the method's conceptual appeal.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a discounted Poisson–Gamma Thompson sampling algorithm for adaptive electrode selection in high-density microelectrode array (HD-MEA) recordings under a fixed channel budget. The method models per-electrode spike counts with a Poisson–Gamma conjugate model, discounts past observations to handle non-stationarity, and uses Thompson sampling with a temperature parameter to balance exploration and exploitation. The authors evaluate the method by offline replay of nine 34-hour dense recordings, in which 100 electrodes are selected from 529 candidates in a local 23×23 region, and compare it with static, random, and combinatorial ε-greedy baselines using an oracle-relative attainment metric. They report that the proposed method exceeds static selection by 17.2 percentage points at the final time point (two-sided exact paired sign-flip test, p=0.004, n=9). They also present a single representative online recording using 1,024 routed electrodes, in which the method captured the first synchronized burst, and they derive an algebraic relation (Eqs. 16–17) linking the fraction of captured spikes to the centroid bias of the estimated activity center.","tokens_in":15664,"tokens_out":6159,"duration_ms":47433,"significance":"If the central quantitative claim holds, the paper provides a practical and principled solution to a genuine bottleneck in switch-matrix HD-MEA recordings: the readout channel limit. The evaluation has several strengths: it uses an external oracle computed from dense reference data, so the main comparison is not circular by construction; it aggregates nine independent recordings; stochastic policies are averaged over 50 seeds; and the paired comparison uses an exact sign-flip randomization test. The centroid-bias identity (Eqs. 16–17) is algebraically correct and usefully connects spike-capture fraction to a downstream spatial summary. The paper also makes core analysis code publicly available. However, as detailed in the major comments, the out-of-sample validity of the headline improvement is compromised by the apparent selection of hyperparameters on the same offline recordings used for the comparison, and the downscaled evaluation region may not represent the full-array problem. These issues matter because the 17.2 pp advantage and its p-value are the primary evidence for the method's practical value.","major_comments":[{"comment":"Section 3.3 states that the online recording used 'the Bayesian hyperparameters that performed best in the offline replay,' and Table 1 fixes the same hyperparameter values (ρ=0.7, τ=1.0, α0=5.0, β0=1.0) for the offline replay that produces the headline 17.2 pp improvement. If these values were selected by searching over the same nine recordings used for the sign-flip test, the offline evaluation is not out-of-sample: the tuning process can exploit recording-specific idiosyncrasies, and the p=0.004 result (the minimum attainable for n=9, corresponding to the proposed method winning in all nine recordings) is exactly the pattern a tuned policy would tend to produce even without a true advantage. The manuscript does not disclose a validation split or a hyperparameter-selection protocol, and the supplementary sensitivity analyses are not shown in the main text. This concern is load-bearing because the 17.2 pp gain and its p-value are the paper's primary evidence for adaptive selection; a held-out validation (e.g., tuning on a subset of recordings and testing on the rest, or a documented a priori choice of hyperparameters) is needed to support the central claim.","section":"§3.3 and Table 1"},{"comment":"The offline evaluation is conducted exclusively in a downscaled setting: 100 electrodes are selected from 529 candidates within a local 23×23 region, whereas the real hardware problem is to select 1,024 from 26,400 electrodes. The paper does not provide evidence that the non-stationarity and spatial correlation structure of this local region are representative of the full array. The single online demonstration (§3.3, n=1) shows feasibility but cannot quantitatively validate the 17.2 pp advantage in the full-array setting. The Discussion (Section 4.3) acknowledges this limitation, but the abstract and conclusion present the 17.2 pp result without this caveat. I recommend either providing supporting analyses (e.g., comparing multiple local regions or coarser full-array maps) or explicitly qualifying the headline claim as a demonstration in a downscaled proxy.","section":"§2.3 and Figure 1(F)"}],"minor_comments":[{"comment":"There are multiple unresolved placeholder references to 'supplementary table??', 'supplementary figure??', and 'supplementary note??' throughout the text (Sections 2.2, 2.3, 3.1, 3.2, 3.4); these need to be filled in before publication.","section":"Throughout"},{"comment":"The statistical test is reported only at the final time point (34 h). Since the method's advantage develops over time, consider also reporting a test at an intermediate time point or a summary of the full temporal curve, to avoid relying on a single (potentially most favorable) time point.","section":"§3.2 and Figure 3(D)"},{"comment":"The statement that turnover is '0% when the same top-k electrodes are retained' should clarify that the measure uses the top-k set, not the ordering within that set.","section":"§2.3"},{"comment":"The code availability statement says the repository 'will be updated with the final analysis scripts upon publication'; if the analysis code is not yet fully available, this should be stated clearly in the main text.","section":"Data Availability"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and the proposed method has potential practical value. The main issue is the lack of a held-out validation for hyperparameter selection, which directly affects the credibility of the headline 17.2 pp improvement. The authors should be asked to either provide a proper validation split or clearly state if the hyperparameters were fixed a priori. The downscaled evaluation limitation is acknowledged, but the abstract and conclusion currently overstate the result."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take: this is a solid, practical methods paper. The central claim—adaptive electrode selection beats scan-and-commit for long-term HD-MEA recording—is probably true, but the reported 17.2 percentage-point advantage is not yet an out-of-sample result, because the hyperparameters were chosen on the same offline recordings used for the headline comparison.\n\nWhat is actually new: the authors frame long-term spontaneous in vitro recording on a switch-matrix HD-MEA as a non-stationary subset-selection problem, and adapt discounted Poisson-Gamma Thompson sampling to it. The application is new; the evaluation framework is the real contribution. They use densely recorded 529-electrode regions as ground truth, replay partial observation through each policy, and compare against an oracle that sees the full window. That is a clean way to measure a routing policy, and the 47.8% top-100 turnover at 34 h convincingly shows the problem is real. The centroid-bias identity (Eqs. 16-17) is correct and gives a nice reason to care about spike-capture fraction.\n\nThe offline evaluation is mostly careful: nine recordings, 50 seeds per stochastic policy, exact sign-flip test, oracle computed from full reference. I believe the comparison as run. The soft spot is the one the stress-test flags, and it is load-bearing. Section 3.3 says the online run used \"the Bayesian hyperparameters that performed best in the offline replay,\" and Table 1 fixes the same values for the offline evaluation that produces the headline number. No validation split or selection protocol is disclosed. With nine recordings, p = 0.004 is the minimum attainable for the sign-flip test, so the method won in all nine—exactly the pattern a tuned policy would show even without a real advantage. The supplementary sensitivity analyses are cited but not in the main text, and they cannot undo selection bias. This needs to be fixed with a held-out split, nested tuning, or honest reporting of the parameter grid and selection rule.\n\nTwo smaller caveats: the offline replay uses 100 of 529 electrodes from a local region, so transfer to 1,024 of 26,400 across the full array is not directly demonstrated; the single online run (n = 1) is illustrative, not validation. Data are not public yet, though core code exists. None of this kills the paper. The citation pattern is fair, and the limitations section is honest.\n\nBottom line: this deserves a serious referee. I would send it out, and I would ask for a validation protocol and data release before the quantitative claim is treated as settled. I would also bring it to a reading group as a good example of adaptive sensing under hardware constraints.","headline":"A solid, practical adaptive-recordings methods paper whose central quantitative claim is probably true but not yet out-of-sample, because the hyperparameters were selected on the same offline recordings used for the headline comparison.","tokens_in":16268,"tokens_out":3179,"would_cite":true,"duration_ms":28697,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["92C20","62L05","62C10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Adaptively reallocating which electrodes record captures 17.2 percentage points more neural activity than picking once and holding still.","keywords":["high-density microelectrode array","adaptive electrode selection","Thompson sampling","non-stationary spontaneous activity","Poisson-Gamma model","synchronized burst","center-of-activity trajectory","channel budget"],"falsifier":"Re-run the adaptive policy online at full array scale on several cultures, interrupting every few hours for the 26-minute full-array scan needed to compute the true oracle subset: if the attainment gap over static selection disappears, or if the top-1,024 active set turns out to be far more stable than the top-100 set in the 23x23 patch, the downscaled result does not generalize. A cheaper offline check is to replay the same algorithms with candidate sets composed of several spatially separated patches, so that the correlation structure resembles the full array.","tokens_in":15022,"feed_emoji":"🧠","tokens_out":15061,"duration_ms":110322,"temperature":0.7,"pith_summary":"High-density microelectrode arrays can sense tens of thousands of sites but can record only about a thousand channels at once, so an experimenter must choose which electrodes stay connected. The standard approach — scan the culture, lock onto the most active sites, and keep them fixed — assumes the best electrodes stay the best. This paper shows that assumption fails: across nine 34-hour recordings of dissociated cortical networks, nearly half (47.8%) of the top-100 active electrodes at the end were not among the initially most active. The paper's method treats channel assignment as a sequential decision problem, modeling each electrode's spike count with a discounted Poisson-Gamma posterior and re-allocating the fixed budget every 30 minutes via Thompson sampling. In offline replay against a per-window oracle, this adaptive policy captures 17.2 percentage points more of the achievable spike yield than static selection at the final time point, suggesting that uncertainty-aware exploration can keep long-term recordings aligned with an evolving network.","feed_headline":"17.2-point win: adaptive electrodes track shifting neural activity","feed_subtitle":"Active electrodes turn over by nearly half in 34 hours, so re-choosing sites every 30 minutes pays off.","key_machinery":"The carrying mechanism is the conjugate Poisson-Gamma model with a discounting step and Thompson sampling. Spike counts are modeled as $y_{c,t} \\sim \\mathrm{Poisson}(\\lambda_c \\Delta)$ with rate prior $\\lambda_c \\sim \\mathrm{Gamma}(\\alpha_c, \\beta_c)$, so posterior updates reduce to $\\alpha_c \\leftarrow \\alpha_c + y_{c,t}$ and $\\beta_c \\leftarrow \\beta_c + \\Delta$ for observed electrodes. Before each window all statistics are scaled by a discount factor $\\rho = 0.7$ ($\\alpha \\leftarrow \\rho\\alpha$, $\\beta \\leftarrow \\rho\\beta$), which preserves each electrode's estimated mean $\\alpha/\\beta$ while inflating its uncertainty, so the policy can revisit electrodes whose activity may have drifted. Selection is Thompson sampling: draw one plausible rate $\\tilde\\lambda_c \\sim \\mathrm{Gamma}(\\alpha_c/\\tau, \\beta_c/\\tau)$ per electrode and route the top-$k$ draws, with temperature $\\tau$ tuning exploration. A supporting identity links the spike-count objective to spatial readouts: if the selected electrodes capture fraction $p_s$ of events and the missed-event centroid is $m_c$, then $\\|m_s - m\\| = \\frac{1-p_s}{p_s}\\|m - m_c\\|$, so raising the captured fraction directly bounds the bias of the center-of-activity trajectory. The evaluation apparatus — dense reference data that allow a per-window oracle subset and an attainment score $100 R_t / R_t^*$ — is what turns the comparison into a quantitative claim.","core_discovery":"The paper's central claim is that adaptive electrode selection under a fixed readout budget can track non-stationary spontaneous activity better than fixed or heuristic policies, because the neural signal itself is restless. The evidence is a downscaled offline replay: from nine dense 34 h recordings of 529 electrodes, each algorithm picks 100 electrodes per 30 min window and is scored by the fraction of the oracle top-100 spike yield it captures. The discounted Poisson-Gamma Thompson sampling policy attains the highest oracle-relative score among static, random, and combinatorial epsilon-greedy baselines, and exceeds static selection by 17.2 percentage points at 34 h (two-sided exact paired sign-flip randomization test over the nine recordings, p = 0.004). The paper further documents the non-stationarity that motivates the method — top-100 active-electrode turnover reaches 47.8% by 34 h — and demonstrates in a single online recording that the policy can be executed in real time, capture the first synchronized burst, and support center-of-activity trajectory analysis of later recurrent bursts.","pith_inferences":["If the 17.2 percentage-point advantage is real at full scale, the natural next comparison the paper does not run is adaptive policy versus periodic full-array rescanning: a 26-minute rescan cycle could itself serve as the exploration mechanism, and the trade-off between losing continuous recording during scans and the policy's partial observability would decide which is preferable.","The evaluation selects 100 from a 529-electrode local patch; because the patch is spatially contiguous, its correlation structure may make activity shifts easier or harder to track than across the full 26,400-electrode array — a re-run of the same replay with candidate sets drawn from multiple distant patches would test whether the advantage generalizes.","The discount factor $\\rho = 0.7$ is fixed; an extension the paper leaves implicit is to tie $\\rho$ to the measured turnover or the decorrelation time of the firing-rate maps so the algorithm's forgetting matches the culture's own rate of change, which could improve or stabilize attainment across cultures.","The centroid-bias identity predicts a quantitative scaling — centroid error should grow roughly as $(1-p_s)/p_s$; checking that predicted scaling on the dense reference data would validate spike-count maximization as a proxy for preserving spatial summaries, without any new hardware."],"forward_implications":["Long-term HD-MEA recordings can track evolving activity without repeated full-array scans: reconfiguring the routed subset every 30 minutes keeps the recorded channels near the current activity peak under the same channel budget.","For spatial summaries such as center-of-activity trajectories, maximizing captured spike count is aligned with trajectory fidelity: the centroid-error identity shows that a higher captured fraction $p_s$ directly shrinks the bound on centroid bias.","Uncertainty-directed exploration, not random exploration, is the load-bearing ingredient: a combinatorial $\\varepsilon$-greedy baseline that explores at a fixed random rate does not match the Bayesian method's attainment.","Active-set turnover is a usable diagnostic: cultures with high top-100 turnover lose yield under scan-and-commit, and the 47.8% figure gives experimenters a quantitative threshold for deciding when adaptive routing is warranted.","The same formulation — temporal discounting plus Thompson sampling under a fixed channel budget — carries over to other high-density recording platforms with more sensing sites than readout channels, with the reward redefined to the relevant objective such as unit yield or information gain."],"supporting_citations":[{"why":"Introduces the switch-matrix HD-MEA architecture whose routing constraint creates the electrode-selection problem.","marker":"[15]"},{"why":"Describes the 1,024-channel, 26,400-electrode CMOS MEA used here, fixing the channel budget the method works under.","marker":"[17]"},{"why":"Surveys recording strategies for high channel-count densely spaced arrays, framing the readout bottleneck as a general problem.","marker":"[22]"},{"why":"Reports chronic co-variation between network configuration and activity, supporting the claim that electrode usefulness shifts over long recordings.","marker":"[26]"},{"why":"Defines the center-of-activity trajectory that motivates spike count as the selection objective and supplies the downstream analysis used in the online recording.","marker":"[29]"},{"why":"Provides the discounting approach for switching bandits that the algorithm uses to forget old observations.","marker":"[36]"},{"why":"Supplies the non-stationary Thompson sampling formulation the method is built on.","marker":"[37]"},{"why":"Gives the bandit-theoretic framing and Thompson sampling definition underlying the selection rule.","marker":"[38]"},{"why":"Provides the ISIN burst-detection framework used to identify synchronized bursts in the online recording.","marker":"[39]"}],"fun_headline_variants":["Adaptive electrodes beat static by 17.2 points","Reallocating electrodes every 30 min tracks shifting activity","Neural dynamics demand adaptive electrode sampling","Thompson sampling improves electrode selection for neurons","Electrode turnover hits 47.8%: adapt to keep up"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's quantitative evidence comes from a downscaled replay — choosing 100 of 529 electrodes in one local 23x23 patch — and the 17.2 percentage-point advantage is assumed to transfer to the real task of choosing 1,024 of 26,400 electrodes across the whole array, a transfer the single online demonstration does not establish.","fun_headline_variants_meta":{"raw":{"variants":["Adaptive electrodes beat static by 17.2 points","Reallocating electrodes every 30 min tracks shifting activity","Neural dynamics demand adaptive electrode sampling","Thompson sampling improves electrode selection for neurons","Electrode turnover hits 47.8%: adapt to keep up"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000548,"raw_usage":{"total_tokens":2641,"prompt_tokens":991,"completion_tokens":1650,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":1574}},"tokens_in":607,"tokens_out":1650,"duration_ms":11997,"temperature":1.0,"reasoning_tokens":1574,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:26:22.793220+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the adaptive policy online at full array scale on several cultures, interrupting every few hours for the 26-minute full-array scan needed to compute the true oracle subset: if the attainment gap over static selection disappears, or if the top-1,024 active set turns out to be far more stable than the top-100 set in the 23x23 patch, the downscaled result does not generalize. A cheaper offline check is to replay the same algorithms with candidate sets composed of several spatially separated patches, so that the correlation structure resembles the full array.","supporting_citations":[{"cited_title":"Switch-matrix-based high-density microelectrode array in cmos technology.IEEE Journal of Solid-State Circuits, 45(2):467–482, 2010","cited_arxiv_id":null,"evidence_quote":"Introduces the switch-matrix HD-MEA architecture whose routing constraint creates the electrode-selection problem."},{"cited_title":"A 1024-channel cmos microelectrode array with 26,400 electrodes for recording and stimulation of electrogenic cells in vitro.IEEE journal of solid-state circuits, 49(11):2705–2719,","cited_arxiv_id":null,"evidence_quote":"Describes the 1,024-channel, 26,400-electrode CMOS MEA used here, fixing the channel budget the method works under."},{"cited_title":"Chronic co-variation of neural network configuration and activity in mature dissociated cultures","cited_arxiv_id":null,"evidence_quote":"Reports chronic co-variation between network configuration and activity, supporting the claim that electrode usefulness shifts over long recordings."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the center-of-activity trajectory that motivates spike count as the selection objective and supplies the downstream analysis used in the online recording."},{"cited_title":"On upper-confidence bound policies for switching bandit problems","cited_arxiv_id":null,"evidence_quote":"Provides the discounting approach for switching bandits that the algorithm uses to forget old observations."},{"cited_title":"Thompson sampling for non-stationary bandit problems.Entropy, 27(1):51, 2025","cited_arxiv_id":null,"evidence_quote":"Supplies the non-stationary Thompson sampling formulation the method is built on."},{"cited_title":"Parameters for burst detection.Frontiers in Computational Neuroscience, 7:193, 2014","cited_arxiv_id":null,"evidence_quote":"Provides the ISIN burst-detection framework used to identify synchronized bursts in the online recording."}],"review_version":2}