{"id":"3fb78d2b-3807-4899-a1a1-6296c9cc823d","arxiv_id":"1906.11626","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"Sparse MLPs trained via SET plus neuron pruning achieve competitive performance on 15 datasets while pruning ~50% of hidden neurons and keeping parameter count linear in neuron count.","lead":"The paper empirically shows that sparse neural networks using adaptive sparse connectivity generalize better than dense fully-connected networks on limited data. A smart generalist might read it to see if sparsity techniques can make deep learning more efficient and less data-hungry.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"No ablations isolate adaptive sparse connectivity from reduced parameter count or optimization differences","rationale":"The reader's weakest_assumption correctly identifies the missing isolation of the mechanism. The full text does not supply the required controls or ablations, so the concern stands and the unverdicted status is unaffected.","tokens_in":1637,"tokens_out":259,"duration_ms":19311,"concrete_test":"On the same 15 datasets and MLP architectures, replace the SET evolutionary rewiring with a static random sparse mask of identical final density and retrain; if the generalization advantage over dense models disappears, the adaptivity is not the operative factor.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that generalization gains are caused by the adaptive sparse connectivity (SET + 50% neuron pruning) rather than by the strict parameter budget itself, the 15 chosen datasets, or unstated differences in training dynamics. The experiments compare sparse MLPs to fully-connected counterparts on 15 datasets but provide no matched-parameter dense baselines, no fixed-sparsity (non-adaptive) controls, and no ablation that disables the evolutionary rewiring while preserving the final sparsity level. Without these, the attribution to the adaptive mechanism remains untested.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that intrinsically sparse neural networks using adaptive sparse connectivity (via a combination of Sparse Evolutionary Training (SET) and neuron pruning) achieve better generalization than fully-connected counterparts. The proposed method is tested on MLPs across 15 datasets, where it zeros out approximately 50% of hidden neurons during training while maintaining a linear number of parameters with respect to the number of neurons, yielding competitive classification and generalization performance.","tokens_in":1746,"tokens_out":484,"duration_ms":23233,"significance":"If substantiated with proper controls, the result would provide empirical support for the idea that adaptive sparsity mechanisms can enhance generalization under parameter budgets, offering a potential route to more efficient deep learning models in data-limited settings. The multi-dataset evaluation and the explicit combination of SET with pruning are concrete strengths that could be built upon.","major_comments":[{"comment":"Experiments section: the central claim that adaptive sparse connectivity improves generalization over fully-connected networks is not isolated from the effect of the strict parameter budget, as the manuscript reports no matched-parameter dense baselines (i.e., dense MLPs with the same number of parameters as the final sparse models) and no fixed-sparsity (non-adaptive) controls that disable the evolutionary rewiring while preserving the final sparsity level.","section":"Experiments"},{"comment":"Experiments section: no ablation is presented that disables the adaptive rewiring component of SET while keeping the 50% neuron pruning and parameter budget fixed, leaving open whether any observed gains are attributable to the adaptive mechanism itself rather than reduced parameter count or optimization dynamics.","section":"Experiments"},{"comment":"Results section: the reported performance on the 15 datasets lacks details on statistical tests, variance across multiple runs, exact data splits, and baseline architectures, making it impossible to assess whether the generalization advantage is robust or reproducible.","section":"Results"}],"minor_comments":[{"comment":"The abstract states that the method 'zeros out around 50% of the hidden neurons' but does not clarify whether this is a fixed target or an emergent outcome of the combined SET+pruning procedure; a precise description of the pruning schedule would improve clarity.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments, which highlight important controls needed to strengthen the isolation of adaptive sparsity effects. We address each major comment below and will incorporate revisions to improve the manuscript.","responses":[{"response":"We agree that matched-parameter dense baselines and fixed-sparsity controls are necessary to better isolate the role of adaptive connectivity from the parameter budget itself. The original comparisons were to standard fully-connected MLPs (which use more parameters), and the sparse models achieve competitive results under a strict budget. In revision, we will add dense MLPs with parameter counts matched to the final sparse models and non-adaptive fixed-sparsity controls that preserve the same sparsity level without rewiring.","revision_made":"yes","referee_comment":"[Experiments] Experiments section: the central claim that adaptive sparse connectivity improves generalization over fully-connected networks is not isolated from the effect of the strict parameter budget, as the manuscript reports no matched-parameter dense baselines (i.e., dense MLPs with the same number of parameters as the final sparse models) and no fixed-sparsity (non-adaptive) controls that disable the evolutionary rewiring while preserving the final sparsity level."},{"response":"This is a fair observation. While SET's adaptive rewiring is a core component of the proposed combination with neuron pruning, an explicit ablation would clarify its contribution. We will add this ablation in the revised manuscript by comparing the full adaptive SET+pruning approach against a variant that applies the same neuron pruning and parameter budget but disables evolutionary rewiring after initialization.","revision_made":"yes","referee_comment":"[Experiments] Experiments section: no ablation is presented that disables the adaptive rewiring component of SET while keeping the 50% neuron pruning and parameter budget fixed, leaving open whether any observed gains are attributable to the adaptive mechanism itself rather than reduced parameter count or optimization dynamics."},{"response":"We acknowledge that these experimental details were not sufficiently reported. The experiments involved multiple runs, but variance, statistical tests, exact splits, and architecture specifications were omitted from the results section. In the revision, we will include standard deviations across runs, specify data splits and preprocessing, detail all baseline architectures, and report statistical significance tests (such as paired t-tests) to demonstrate robustness.","revision_made":"yes","referee_comment":"[Results] Results section: the reported performance on the 15 datasets lacks details on statistical tests, variance across multiple runs, exact data splits, and baseline architectures, making it impossible to assess whether the generalization advantage is robust or reproducible."}],"tokens_in":1300,"tokens_out":556,"duration_ms":22982,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"This paper's core observation is that MLPs trained with SET combined with 50% neuron pruning reach competitive accuracy on 15 classification datasets while using a strict parameter budget. The new element is the addition of neuron pruning to the existing SET procedure, which keeps the number of parameters linear in the number of neurons and removes half the hidden units during training. That extension is straightforward and the multi-dataset evaluation is a reasonable check on whether the approach holds up across tasks. The work sits squarely in the sparse-training literature and offers a practical tweak rather than a new framework. The main limitation is that the experiments compare the resulting sparse models only to fully-connected counterparts. There are no matched-parameter dense baselines, no fixed-sparsity controls without the evolutionary rewiring, and no ablation that disables adaptation while preserving the final sparsity level. Without those, it is not possible to attribute any generalization difference specifically to the adaptive connectivity mechanism instead of the reduced parameter count or other optimization details. The paper is empirical throughout, so the strength of the central claim depends on those controls being added. Readers working on sparse neural network training will find the reported numbers and the simple implementation useful. The work is coherent enough on its own terms to merit peer review, though the authors should address the missing ablations before publication.","headline":"SET plus neuron pruning gives competitive sparse MLP results on 15 datasets, but the experiments do not isolate adaptive rewiring from simple parameter reduction.","tokens_in":2221,"tokens_out":330,"would_cite":false,"duration_ms":17443,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Empirical sparsity regularization in MLPs has no overlap with RS forcing chain","alignment":"orthogonal","rationale":"Paper central machinery is SET evolutionary rewiring + neuron pruning (α=0.04, β=10, γ=40) on MLPs to enforce strict parameter budget and implicit regularization, tested on 15 tabular datasets. No ratio-symmetric cost, J(x) functional equation, φ-ladder, 8-tick periodicity, or parameter-free constant derivation appears. Domain (cs.NE empirical generalization) lies outside RS scope; no RS theorem is invoked or contradicted.","tokens_in":45118,"confidence":"high","tokens_out":141,"duration_ms":6077,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Intrinsically sparse neural networks with adaptive sparse connectivity generalize better than fully-connected networks under a strict parameter budget.","keywords":["sparse neural networks","generalization","adaptive sparse connectivity","neuron pruning","Sparse Evolutionary Training","multilayer perceptron","deep learning","parameter budget"],"falsifier":"Run controlled experiments training sparse adaptive and dense networks on identical datasets with matched parameter counts and hyperparameters, then check if the sparse version has higher test accuracy; failure to do so would disprove the claim.","tokens_in":2543,"feed_emoji":"","tokens_out":611,"duration_ms":23193,"temperature":0.7,"pith_summary":"The paper empirically demonstrates that neural networks kept sparse by design during training, through adaptive connection changes, achieve better generalization than dense networks on classification tasks with limited data. It introduces a training method that merges Sparse Evolutionary Training with neuron pruning to eliminate about half the hidden neurons while keeping the number of parameters linear in the neuron count. Experiments on multilayer perceptrons across 15 datasets show competitive accuracy and generalization. If correct, this indicates that maintaining a fixed parameter budget via sparsity can serve as an effective regularizer.","feed_headline":"Adaptive sparse networks generalize better than dense ones","feed_subtitle":"They enforce a fixed parameter budget in training and deliver competitive results after pruning half the neurons on 15 datasets.","key_machinery":"Adaptive sparse connectivity enforced by combining Sparse Evolutionary Training (SET) with neuron pruning, which maintains a strict parameter budget and prunes half the hidden neurons.","core_discovery":"Intrinsically sparse neural networks with adaptive sparse connectivity, which by design have a strict parameter budget during the training phase, have better generalization capabilities than their fully-connected counterparts. The proposed technique combines the Sparse Evolutionary Training procedure with neurons pruning to zero out around 50% of the hidden neurons during training, while having a linear number of parameters to optimize with respect to the number of neurons, yielding competitive classification and generalization performance on 15 datasets.","pith_inferences":["This approach might apply to other network types like CNNs for vision tasks.","Sparsity during training could complement or replace techniques like weight decay or dropout.","Fixed parameter budgets may help in resource-constrained training scenarios.","Further tests could vary the pruning rate to find optimal sparsity levels."],"forward_implications":["Sparse models show improved generalization compared to dense counterparts.","The method achieves competitive classification performance on 15 datasets.","Parameter count remains linear with respect to the number of neurons.","About 50% of hidden neurons can be pruned without harming performance."],"fun_headline_variants":["Sparse adaptive connectivity leads to better generalization","Strict parameter budget helps sparse nets generalize","Combining SET and pruning for sparse MLP generalization","Adaptive sparse models competitive after pruning half neurons"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Observed generalization improvements result from the adaptive sparse connectivity itself, not from specific dataset properties, the 50% pruning rate, or unmentioned differences in training dynamics.","fun_headline_variants_meta":{"raw":{"variants":["Sparse adaptive connectivity leads to better generalization","Strict parameter budget helps sparse nets generalize","Combining SET and pruning for sparse MLP generalization","Adaptive sparse models competitive after pruning half neurons"]},"model":"grok-4.3","cost_usd":0.010337,"raw_usage":{"total_tokens":4535,"prompt_tokens":585,"num_sources_used":0,"completion_tokens":51,"cost_in_usd_ticks":103374500,"prompt_tokens_details":{"text_tokens":585,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3899,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":585,"tokens_out":51,"duration_ms":29773,"temperature":1.0,"reasoning_tokens":3899,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-25T13:57:38.967145+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Run controlled experiments training sparse adaptive and dense networks on identical datasets with matched parameter counts and hyperparameters, then check if the sparse version has higher test accuracy; failure to do so would disprove the claim.","supporting_citations":[],"review_version":1}