{"id":"e43fda8b-3cfd-4894-b4bb-1c86d63a19ed","arxiv_id":"2510.09174","paper_version":3,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Hierarchical Re-Basin merging induces stronger adversarial and perturbation robustness in combined models as more participants are added, but produces larger performance drops than previously reported.","lead":"This paper proposes a hierarchical version of Git Re-Basin for merging multiple trained neural network models, claiming it outperforms the standard MergeMany approach. It reports that the method induces increasing adversarial and perturbation robustness as more models are included, though with a larger accuracy drop than in prior work.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Robustness gains may stem from model selection/training details rather than the hierarchical Re-Basin procedure itself","rationale":"The reader's weakest assumption directly identifies the same attribution gap. Because the paper is empirical with no formal verification or parameter-free derivations, isolating the merging operator from base-model confounders is the minimal requirement for the scaling claim. The proposed test is a direct, minimal change to the existing experimental pipeline that would falsify or support the causal link.","tokens_in":1545,"tokens_out":312,"duration_ms":41184,"concrete_test":"Fix the identical collection of pre-trained models used in the standard MergeMany baseline and re-apply only the hierarchical Re-Basin steps at increasing depths; recompute adversarial robustness (PGD or AutoAttack accuracy) and clean accuracy for each depth. If the robustness-vs-depth slope flattens or reverses while clean accuracy remains matched, the claim does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that increasing the number of models in the hierarchy causally strengthens adversarial/perturbation robustness via the Re-Basin alignment and merging steps. The abstract reports a larger performance drop than prior work, which could interact with robustness metrics. Without explicit controls (e.g., holding the exact set of base models fixed across standard MergeMany vs. hierarchical variants, or ablating across independent training seeds while varying only the merging depth), the scaling effect could be an artifact of how the participating models were chosen or trained rather than the procedure.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript proposes a hierarchical variant of the Git Re-Basin model-merging procedure that is claimed to outperform the standard MergeMany baseline. It reports that Re-Basin merging induces adversarial and perturbation robustness, with the magnitude of this robustness increasing as more models participate in the hierarchy, while also documenting a substantially larger clean-data performance drop than was reported in the original Re-Basin work.","tokens_in":1660,"tokens_out":490,"duration_ms":80054,"significance":"If the robustness scaling is shown to be causally attributable to the hierarchical alignment-and-merge steps rather than to model-selection or training artifacts, the result would be of moderate interest to the model-merging community: it would identify a new, training-free regularization pathway whose strength can be tuned by hierarchy depth. The larger performance drop observation is also potentially useful for understanding the robustness-accuracy trade-off in merging, provided it is placed in context with prior baselines.","major_comments":[{"comment":"§4 Experiments (and associated tables): the central claim that robustness strengthens with hierarchy depth requires an ablation in which the exact set of base models and training seeds is held fixed while only the merging depth is varied. No such controlled comparison is described; without it the scaling effect cannot be confidently attributed to the hierarchical Re-Basin procedure itself rather than to incidental differences in the participating models.","section":"§4 Experiments"},{"comment":"§4.2 and Table 3: the reported adversarial and perturbation robustness numbers are presented without error bars across independent training runs or statistical significance tests. Given that the abstract already notes a larger performance drop than prior work, the absence of these controls makes it impossible to judge whether the robustness gains are reliable or simply correlated with the larger accuracy degradation.","section":"§4.2"}],"minor_comments":[{"comment":"The abstract states empirical findings without any reference to experimental setup, number of models, datasets, or evaluation protocols; a one-sentence summary of the experimental regime would improve readability.","section":null},{"comment":"Notation for the hierarchical merging levels is introduced without a clear diagram or pseudocode; a small figure illustrating the tree structure would clarify the algorithm.","section":"§3"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments, which help clarify the strength of our claims. We address each major point below and indicate the revisions we will incorporate.","responses":[{"response":"We agree that the current experimental design does not fully isolate the effect of hierarchy depth from the number of participating models. In our reported results, deeper hierarchies incorporate additional models by construction, which could introduce confounding factors. We will add a controlled ablation in the revised manuscript that fixes the exact set of base models and training seeds while varying only the merging depth (e.g., by constructing hierarchies of different depths from the same pool of models). This will allow a direct attribution of any robustness scaling to the hierarchical alignment-and-merge steps.","revision_made":"yes","referee_comment":"[§4 Experiments] §4 Experiments (and associated tables): the central claim that robustness strengthens with hierarchy depth requires an ablation in which the exact set of base models and training seeds is held fixed while only the merging depth is varied. No such controlled comparison is described; without it the scaling effect cannot be confidently attributed to the hierarchical Re-Basin procedure itself rather than to incidental differences in the participating models."},{"response":"We acknowledge that the absence of error bars and statistical tests limits the ability to assess reliability, particularly in light of the larger clean accuracy drop we report. We will rerun the primary experiments across multiple independent training seeds, report standard deviations or confidence intervals for both clean accuracy and robustness metrics, and include statistical significance tests (e.g., paired t-tests) comparing hierarchical Re-Basin against baselines. These additions will be incorporated into the revised §4.2 and Table 3.","revision_made":"yes","referee_comment":"[§4.2] §4.2 and Table 3: the reported adversarial and perturbation robustness numbers are presented without error bars across independent training runs or statistical significance tests. Given that the abstract already notes a larger performance drop than prior work, the absence of these controls makes it impossible to judge whether the robustness gains are reliable or simply correlated with the larger accuracy degradation."}],"tokens_in":1227,"tokens_out":455,"duration_ms":33663,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing here is that the hierarchical merging scheme beats standard MergeMany and appears to build adversarial and perturbation robustness that grows with more models, though it also triggers a bigger accuracy drop than the original Re-Basin paper reported. The new element is the hierarchical structure itself, which they position as a practical way to get robustness as a side effect of merging. They are straightforward about the performance cost, which is useful to see. The scaling observation is the part that could matter for people trying to make merged models more reliable without extra training. On the downside, the abstract gives almost no experimental details, so it is hard to judge whether the robustness really traces to the hierarchical steps or to how the base models were chosen and trained. Without ablations that hold the exact set of models fixed while changing only the merging depth, the scaling effect could be driven by model selection rather than the procedure. The stress-test concern about that point still looks live based on what is shown. This paper is aimed at the small group already working on Re-Basin and model merging. A reader in that niche might pick up the hierarchical idea and the robustness observation as something to test further, but it is not yet strong enough to shift how most people merge models. I would send it to peer review so the authors can supply the controls, baselines, and statistical checks that are currently absent.","headline":"Hierarchical Re-Basin produces scaling robustness but the larger performance drop and missing controls leave open whether the gains are truly from the merging procedure.","tokens_in":2142,"tokens_out":346,"would_cite":false,"duration_ms":36253,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"We propose a hierarchical model merging scheme that significantly outperforms the standard MergeMany algorithm... Re-Basin induces adversarial and perturbation robustness... stronger the more models participate"},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/AlphaCoordinateFixation.lean","rs_theorem":"J_uniquely_calibrated_via_higher_derivative","paper_passage":"Re-Basin seems to act as a sort of regularization, positively impacting adversarial and perturbation robustness"}],"headline":"Hierarchical Re-Basin model merging and robustness observations lie outside RS forcing chain","alignment":"orthogonal","rationale":"Paper's central machinery (permutation-based Re-Basin interpolation, hierarchical MergeMany variant, empirical robustness/Lipschitz/weight-norm trends on CIFAR-10 MLPs) operates in the domain of neural-network loss landscapes and empirical regularization. RS framework derives J-cost, φ-ladder, 8-tick periodicity, and physical constants from a single distinction with zero adjustable parameters (see reality_from_one_distinction, AbsoluteFloorClosure, Cost.FunctionalEquation). No shared structure (cosh-cost, ratio symmetry, parameter-free constant derivation) is present; the work is a standard ML experiment with no opinion from RS.","tokens_in":42988,"confidence":"high","tokens_out":333,"duration_ms":12542,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Hierarchical Re-Basin merges models while building resistance to adversarial attacks and input perturbations.","keywords":["Re-Basin","model merging","hierarchical merging","adversarial robustness","perturbation robustness","neural network ensembles","regularization"],"falsifier":"Run the same set of base models through both flat Re-Basin and the hierarchical version, then measure whether the hierarchical version still shows higher adversarial and perturbation accuracy.","tokens_in":2452,"feed_emoji":"🛡️","tokens_out":532,"duration_ms":29515,"temperature":0.7,"pith_summary":"The paper develops a hierarchical scheme for merging trained neural networks with the Re-Basin method and shows it beats the standard flat MergeMany algorithm. The merged models gain resistance to adversarial examples and small perturbations, and this resistance grows stronger when more base models participate in the hierarchy. The same procedure produces a larger drop in clean-data performance than earlier Re-Basin work reported.","feed_headline":"Hierarchical merging builds attack resistance into models","feed_subtitle":"Robustness to adversarial examples and perturbations grows with more models in the tree, though clean accuracy drops more than prior reports","key_machinery":"The hierarchical merging scheme, which applies Re-Basin recursively to successive groups of models in a tree structure instead of merging all models at once.","core_discovery":"Re-Basin applied through a hierarchical merging procedure induces adversarial and perturbation robustness into the resulting models, with the robustness effect becoming stronger the more models participate in the hierarchy. The hierarchical algorithm also delivers better merged-model performance than the flat MergeMany baseline, although the accuracy cost on standard tasks is larger than previously observed.","pith_inferences":["The hierarchy may act as a built-in regularizer that trades some accuracy for robustness.","Similar robustness patterns could be tested by applying hierarchy to other merging algorithms.","Deployment settings that value robustness over peak accuracy might benefit from deeper hierarchies."],"forward_implications":["Merged models gain resistance to adversarial perturbations that scales with the number of base models.","The hierarchical scheme produces stronger overall merged performance than flat merging.","Clean-data accuracy falls more than earlier Re-Basin reports indicated.","Robustness benefits appear consistently across the tested merging depths."],"fun_headline_variants":["Re-Basin hierarchy induces robustness to adversarial attacks","Robustness increases as more models join Re-Basin hierarchy","Hierarchical Re-Basin outperforms standard MergeMany merging","Accuracy drops more than reported with hierarchical Re-Basin"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The observed robustness gains and performance drop arise from the hierarchical merging procedure itself rather than from the particular models, training runs, or evaluation protocols used.","fun_headline_variants_meta":{"raw":{"variants":["Re-Basin hierarchy induces robustness to adversarial attacks","Robustness increases as more models join Re-Basin hierarchy","Hierarchical Re-Basin outperforms standard MergeMany merging","Accuracy drops more than reported with hierarchical Re-Basin"]},"model":"grok-4.3","cost_usd":0.007193,"raw_usage":{"total_tokens":3154,"prompt_tokens":500,"num_sources_used":0,"completion_tokens":55,"cost_in_usd_ticks":71928000,"prompt_tokens_details":{"text_tokens":500,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2599,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":500,"tokens_out":55,"duration_ms":28760,"temperature":1.0,"reasoning_tokens":2599,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-21T20:29:50.437867+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Run the same set of base models through both flat Re-Basin and the hierarchical version, then measure whether the hierarchical version still shows higher adversarial and perturbation accuracy.","supporting_citations":[],"review_version":1}