{"id":"20386df2-8fc9-4b92-b5ce-638fff24de27","arxiv_id":"2607.06628","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":2,"one_line_summary":"In grokking modular arithmetic, weight direction portably carries circuit identity across independent runs while weight norm only sets susceptibility to overwrite and a weak delay effect.","lead":"Swapping the direction of a neural network's weights between two independent training runs transfers which circuit the network ends up using; the weight magnitude only controls how easily that identity can be overwritten. The result gives a clean causal split between which solution a network approaches and how committed it is to that solution.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged proxy metric.","rationale":"The reader's strongest claim accurately restates the paper's headline results and the weakest assumption correctly identifies the embedding-spectrum proxy. That proxy is the softest point, but it is already disclosed, does not create an internal contradiction (the same metric is used for both endpoints and the chimera), and is buttressed by the matched-random control and the optimizer ablation. No additional load-bearing concern (statistical, geometric, or experimental) rises to the level that would move the verdict. Therefore the CONDITIONAL verdict with high confidence remains appropriate; no adjustment is warranted.","tokens_in":11391,"tokens_out":475,"duration_ms":7053,"concrete_test":"On the three pairs already used for layer localization (Appendix A), recompute CFS lean after the reverse-radial intervention using an independent non-embedding fingerprint (e.g., Fourier components of the attention or MLP weight matrices, or logit-lens Fourier peaks on held-out inputs). If the 3/3 sign agreement with the embedding-spectrum CFS lean holds, the proxy concern does not overturn the 40/40 claim; if signs flip on any pair, the identity-transfer result weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (40/40 direction-driven identity transfer, donor-specific vs matched-random, threshold location predicted by recipient norm with perfect 20/20 class separation) is internally well-supported by the reported endpoint swaps, geodesic interpolations, bisection localization, mid-norm and matched-random controls, and optimizer-state ablation. The reader's weakest assumption—that the embedding power-spectrum cosine similarity is a faithful proxy for full circuit identity—is already the paper's own stated limitation (Section 7, Appendix A circularity note, Limitations). No stronger load-bearing flaw is present: the metric is applied consistently to both parents and chimeras, the discrete-log reordering is correctly task-matched, the matched-random control isolates content from angular size, and the threshold–norm separation is a pure ordinal claim under the intervention that holds the recipient norm fixed. The non-cyclic-task and single-architecture limits affect scope, not the correctness of the reported dissociation.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces cross-trajectory chimera interventions that recombine the weight norm of one independently trained run with the unit direction of another, then continue training under AdamW. On modular addition and multiplication (p=59, one-layer transformer), direction carries a transferable, donor-specific circuit identity: reverse-radial chimeras adopt the donor circuit in 40/40 recombinations (CFS lean sign-consistent), while an angle-matched random control produces no lean. Identity transfer is threshold-like under geodesic (slerp) interpolation of direction at fixed recipient norm; the flip location t* is predicted by recipient norm, separating perfectly by high/low-norm class over all 20 pairs (joint exact permutation p=1.9e-4). Norm itself produces only a modest, non-localizable delay effect (~30% fractional displacement) and no identity signal. An adaptive bisection localizes t* to ±1/64; an optimizer-state ablation (reset / recipient / donor Adam moments) leaves both the identity transfer and the threshold–norm separation intact.","tokens_in":11577,"tokens_out":947,"duration_ms":8384,"significance":"If the results hold, the work supplies a clean causal dissociation between portable circuit identity (direction) and state-dependent susceptibility (norm) that single-trajectory interventions cannot establish. The chimera construction, matched-random control, and reusable bisection procedure are concrete methodological contributions for probing basin membership across runs. Strengths include fully independent (disjoint-seed) pairs, exact sign tests, perfect class separation under a pure ordinal claim, and an optimizer-state ablation that rules out moment history as the driver. The findings are limited to two cyclic-group tasks and a single architecture, but within that scope they give a falsifiable geometric division of labour that is of clear interest to the grokking and mechanistic-interpretability communities.","major_comments":[{"comment":"The circuit-identity metric (normalized power spectrum of token-embedding rows, with discrete-log reordering for multiplication) is only a proxy for full circuit equivalence (Limitations; §7 / Appendix A). While the paper correctly flags circularity for embedding-swap cells and applies the metric consistently to parents and chimeras, the central 40/40 and threshold claims rest on this proxy. A modest additional check—e.g., attention-pattern or logit-lens agreement on a subset of pairs—would substantially strengthen the claim that CFS lean reports basin membership rather than embedding-spectrum coincidence alone.","section":null},{"comment":"Pair selection deliberately favours large recipient-norm differences (§5, “Scope of the norm–threshold relationship”), producing two well-separated classes rather than a continuum. The perfect 20/20 separation and joint p=1.9e-4 are therefore an ordinal claim under that design. The manuscript should either (a) add a denser sweep of intermediate-norm recipients or (b) more explicitly bound the claim to class separation, so that readers do not over-read a continuous law t*(r).","section":null}],"minor_comments":[{"comment":"Figure 4 shaded bands are described as visual aids only; a short caption note that they are not a fitted relationship would prevent misreading.","section":null},{"comment":"The non-additivity of layer-group identity signals (Appendix A) is left as an open puzzle; a sentence in the Discussion noting that combined swaps still re-grok to full accuracy (already stated) would help readers rule out instability.","section":null},{"comment":"Sparse-parity exploration is mentioned only in Limitations; a brief footnote on the hyperparameter mismatch would make the scope decision fully transparent.","section":null},{"comment":"Table 1 and Figure 6 are clear; ensuring that the ±1/64 half-width is printed on every t* panel would aid reproducibility.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The central experimental package (40/40, matched-random, perfect class separation, optimizer ablation) is unusually clean for a grokking paper. The proxy-metric and pair-selection issues are real but already partially acknowledged; they do not undermine the reported dissociation and are fixable with modest additional checks or clearer bounding language. Fit for a solid ML venue is good."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The new piece is the cross-trajectory chimera itself: recombine one run’s norm with another’s unit direction, continue training, and ask what actually travels. Direction carries donor-specific circuit identity in 40/40 endpoint swaps; the angle-matched random control stays near zero lean, so it is content, not just angular size. Transfer is threshold-like under geodesic interpolation, and that threshold is predicted by the recipient’s own norm—perfect slow/fast separation over all 20 pairs (joint exact p = 1.9e-4). Norm itself only weakly shifts delay and carries no identity. Optimizer-state ablation (reset / recipient / donor moments) leaves both results intact.\n\nWhat works: disjoint seed pairs, pre-registered decision thresholds, mid-norm and matched-random controls, adaptive bisection to ±1/64, and honest scope statements. The discrete-log reordering for multiplication is correctly task-matched. The paper does not overclaim a continuous t*(r) law or full circuit equivalence.\n\nSoft spots are real but already flagged and proportionate. The identity metric is an embedding power-spectrum proxy; the authors correctly refuse to interpret the circular embedding-swap cells and treat the non-additive hidden-layer pattern as suggestive only. Scope is two cyclic-group tasks and one small transformer; sparse parity is left for later because the hyperparameter regime differs. Pair selection favors large norm gaps, so the threshold result is a clean ordinal separation, not a fitted curve. None of this overturns the reported dissociation.\n\nThis is for people who care about basin selection, portable circuits, and causal probes in grokking / mech-interp. The intervention template is reusable. I would send it to peer review; the evidence is sharp enough to deserve referee time even if the metric and scope get tightened. Worth reading and citing if you work in this corner.","headline":"Clean causal dissociation: direction portably carries donor-specific circuit identity (40/40), and the flip threshold is predicted by recipient norm with perfect class separation.","tokens_in":12246,"tokens_out":475,"would_cite":true,"duration_ms":4769,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Direction of the weights decides which circuit a grokking run will adopt; their magnitude only decides how easily that choice can be overwritten.","keywords":["grokking","chimera interventions","weight norm","weight direction","circuit identity","modular arithmetic","cross-trajectory transfer","basin selection"],"falsifier":"Find a pair of modular-arithmetic runs whose embedding spectra look highly similar yet whose full circuits (attention patterns, MLP features, unembedding) differ, or a chimera that lands in the donor's spectrum but not the donor's actual circuit; either would break the identity metric that drives the 40/40 result.","tokens_in":12600,"feed_emoji":"🔀","tokens_out":917,"duration_ms":13543,"temperature":0.7,"pith_summary":"When two neural networks train on the same modular-arithmetic task from different random seeds, they often settle into different structured solutions after a long delay known as grokking. This paper asks which pieces of a partially trained network can be transplanted into a second, independently trained network and still determine its outcome. By splitting every weight vector into a size (norm) and a pure direction, then recombining one run's size with the other's direction, the authors build chimeras that continue training. Across forty independent recombinations the final circuit always follows the direction donor, not the size donor; a control that matches the angular distance but uses a random direction produces no such transfer. The switch itself is threshold-like rather than a smooth blend, and the location of that threshold is predicted by how large the recipient's weights already are: high-norm recipients flip early, low-norm ones resist. Size itself carries only a weak, whole-network effect on how long generalization takes. Direction therefore indexes which solution a trajectory approaches, while magnitude governs how susceptible that identity is to being overwritten.","feed_headline":"Weight direction picks the circuit; size only sets how easy it is to flip","feed_subtitle":"Swapping directions across independent grokking runs transfers circuit identity in 40 of 40 trials","key_machinery":"Cross-trajectory chimera interventions: given two independently trained runs, decompose each weight vector into norm r and unit direction u, recombine one run's norm with the other's direction (or interpolate directions on the geodesic at fixed recipient norm), continue training, and measure whether circuit identity (cosine similarity of embedding power spectra) follows the direction donor and delay follows the norm donor.","core_discovery":"Cross-trajectory chimera interventions dissociate the portable roles of weight magnitude and direction in grokking. Implanting a donor's unit direction at the recipient's norm drives the continued run to the donor's eventual circuit identity in 40/40 independent recombinations on two modular-arithmetic tasks; the transfer is donor-content-specific and threshold-like. The interpolation threshold at which identity flips is predicted by the recipient's weight norm, separating perfectly by norm class over all 20 pairs. Norm carries only a modest distributed delay effect and no identity signal.","pith_inferences":["The same norm-versus-direction dissociation may appear in other delayed-generalization regimes once a reliable circuit fingerprint exists.","If high-norm networks sit in shallower landscape regions, early training interventions that keep norms large could enlarge the set of reachable circuits.","Non-additivity of the identity signal across layers hints that circuit membership is a collective property of the weight configuration rather than a sum of layer-wise votes.","The chimera construction itself is architecture-agnostic and could test portability of other geometric or spectral features beyond modular arithmetic."],"forward_implications":["Direction alone can be used as a portable control knob to select among known circuits without re-training from scratch.","Recipient weight norm becomes a measurable state variable that predicts how easily a network's identity can be overwritten by a foreign direction.","The adaptive bisection procedure supplies a cheap, reusable way to localize stability thresholds for any intervention that is stable under its training protocol.","Norm and direction play non-interchangeable causal roles, so single-trajectory rescaling or freezing experiments that mix the two axes cannot isolate portability.","Identity transfer is a basin switch, not a continuous blend, so intermediate directions will typically collapse to one parent circuit or the other."],"fun_headline_variants":["Direction implants donor circuits across runs in 40/40 swaps","Weight direction sets circuit identity; norm only gates the flip","Cross-trajectory direction swaps transfer circuits; norms set thresholds","Unit direction carries portable circuit ID; magnitude predicts overwrite","Norm class perfectly predicts direction-swap flip thresholds"],"cache_read_input_tokens":12544,"weakest_assumption_plain":"The claim treats the power spectrum of the token embeddings as a faithful stand-in for the whole circuit's identity, so that spectral similarity correctly reports which solution a chimera has entered.","fun_headline_variants_meta":{"raw":{"variants":["Direction implants donor circuits across runs in 40/40 swaps","Weight direction sets circuit identity; norm only gates the flip","Cross-trajectory direction swaps transfer circuits; norms set thresholds","Unit direction carries portable circuit ID; magnitude predicts overwrite","Norm class perfectly predicts direction-swap flip thresholds"]},"model":"grok-4.5","effort":"low","cost_usd":0.005866,"raw_usage":{"total_tokens":1573,"prompt_tokens":797,"num_sources_used":0,"completion_tokens":81,"cost_in_usd_ticks":58660000,"prompt_tokens_details":{"text_tokens":797,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":695,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":797,"tokens_out":81,"duration_ms":8212,"temperature":1.0,"reasoning_tokens":695,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T01:07:37.864283+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Find a pair of modular-arithmetic runs whose embedding spectra look highly similar yet whose full circuits (attention patterns, MLP features, unembedding) differ, or a chimera that lands in the donor's spectrum but not the donor's actual circuit; either would break the identity metric that drives the 40/40 result.","supporting_citations":[],"review_version":1}