{"id":"ae8afb76-ebde-45da-8949-b69ea66faefe","arxiv_id":"2606.28525","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A history-defined direction v_rev in activation space mediates early fine-tuning reversion; blocking motion along it reduces harmful reversion while preserving task performance.","lead":"The paper frames fine-tuning reversion in AI models as a geometric 'gravitational' pull back toward dominant behavioral patterns set in early training. If the view holds, it offers a concrete direction to monitor and block when trying to keep safety properties stable after alignment.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Blocking experiment supports correlation with v_rev but does not isolate dominant-manifold causation from other trajectory correlations","rationale":"The reader’s weakest assumption directly names the gap between the observed intervention effect and the specific manifold-based causal story; the blocking result is real but its attribution to the dominant early manifold versus other correlated geometric or optimization properties is the precise point that remains under-determined from the reported evidence.","tokens_in":1853,"tokens_out":351,"duration_ms":22875,"concrete_test":"Recompute the T=100 blocking results using three control directions per run: (1) a random vector drawn from the same activation distribution at step 0, (2) the top principal component of activations at step 20, and (3) the gradient direction at step 20 projected orthogonal to v_rev; if any control produces a comparable shift in alignment or harmfulness, the specificity to the history-defined reversion direction is weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The intervention (selectively blocking motion along v_rev) alters final alignment (0.648 → -0.211) and harmfulness (19% → 8.5%), yet v_rev is constructed from early-training history; any direction that is (a) stable early, (b) aligned with later gradients, or (c) lies in a high-variance subspace of the activation geometry could produce similar blocking effects. The paper does not report controls that orthogonalize v_rev against such alternatives (e.g., random early directions, late-training principal components, or gradient subspaces at the same step). Without those, the geometric “gravitational” interpretation remains one of several consistent explanations for the observed causal mediation.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes a geometric 'gravitational' interpretation of fine-tuning reversion phenomena, arguing that early training phases create dominant behavioral manifolds while later alignment phases are shallower displacements; subsequent updates therefore inherit a reversion component toward a history-defined direction v_rev. It reports that alignment with v_rev rises rapidly (from cos = 0.429 +/- 0.052 to 0.647 +/- 0.021 by step 20) and exceeds the p99 of an isotropic null across 24 run-step pairs, and that selectively blocking motion along v_rev changes final alignment at T=100 from 0.648 +/- 0.009 to -0.211 +/- 0.021 while reducing harmfulness from 19.0% +/- 4.0% to 8.5% +/- 1.5% with little task cost.","tokens_in":2036,"tokens_out":378,"duration_ms":21811,"significance":"If the causal mediation result holds, the work supplies a concrete, history-defined direction that partially explains and controls early post-alignment reversion, offering a unifying lens on safety erosion, capability re-emergence, and related fragility. The controlled blocking intervention with reported error bars is a methodological strength that moves beyond pure correlation.","major_comments":[{"comment":"The blocking experiment (abstract and results section) establishes that v_rev is a causally relevant mediator, yet the manuscript does not report controls that orthogonalize v_rev against plausible alternatives such as random stable early directions, late-training principal components, or gradient subspaces at the same step. Without these, the specific attribution to a dominant early-training manifold (as opposed to other trajectory correlations) remains under-supported for the gravitational interpretation.","section":"results on blocking intervention"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed review and for highlighting the need for stronger controls on the blocking intervention. We address the major comment below and indicate where revisions will be made.","responses":[{"response":"We agree that the manuscript does not include explicit orthogonalization controls against random stable early directions, late-training principal components, or same-step gradient subspaces. v_rev is defined specifically from the observed representational drift in the early post-alignment phase and shows rapid, statistically elevated alignment with the history-defined direction; the blocking intervention then demonstrates a causal effect on both alignment and harmfulness. These elements provide support for v_rev as a mediator, but we acknowledge that without direct comparisons to the suggested alternative directions the attribution to a dominant early-training manifold (rather than other trajectory correlations) is not fully isolated. We will revise the results and discussion sections to add such controls where computationally feasible or, at minimum, to include an explicit limitations paragraph addressing this gap.","revision_made":"yes","referee_comment":"[results on blocking intervention] The blocking experiment (abstract and results section) establishes that v_rev is a causally relevant mediator, yet the manuscript does not report controls that orthogonalize v_rev against plausible alternatives such as random stable early directions, late-training principal components, or gradient subspaces at the same step. Without these, the specific attribution to a dominant early-training manifold (as opposed to other trajectory correlations) remains under-supported for the gravitational interpretation."}],"tokens_in":1469,"tokens_out":316,"duration_ms":19669,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's core result is that a direction v_rev built from early training history becomes aligned with later updates, and selectively blocking motion along it flips the final alignment score and cuts measured harmfulness roughly in half. That intervention is the clearest new piece.\n\nWhat stands out is the construction itself: v_rev is defined from the training trajectory rather than from the final model or from harm labels, and the authors then run a controlled block that changes a behavioral outcome while keeping task performance similar. The alignment numbers rising from 0.43 to 0.65 and sitting above the p99 isotropic null across 24 run-step pairs give a reproducible pattern in their setting. The error bars on the blocking condition (alignment 0.648 to -0.211, harm 19% to 8.5%) are reported and the effect size is large enough to notice.\n\nThe soft spot is that the experiment shows v_rev is causally relevant under their procedure, yet it does not yet isolate the claimed mechanism. Any direction that is stable early, lies in a high-variance subspace, or happens to align with later gradients could produce similar blocking effects. The abstract does not describe orthogonal controls against random early directions, late-training PCs, or gradient subspaces at the same step, so the \"gravitational\" reading remains one consistent story among several. Full methods and dataset details would also let a reader check whether the blocking step introduces its own confounds.\n\nThe work is aimed at researchers who track how fine-tuning trajectories interact with earlier training phases, especially in safety settings. Someone already thinking geometrically about optimization will find the v_rev construction and the intervention useful to test or extend. The paper shows clear thinking on its own terms and ships a falsifiable claim with numbers, so it is worth a serious referee even if the interpretation needs tightening.","headline":"The blocking experiment ties v_rev to reduced reversion in their runs, but the dominant-manifold story still needs controls that separate it from other stable trajectory directions.","tokens_in":2525,"tokens_out":447,"would_cite":false,"duration_ms":20251,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Early training creates dominant manifolds that later fine-tuning reverts toward along a history-defined direction v_rev.","keywords":["fine-tuning reversion","alignment","representational drift","behavioral manifolds","safety","machine learning","post-training dynamics"],"falsifier":"An experiment in which blocking motion along v_rev produces no reduction in harmfulness or in which measured alignments with v_rev remain within the isotropic null distribution would falsify the mediation claim.","tokens_in":2757,"feed_emoji":"🧭","tokens_out":682,"duration_ms":26470,"temperature":0.7,"pith_summary":"The paper proposes a geometric view of fine-tuning reversion in which large early training phases establish dominant behavioral manifolds while later alignment phases produce only shallow displacements. Subsequent updates therefore acquire a persistent component of motion back toward a witness of the early manifold, called v_rev. Representational drift is shown to align with this direction within the first 20 steps, and every observed alignment exceeds the 99th percentile of an isotropic null distribution. Selectively blocking motion along v_rev reverses the sign of final alignment and cuts measured harmfulness roughly in half with negligible task degradation. The results position v_rev as a causally relevant mediator of early post-alignment reversion in the reported settings.","feed_headline":"Blocking reversion direction halves model harmfulness","feed_subtitle":"Early training manifolds pull fine-tuned models back, but steering away from the defined direction preserves alignment at low cost.","key_machinery":"v_rev, a direction in activation space computed from training history that serves as a witness for the dominant early manifold and mediates reversion dynamics.","core_discovery":"Large early training phases establish dominant behavioral manifolds; subsequent fine-tuning inherits a reversion component along a history-defined direction v_rev that witnesses these manifolds, and selectively blocking motion along v_rev changes final alignment from 0.648 to -0.211 while reducing harmfulness from 19.0 percent to 8.5 percent.","pith_inferences":["The same history-defined direction might be recoverable in other post-training regimes such as continued pre-training or multi-task adaptation.","If v_rev can be estimated early, safety interventions could be applied at the start of any fine-tuning run rather than after reversion has occurred.","The manifold picture suggests that latent trait transfer through unrelated supervision may also be explainable by shared reversion components."],"forward_implications":["Representational drift rapidly acquires a component along v_rev, rising from cos 0.429 after the first update to 0.647 by step 20.","Across 24 run-step pairs every observed alignment with v_rev exceeds the p99 of an isotropic activation-space null.","Blocking motion along v_rev changes final alignment at T=100 from positive 0.648 to negative 0.211.","The same intervention reduces harmfulness from 19.0 percent to 8.5 percent while leaving task performance largely intact."],"fun_headline_variants":["Dominant early manifolds revert fine-tuned models","v_rev direction mediates alignment reversion","Blocking v_rev reduces harmfulness after fine-tuning","History-defined reversion explains safety fragility"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The observed increase in alignment with v_rev and the effect of blocking it are caused by the existence of a dominant early-training manifold rather than by other correlated properties of the optimization trajectory or activation geometry.","fun_headline_variants_meta":{"raw":{"variants":["Dominant early manifolds revert fine-tuned models","v_rev direction mediates alignment reversion","Blocking v_rev reduces harmfulness after fine-tuning","History-defined reversion explains safety fragility"]},"model":"grok-4.3","cost_usd":0.008004,"raw_usage":{"total_tokens":3683,"prompt_tokens":748,"num_sources_used":0,"completion_tokens":50,"cost_in_usd_ticks":80037000,"prompt_tokens_details":{"text_tokens":748,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2885,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":748,"tokens_out":50,"duration_ms":35427,"temperature":1.0,"reasoning_tokens":2885,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T00:53:50.619401+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An experiment in which blocking motion along v_rev produces no reduction in harmfulness or in which measured alignments with v_rev remain within the isotropic null distribution would falsify the mediation claim.","supporting_citations":[],"review_version":1}