{"id":"73642187-1fc6-44f5-ac71-5a505a7b6afe","arxiv_id":"2608.09490","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Functional interference between task vectors is real but conditional: it persists across scales and model families only for coarse, input- and format-specific comparisons, not for benchmark predictions.","lead":"This paper tests when adding two task vectors in a language model's weights behaves predictably. It finds the result depends heavily on the input prompt and its format, so task arithmetic is not a universal merging rule.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing per-adapter task accuracy leaves the semantic identity of the task vectors unverified: the central code+safety vs. code+math contrast could reflect adapter quality rather than task-vector interference.","rationale":"The reader's weakest assumption already identifies the same load-bearing concern: the interaction metric is defined in logit/first-token space, and the paper does not report task accuracy for individual adapters, so the semantic labels are unverified. I agree with that diagnosis and add that the missing same-task and random-direction anchors are the natural control for the same issue. This is not a demonstrated flaw; it is a missing verification that conditions the central interpretation. The paper is unusually transparent about the limitation, and the estimator audits, norm-matching controls, prospective specifications, and cross-scale/cross-architecture stress tests are strong. The exact OOS p-value of 0.083 is explicitly bounded by the design and is treated as directional evidence, and the pass@1 failure is presented as a validity boundary rather than a contradiction. Therefore the appropriate verdict remains CONDITIONAL, unchanged from the reader: the central conditional claim is plausible and well-supported except for one verifiable condition, namely that the task labels correspond to real, non-degenerate capabilities and that the contrast is not reproducible with random norm-matched directions. If the proposed per-adapter evaluation resolves that condition, the paper should be acceptable; if not, the central semantic conclusion would need to be weakened.","tokens_in":16281,"tokens_out":5348,"duration_ms":67775,"concrete_test":"Release per-adapter capability evaluations: for each seed and base model, report held-out task accuracy for the base model and each single adapter on its intended task (GSM8K for math, HumanEval+/MBPP+ pass@1 for code, XSTest safe/unsafe behavior for safety, AlpacaEval for instruction), plus a same-task/same-seed control and a norm-matched random-direction anchor evaluated through Eq. (5) on the same prompt strata. Check that (i) each adapter beats the base model on its target metric by a meaningful margin, and (ii) the code+safety minus code+math contrast for random norm-matched directions is centered near zero. If the safety adapter is near chance or random directions reproduce the contrast, the semantic interpretation of the central claim is not supported; if both checks pass, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central positive result is the input-conditioned non-additivity contrast R(code+safety) - R(code+math), interpreted as task-vector interference. Equation (5) defines R on first-token JSDs, but the labels 'code', 'math', and 'safety' are attached to adapters solely by their training data. The paper states in the Discussion: 'We do not report task accuracy for the individual adapters, so each task label describes training data rather than verified capability.' This leaves two unseparated explanations for the contrast: task-vector semantics, or idiosyncratic adapter quality and direction. Because code+safety and code+math share the same code checkpoint within each seed, code-side quality is controlled; the uncontrolled side is the safety adapter versus the math adapter. If the safety adapter is degenerate, poorly calibrated, or merely learns a generic format shift, the large R for code+safety on code/instruction prompts could be a quality artifact rather than an interaction between 'code' and 'safety' directions. The acknowledged absence of same-task and random-direction anchors (Expanded Methods) reinforces the issue: without a norm-matched random-direction baseline, we cannot tell whether R is specific to trained task vectors or reflects generic properties of norm-matched displacements. This is load-bearing because the prospective OOS ordering and the cross-scale/architecture persistence all inherit the task labels; if the labels are not validated, the claimed conditional generalization across task pairs is unsupported. The first-token-only limitation and the pass@1 boundary are explicitly acknowledged and are not the main vulnerability; the missing semantic verification is.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies functional non-additivity of task-vector composition by defining an input-conditioned interaction ratio R (Eq. 5) that compares the observed merged logit distribution against a logit-space additive counterfactual built from the two axis paths, normalized by displacement from the base. On Qwen2.5-1.5B with norm-matched LoRA vectors, the paper reports that code+safety has higher R than code+math on code and instruction prompts but not math prompts; a prospectively specified six-task expansion yields 8/8 high-vs-low sign agreements; the contrast persists under 0.5B full fine-tuning, Qwen2.5-7B LoRA, and Llama-3.1-8B; and public raw prompts preserve but an instruction wrapper collapses the code-prompt contrast, while EvalPlus pass@1 interactions do not robustly reproduce it. The paper concludes that weight-space composition supports coarse, input- and format-conditioned functional statements but is not a universal merging-performance predictor.","tokens_in":16535,"tokens_out":9624,"duration_ms":113920,"significance":"The result, if it holds, is a useful boundary result for task arithmetic: it separates parameter geometry from functional geometry and provides a concrete measurement surface rather than a fitted merge-quality predictor. The methodology has notable strengths: matched checkpoints within seed, core median norm matching, frozen prospective specifications with dated summaries, a cached estimator audit that reproduces the frozen bootstrap values, explicit denominator-boundary checks, a design-preserving exact randomization test, and a claim-boundary checklist that distinguishes supported from unsupported conclusions. The paper is honest about its limits, including the absence of per-adapter task accuracy and the failure of the metric to predict pass@1. Its significance is therefore conditional and primarily negative/descriptive: it maps where composition behaves non-additively rather than offering a new merging algorithm.","major_comments":[{"comment":"The absence of any per-adapter capability validation is load-bearing for the interpretation of the central code+safety vs. code+math contrast as task-vector interference. Because the code checkpoint is shared within each seed, the uncontrolled comparison is between the safety adapter and the math adapter; if the safety adapter is degenerate, miscalibrated, or dominated by a generic refusal/format shift, the elevated R on code and instruction prompts could reflect properties of that adapter's output distribution rather than an interaction between task directions. This concern propagates to the OOS ordering and to every stress-test row, which inherit the same task labels. The Discussion explicitly acknowledges \"We do not report task accuracy for the individual adapters,\" but the acknowledgement does not resolve the ambiguity. I request per-adapter validation (for example, held-out loss or task accuracy for each adapter on its own evaluation split, plus perhaps a norm-matched random-direction control) or, alternatively, a systematic reframing of all task-level conclusions as statements about data-conditional fine-tuning displacements rather than about \"tasks.\"","section":"Discussion and Limitations; Eq. (5)"},{"comment":"The paper's own limitation statement notes the absence of same-task and random-direction anchors. Without a norm-matched random-direction baseline, the paper cannot distinguish task-vector-specific interference from generic properties of any large norm-matched displacement in the model's weight space. The central claim that \"task-vector interference is measurable\" requires at least one such control; the existing comparisons between task pairs are informative but do not establish that the effect is specific to task vectors. I recommend adding a random-direction condition (e.g., Gaussian or permuted directions matched in norm) or softening the conclusion accordingly.","section":"Expanded Methods; Claim Boundary Checklist"}],"minor_comments":[{"comment":"R is a dimensionless ratio, but the text reports it in \"percentage points\" and \"percentages\" (e.g., Table 3 caption). Please state explicitly that all reported values are 100 times R.","section":"Results, Tables 2-5 and Eq. (5)"},{"comment":"The section heading overstates the 1.5B finding, where the clustering diagnostics are significant (cosine gap p=0.0002); the heading applies to the 0.5B block. Consider rewording to \"Persistence without Detectable Clustering at 0.5B.\"","section":"Functional Structure Persists Without Detectable Parameter Clustering"},{"comment":"The sentence \"Accuracy for comparisons between high and low bins is 37.5% on math prompts, 50% on code prompts...\" uses \"accuracy\" for a sign-agreement rate; rename to avoid confusion with model task accuracy, especially given the paper's own caution about missing task accuracy.","section":"Prospective Directional Test on Unseen Pairs"},{"comment":"The sentence \"Table 7 shows that the absolute numerator and denominator vary substantially..., so the main result is not inferred from a contrast alone\" is unclear; the displayed variation does not by itself support that conclusion, and the sentence could be removed or rewritten.","section":"Estimator and Denominator Audit"},{"comment":"The abstract's \"all eight ... predicted sign\" should be accompanied by the exact design-preserving p-value (0.083) or a qualifier such as \"directionally, though not statistically definitive,\" to prevent overinterpretation by casual readers.","section":"Abstract and Prospective Directional Test"}],"recommendation":"major_revision","confidential_remarks":"The paper is methodologically careful and unusually honest about its boundaries. The main risk is that the central semantic interpretation rests on unvalidated task labels; if the authors add adapter-level validation and a random-direction control, I would be comfortable with acceptance. There is also a fit question: the paper is more of a diagnostic/negative result than a new method, but it should be of interest to the task-arithmetic and model-merging community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is the rare task-arithmetic paper that treats the input distribution as part of the estimand rather than an afterthought. Second, its main claim—interference is conditional, not global—holds up, but the semantic identity of the task adapters is never directly verified, and that leaves the headline contrast one step from fully landing.\n\nWhat's new: an input-conditioned functional interaction surface built from logit-space additive counterfactuals, with norm-matched controls and prospective transfer tests. The paper ships unusually disciplined empirics: a cached estimator audit, frozen protocols, three seeds, hierarchical bootstrap intervals, and honest reporting of what is descriptive versus confirmatory. The primary contrast (code+safety more non-additive than code+math on code and instruction prompts but not math) persists across full fine-tuning at 0.5B, Qwen2.5 scale to 7B, and a Llama-3.1-8B audit. The external boundary is interesting: raw public prompts preserve the contrast, a wrapper collapses it, and pass@1 does not robustly reproduce it. That is a substantive validity boundary, not a failure.\n\nThe soft spots are real but mostly acknowledged. The Discussion states that task labels describe training data, not verified capability; no adapter task accuracy is reported. Since the code checkpoint is shared within seed, the uncontrolled side is the safety adapter versus the math adapter. If the safety adapter is degenerate or just learns a format shift, the contrast could partly reflect quality rather than task-vector semantics. The missing same-task and random-direction anchors make absolute-scale readings unsafe; the authors say so. The prospective 8/8 OOS sign match has an exact design-preserving p of 0.083, and the finer ordering (5/8 high versus middle) is unreliable. Those limits are in the paper.\n\nThe paper does not overclaim. It explicitly warns against treating the ratio as a universal merge-performance predictor. For a model-merging researcher, or anyone deciding when to trust weight-space composition, it is a useful boundary map.\n\nI would send this to peer review. The missing adapter accuracy should be the first thing a reviewer asks for; adding even a few held-out accuracy numbers per adapter plus a random-direction baseline would close the main gap. As is, it is a solid, honest empirical contribution that deserves serious referee time.","headline":"A careful empirical mapping of task-vector interference: input- and format-conditioned, honestly bounded, but the adapters' semantic identity is left unverified.","tokens_in":17085,"tokens_out":2409,"would_cite":true,"duration_ms":28669,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Task-vector interference is measurable and input-dependent, not a universal semantic.","keywords":["task arithmetic","task vectors","model merging","functional interference","weight-space composition","LoRA","input-conditioned geometry","validity boundaries"],"falsifier":"Recompute the central code+safety versus code+math contrast with a different operationalization of functional interaction-for example, full-sequence likelihood or fine-grained task accuracy on code and instruction benchmarks-using the same seeds and norm-matched endpoints; if the direction reverses or vanishes on code and instruction prompts, the claimed boundary is an artifact of the first-token JSD ratio.","tokens_in":1583,"feed_emoji":"🧩","tokens_out":6770,"duration_ms":94546,"temperature":0.7,"pith_summary":"This paper asks when adding fine-tuning task vectors in weight space produces predictable changes in model behavior. It separates parameter geometry from functional geometry and measures pairwise non-additivity over a two-task composition surface, using a first-token logit-space interaction ratio normalized by base displacement. On Qwen2.5-1.5B, code+safety is more non-additive than code+math on code and instruction prompts but not on math prompts. A prospectively specified six-task expansion yields all eight predicted high-versus-low signs for unseen task pairs, and the contrast persists under full fine-tuning, larger scales, and a second model family. The conclusion is that weight-space composition supports coarse, input- and format-conditioned functional statements, not a universal merging-performance predictor.","feed_headline":"Task-vector merging works only under input and format conditions","feed_subtitle":"The same adapter pair interferes on code prompts but not math prompts; a wrapper template erases the effect.","key_machinery":"The central object is the input-conditioned interaction ratio $R^{(s)}_{a,b}(X)$ from Eq. (5): the ratio of the expected first-token Jensen-Shannon divergence between the observed composed distribution and an additive counterfactual, to the expected divergence from the base. The counterfactual is built in logit space as $\\ell_{\\alpha,0}+\\ell_{0,\\beta}-\\ell_{0,0}$, which removes the marginal nonlinearity of each axis before measuring their joint effect. Norm-matched controls equalize vector radii across tasks and seeds, prospective bins and frozen success rules provide transfer tests on unseen pairs, and a hierarchical bootstrap over seeds and prompts supplies uncertainty intervals.","core_discovery":"The central claim is that functional interference between task vectors is jointly determined by the task pair and the input distribution, and that this interference can be measured by a first-token interaction ratio $R^{(s)}_{a,b}(X)$: the ratio of Jensen-Shannon divergence between the actual merged distribution and an additive logit-space counterfactual, to the divergence from the base. The paper establishes that the code+safety pair exceeds the code+math pair by 5.96 and 7.58 percentage points on code and instruction prompts but only 0.10 on math prompts; all eight prospective high-versus-low comparisons on unseen task pairs have the predicted sign; and the ordering persists under full-parameter fine-tuning at 0.5B, Qwen2.5 scales up to 7B, and a Llama-3.1-8B cross-architecture audit. The same signal collapses when identical public code prompts are wrapped in an instruction template and is not robustly reproduced by EvalPlus pass@1. Therefore task-vector interference is conditionally generalizable, not a global semantic coordinate system in weight space.","pith_inferences":["The same ratio machinery could be applied layer-wise or at multiple decoding steps to localize where interference enters the network, which the paper names as the natural next test.","The results suggest that a safety adapter is not a semantically clean axis: the code+safety contrast may partly reflect training-data characteristics or refusal behavior rather than a stable safety direction.","If the boundary is correct, task-arithmetic methods that optimize a single global coefficient set should be re-evaluated per prompt stratum, since a fixed coefficient set cannot be optimal across formats.","A testable extension is to use the interaction ratio R as a screening score on many task pairs and compare its ranking with measured merge accuracy on a fixed prompt distribution; the paper's pass@1 result suggests the ranking may fail on discrete metrics."],"forward_implications":["If a practitioner needs to predict whether two adapters interfere, the input prompt distribution and its serialization must be specified; there is no input-free answer.","Parameter-space summaries such as cosine similarity or clustering are not reliable proxies for functional interference: at 0.5B the functional contrast persists without detectable task clustering.","The code+safety versus code+math hierarchy transfers across rank-16 LoRA, full fine-tuning at 0.5B, Qwen2.5 scaling to 7B, and Llama-3.1-8B, so the coarse ordering is robust to adaptation method, scale, and one additional model family.","Continuous first-token interaction does not predict discrete benchmark merge performance: EvalPlus pass@1 interactions are wide and inconsistent, so evaluation on intended prompts remains necessary.","Because prompt format flips the result on identical prompts, wrapper templates are part of the conditioning distribution rather than a neutral serialization."],"supporting_citations":[{"why":"Defines task arithmetic, the additive weight-space operation whose validity boundaries this paper maps.","marker":"Ilharco et al. 2023"},{"why":"LoRA is the primary adaptation parameterization used for the controlled experiments.","marker":"Hu et al. 2022"},{"why":"TIES-Merging is a representative interference-reduction method that the paper contrasts with plain additive composition.","marker":"Yadav et al. 2023"},{"why":"Closest prior work showing parameter-space angles are not a stable predictor of adapter interference, motivating the input-conditioned functional measure.","marker":"Sivaramakrishnan et al. 2026"},{"why":"Closest prior work predicting pairwise merge performance from checkpoint properties, which the paper distinguishes from its non-additivity target.","marker":"Zhou et al. 2026"},{"why":"EvalPlus pass@1 evaluation is used to establish the discrete-behavior validity boundary.","marker":"Liu et al. 2023"},{"why":"PKU-SafeRLHF data defines the safety adapter used in the central comparisons; the paper does not claim verified safety capability.","marker":"Ji et al. 2024"},{"why":"Llama-3.1-8B is the base model for the cross-architecture audit.","marker":"Dubey et al. 2024"}],"fun_headline_variants":["Input distribution decides when task vectors interfere","Task-vector merging hinges on prompt format and type","No universal merge predictor: task vectors are conditional","Task arithmetic works only for coarse, conditioned statements"],"cache_read_input_tokens":19200,"weakest_assumption_plain":"The measurement rests on the assumption that the first-token interaction ratio R, computed in logit space and normalized by base displacement, captures the meaningful functional interference between task vectors; the paper does not report task accuracy for adapters, so task labels describe training data rather than verified capabilities.","fun_headline_variants_meta":{"raw":{"variants":["Input distribution decides when task vectors interfere","Task-vector merging hinges on prompt format and type","No universal merge predictor: task vectors are conditional","Task arithmetic works only for coarse, conditioned statements"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000256,"raw_usage":{"total_tokens":1616,"prompt_tokens":1026,"completion_tokens":590,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":642,"completion_tokens_details":{"reasoning_tokens":532}},"tokens_in":642,"tokens_out":590,"duration_ms":7123,"temperature":1.0,"reasoning_tokens":532,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:29:52.071729+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the central code+safety versus code+math contrast with a different operationalization of functional interaction-for example, full-sequence likelihood or fine-grained task accuracy on code and instruction benchmarks-using the same seeds and norm-matched endpoints; if the direction reverses or vanishes on code and instruction prompts, the claimed boundary is an artifact of the first-token JSD ratio.","supporting_citations":[{"cited_title":"Demystifying Mergeability: Interpretable Properties to Predict Model Merging Success , journal =","cited_arxiv_id":null,"evidence_quote":"Closest prior work predicting pairwise merge performance from checkpoint properties, which the paper distinguishes from its non-additivity target."}],"review_version":1}