{"id":"b970d489-6515-4221-b37a-4674a42252b5","arxiv_id":"2608.07157","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Update-based estimates of client data heterogeneity in sub-model federated learning are dominated by device capacity, and adaptive allocation adds nothing over a matched-budget random control once parameter coverage is guaranteed.","lead":"This paper shows that in federated learning where clients train smaller versions of a shared model, the standard way to estimate which clients have unusual data is actually measuring device capacity instead, and it can't be fixed by clever math. It also proves that when no client is allowed to train the whole model, the ignored parts of the model slowly break it.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Missing equal-capacity positive control leaves the central 'capacity confound' interpretation undemonstrated; the estimator may simply be an uninformative data signal.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing gap: the paper never runs the equal-capacity positive control that would distinguish a capacity confound from an estimator that is simply uninformative about data heterogeneity. This is the most important issue because the paper's central theoretical contribution is the confound interpretation, not merely the observation that update divergence correlates with capacity. The reported partial correlations in Table 5 are also inconsistent with the abstract's claim that 'no data signal remains once capacity is controlled for'; values as large as -0.60 indicate a remaining, possibly negative, relationship. A secondary concern is that the claimed second 'corrected estimator' is never defined, but that is secondary because the missing control is logically prior. The existing CONDITIONAL verdict already requires addressing this gap, so my stress-test does not move the verdict. If the positive control fails to show a substantial data signal under equal capacities, the conclusion would need to be weakened to 'update divergence is not a reliable data heterogeneity signal under sub-model training, with capacity variation adding a negative bias' rather than the stronger 'capacity confounds an otherwise valid estimator.'","tokens_in":23236,"tokens_out":3201,"duration_ms":31916,"concrete_test":"Run the identical HAS-FL training protocol on the same Dirichlet partitions and seeds with capacities fixed equal for all clients (p_max_i = 1.0, and a second run with fixed p_i = 0.5 for all clients), then compute r(H, TV) between the smoothed estimates and ground-truth label divergence. If r(H, TV) is substantially positive (say above 0.5), the capacity-confound interpretation is supported. If it remains near zero or negative, the headline finding must be revised: update divergence does not recover data heterogeneity even without capacity variation, and capacity is not demonstrated as a confound.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that capacity variation, not data, drives update-divergence estimates, and that controlling for capacity removes the data signal. Section 4.4 validates the estimator against ground-truth label-divergence only in the system-heterogeneous setting, where p_max varies across clients (0.25 to 1.0). No experiment runs the same estimator with capacities held equal. Without such a positive control, the strong negative r(H, p_max) values (-0.72 to -0.84) are consistent with two very different explanations: (a) capacity is a confound that masks a valid data signal, or (b) the estimator is simply a poor or negatively biased proxy for label divergence whenever sub-models are trained, independent of capacity variation. The paper's own partial correlations argue against explanation (a): r(H, TV | p_max) is -0.60, -0.47, -0.40, -0.26, -0.28, and 0.10 across runs, not 'no data signal remains.' The abstract's stronger claim is therefore not supported by the reported numbers, and the title-level conclusion that capacity confounds an otherwise valid signal requires the missing control.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses HAS-FL, an adaptive sub-model federated-learning framework, as a test case for a broader question: can a server estimate client data heterogeneity from sub-model updates when client model width varies with device capacity? The authors report three findings: (i) update-divergence estimates are dominated by capacity rather than by label-distribution divergence, based on correlations with ground-truth TV distance and capacity constraints in Table 5; (ii) capped allocation can leave parameters frozen at initialization, formalized as Proposition 3, and a coverage guarantee removes the resulting collapse; (iii) a matched-budget random-allocation control matches or outperforms the heterogeneity-aware policy on three benchmarks. The paper concludes that the apparent benefits of adaptive sub-model allocation come from capacity budgeting and parameter coverage rather than from heterogeneity estimation.","tokens_in":23398,"tokens_out":4775,"duration_ms":49963,"significance":"If fully established, the capacity-confound result is an important negative result: it would invalidate a natural and repeatedly suggested design direction in system-heterogeneous federated learning. The paper has real strengths: Proposition 3 is a direct and machine-checkable proof, the estimator is validated against ground-truth label divergence computed from actual partitions, experiments are multi-seed and reproducible with released code, and the matched-budget random control is a strong and appropriate comparison. However, the paper's strongest claim is currently overstated relative to its own numbers, and the confound interpretation would be substantially strengthened by an equal-capacity positive control that the paper does not run.","major_comments":[{"comment":"The abstract and Section 4.4 claim that 'no data signal remains once capacity is controlled for,' but Table 5 reports partial correlations r(H, TV | p_max) of -0.60, -0.47, -0.40, -0.26, -0.28, and 0.10 across the six runs. These values are not all near zero, and two are substantial in magnitude; only one is positive. The evidence supports 'no consistent positive data signal' or 'a weak and inconsistent data signal,' not 'no data signal remains.' Similarly, the text's statement that r(p, TV | p_max) is approximately zero is contradicted by the -0.36 and -0.46 entries in the same table. These are the paper's headline quantitative claims, so the wording and conclusions should be revised to match the reported numbers.","section":"Abstract and Section 4.4 (Table 5)"},{"comment":"The central 'confound' interpretation is missing a positive control: an experiment in which client capacities are held equal across clients while label divergence still varies, using the same estimator. Without such a baseline, the strong negative correlations between H and p_max are equally consistent with the alternative explanation that the proposed update-divergence estimator is simply a poor or negatively biased proxy for label divergence under sub-model training, independent of capacity variation. Since Section 4.7 explicitly frames the paper as a warning about capacity changing the meaning of update-based signals, the paper should either add the equal-capacity control or explicitly limit the claim to the system-heterogeneous setting.","section":"Section 4.4"},{"comment":"The abstract and Section 4.4 claim that the confound persists 'across two corrected estimators,' and the text says 'we verified that both the coordinate-restricted and the common-core variants of the estimator exhibit the coupling.' However, only one estimator is defined (Eq. (4)), and Table 5 reports a single set of correlations. The common-core variant is never defined and its results are not reported. Either define the second estimator and give its correlations, or revise the claim to refer to one estimator.","section":"Abstract and Section 4.4"}],"minor_comments":[{"comment":"The text states that HeteroFL provides 'full coverage by construction,' but with 10 of 20 clients sampled per round, a static policy cannot guarantee per-round coverage unless a full-width client is sampled in every round. If 'full coverage' is meant in the union-of-masks sense over the whole federation, please state this explicitly; otherwise clarify how HeteroFL avoids rounds in which no sampled client covers the outer channels.","section":"Table 4 and Section 4.3"},{"comment":"The allocation rule depends on free parameters p_min, gamma, beta, T_adapt, and T_norm, with values p_min=0.4 and gamma=0.25 chosen without a sensitivity analysis. Since the central negative result concerns the estimator rather than the allocation rule, this is not blocking, but a sentence on the robustness of the conclusions to these choices would help.","section":"Section 4.1 and Eq. (6)"},{"comment":"The seed labels s42 through s44 are unusual and the paper does not explain why seeds are not numbered 1-3. Please clarify, since the reader cannot tell whether these are three separate initializations or a subset of a larger set.","section":"Table 5"},{"comment":"The Shakespeare benchmark uses only 20 role-clients, and the paper acknowledges this limitation. It would be useful to state explicitly whether the three seeds correspond to different train/test splits of the same roles or to different roles altogether, since this affects the interpretation of the reported standard deviations.","section":"Section 4.6 and Table 7"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be publishable after revision if the authors add the equal-capacity positive control or carefully restrict the claim, and if they align the abstract and Section 4.4 with the actual partial correlations. The missing control is specifically fixable within the manuscript's scope, so major revision seems appropriate rather than rejection. The direct proof and reproducible code are genuine strengths."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a useful, mostly careful negative result that overshoots its own evidence in the abstract. The empirical pattern is real—update-divergence estimates correlate strongly with device capacity across seeds and datasets, and partial correlations with ground-truth label divergence are weak after controlling for capacity. But the claim that \"no data signal remains\" is not what Table 5 shows (partial rs down to -0.60), and the paper never runs the positive control that would justify the word \"confound\": an equal-capacity federation where the estimator is shown to track label divergence. Without that control, the identification is incomplete. What you can conclude from the reported experiments is that this estimator is not informative in the system-heterogeneous regime; you cannot conclude that capacity masks an otherwise valid signal. That is a meaningful distinction, and the stress-test note has it right.\n\nWhat is genuinely new and good: the matched-budget random control is a clean methodological addition, and the coverage-guarantee ablation plus Proposition 3 turns a known failure mode into a precise structural statement. The Shakespeare result—adaptivity being actively harmful—is a nice counterpoint. The paper also ships code, which makes the experiments reproducible.\n\nSoft spots, in order of severity. First, the missing equal-capacity control, as above; this is the main thing a referee should demand. Second, the paper refers to \"two corrected estimators\" but defines only one; the common-core variant is mentioned but never specified. Third, the abstract and conclusion repeat the \"no data signal remains\" line despite the paper's own partial correlations being as large as -0.60 in magnitude. That is an overstatement, not a fatal flaw, but it should be fixed.\n\nOverall: for people working on sub-model FL or server-side heterogeneity estimation, this is worth reading and will likely become a standard cautionary citation, especially for the coverage result and the matched-budget protocol. It needs one additional experiment before the central claim is bulletproof, but the paper is serious, well-contextualized, and deserves a proper peer review with a request for that control.","headline":"Solid negative result with a load-bearing overclaim: the capacity-confound interpretation needs an equal-capacity positive control that the paper never runs.","tokens_in":23976,"tokens_out":2346,"would_cite":true,"duration_ms":22944,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Sub-model federated learning cannot recover client data heterogeneity from update divergence once device capacities differ: across two corrected estimators, multiple datasets, and all seeds, the estimates are dominated by capacity…","keywords":["federated learning","sub-model training","system heterogeneity","statistical heterogeneity","capacity confound","update divergence","heterogeneity estimation","coverage guarantee"],"falsifier":"Run the same federated training protocol with all clients at equal full capacity while keeping the same ground-truth label-divergence partition, and correlate the update-divergence estimates with total-variation divergence across clients; a clearly positive correlation would support the capacity-confound interpretation, whereas a near-zero or negative one would show the estimator fails to measure data heterogeneity even without capacity variation, and the confound framing would need revision.","tokens_in":22982,"feed_emoji":"⚙️","tokens_out":11609,"duration_ms":85158,"temperature":0.7,"pith_summary":"This paper asks whether a federated server can infer a client's data heterogeneity from the model updates it receives when clients train width-reduced sub-models sized to their device capacity. Using the adaptive allocation framework HAS-FL as a test case, it claims the answer is no: validated against ground-truth label-distribution divergence on reproducible partitions, the update-divergence estimates correlate strongly and negatively with device capacity on every dataset and seed, and no consistent data signal survives once capacity is controlled for. The paper further claims that the apparent benefits of adaptive allocation come from parameter coverage and budget discipline rather than heterogeneity awareness: a coverage guarantee that keeps the largest-capacity clients at full width removes a freeze-at-initialization failure mode, and random allocation to the same average budget performs no differently on the image benchmarks and better on a naturally partitioned text benchmark. If right, these results warn that any method estimating client statistics from sub-model updates in a system-heterogeneous federation is measuring capacity, not data.","feed_headline":"Sub-model FL measures device capacity, not data skew","feed_subtitle":"Update-divergence estimates track capacity (r ≈ -0.8), so coverage, not allocation intelligence, protects accuracy.","key_machinery":"The central object is the normalized gradient divergence estimator $\\hat{H}^{(t)}_i = \\|m^{(t)}_i \\odot (\\Delta^{(t)}_i - \\bar{\\Delta}^{(t)})\\|^2 / (\\|m^{(t)}_i \\odot \\bar{\\Delta}^{(t)}\\|^2 + \\epsilon)$, computed on the coordinates client $i$ trained, smoothed by an exponential moving average, and fed into the allocation rule $p^{(t)}_i = \\min(p^{\\max}_i, p_{\\min} + \\gamma \\tilde{H}^{(t)}_i / (\\bar{H}^{(t)} + \\epsilon))$. The load-bearing pieces are the nested width masks inherited from HeteroFL-style sub-model extraction, the coverage guarantee that restores full width to the largest-capacity clients, and Proposition 3, which proves that under capped allocation the uncovered coordinates stay frozen at initialization. Together they show that sub-model updates differ across width tiers by an order of magnitude in size and norm, dwarfing any data-driven variation; that frozen output-path parameters are what make uniform allocation collapse; and that coverage plus budget, not heterogeneity targeting, carries the accuracy.","core_discovery":"The paper's central claim is that update-based estimators of client heterogeneity are confounded by capacity whenever sub-model width varies with device capability. In the system-heterogeneous setting, the smoothed normalized gradient divergence $\\tilde{H}^{(t)}_i$ correlates with device capacity at $r$ between $-0.84$ and $-0.72$ across two corrected estimators, two datasets, and all seeds, while partial correlations with ground-truth label divergence after removing capacity are near zero or negative, so the allocations do not track true heterogeneity. A second claim is structural: capped allocation freezes uncovered parameters, since coordinates outside the widest trained slice remain at random initialization for the entire run while participating in every forward pass (Proposition 3), progressively corrupting the global model, and a coverage guarantee that assigns full width to the highest-capacity clients eliminates the degradation. A third claim, established by a matched-budget control, is that adaptive allocation contributes nothing beyond its capacity budget: random time-varying allocation at the same average capacity matches HAS-FL on CIFAR-10 and EMNIST and beats it on Shakespeare, where the adaptive policy uses the most capacity and achieves the lowest accuracy.","pith_inferences":["Likely generalization beyond the paper's measurements: the confound should afflict any update statistic, including norms, cosine similarity, and direction agreement, because the mechanism is the order-of-magnitude gap in update size and norm across width tiers, so equal-model-size clustering and client-selection methods remain safe only when every client trains the identical architecture.","Testable extension: an equal-capacity federation with the same label-divergence partition would separate the two readings of the paper's negative result, with a positive correlation there completing the confound story and a null result indicating divergence-based heterogeneity signals should be abandoned outright.","Natural next step the paper names but does not pursue: estimating data statistics within capacity strata, using fixed-width probe batches, or having clients report statistics directly, with the paper's reproducible-partition partial-correlation protocol as the evaluation harness.","Design rule suggested by the coverage analysis: any federated aggregation should guarantee that every parameter participating in inference is eventually updated by some client, with priority on output-path parameters, since frozen interior units degrade less than frozen classifier coordinates."],"forward_implications":["Any method that estimates client data statistics from sub-model updates in a system-heterogeneous federation is measuring capacity rather than data, so heterogeneity-aware allocation built on such estimates cannot work as intended.","Methods that size or shape sub-models from training-derived signals, such as capability-driven pruning ratios, data-driven channel importance, magnitude-based composition, or learned sparse ratios, inherit the same confound and should validate against a capacity-stratified analysis.","Parameter coverage is the decisive design element in sub-model federated learning: preserving coverage of output-path parameters prevents the freeze-at-initialization collapse that destroys uniform allocation on high-class-count tasks.","Adaptive allocation schemes should be reported against a matched-budget random control, which in these experiments fully explained the benefits attributed to the adaptive policy.","Sub-model training still buys the participation of resource-constrained devices at quadratically reduced compute and communication cost, but at a genuine accuracy cost of roughly 8 to 13 points against full-model training."],"supporting_citations":[{"why":"Supplies the nested width-mask sub-model construction and channel-wise aggregation that HAS-FL's masks and coordinate-restricted divergence estimator inherit.","marker":"Diao et al., 2021"},{"why":"Provides the ordered-dropout sub-model scheme used as the random-tier allocation baseline in the main comparisons.","marker":"Horvath et al., 2021"},{"why":"Documents the prior observation that parameters not selected by any client remain unupdated, which the coverage analysis builds on.","marker":"Alam et al., 2022"},{"why":"Defines the FedAvg protocol and serves as the full-model baseline whose accuracy sets the cost of sub-model training.","marker":"McMahan et al., 2017"},{"why":"Supplies the LEAF Shakespeare construction whose natural role-based partitions ground the third benchmark.","marker":"Caldas et al., 2018"},{"why":"Full-model control-variate baseline, and the convergence-bound setting in which gradient divergence is theoretically meaningful.","marker":"Karimireddy et al., 2020"},{"why":"Full-model proximal baseline used to contextualize the sub-model accuracy-capacity trade-off.","marker":"Li et al., 2020b"},{"why":"Provides the CIFAR-10 dataset used for the primary image benchmark.","marker":"Krizhevsky, 2009"}],"fun_headline_variants":["Sub-model FL: capacity, not data, drives update estimates","Adaptive sub-model FL: allocation smarts don't beat random","Coverage, not allocation intelligence, protects sub-model FL","Capacity confound dooms adaptive sub-model allocation","Sub-model FL: update divergence tracks device capacity, not data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the update-divergence estimator would recover data heterogeneity if device capacities were equal, a positive control the paper never runs, leaving open that the estimator is simply a weak data signal in general.","fun_headline_variants_meta":{"raw":{"variants":["Sub-model FL: capacity, not data, drives update estimates","Adaptive sub-model FL: allocation smarts don't beat random","Coverage, not allocation intelligence, protects sub-model FL","Capacity confound dooms adaptive sub-model allocation","Sub-model FL: update divergence tracks device capacity, not data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000232,"raw_usage":{"total_tokens":1552,"prompt_tokens":1070,"completion_tokens":482,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":686,"completion_tokens_details":{"reasoning_tokens":399}},"tokens_in":686,"tokens_out":482,"duration_ms":4841,"temperature":1.0,"reasoning_tokens":399,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:49:11.284939+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same federated training protocol with all clients at equal full capacity while keeping the same ground-truth label-divergence partition, and correlate the update-divergence estimates with total-variation divergence across clients; a clearly positive correlation would support the capacity-confound interpretation, whereas a near-zero or negative one would show the estimator fails to measure data heterogeneity even without capacity variation, and the confound framing would need revision.","supporting_citations":[],"review_version":1}