{"id":"cea6f2fb-535e-4fe3-b968-189d3b4e0d9c","arxiv_id":"2504.21775","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"HetPFL adaptively samples performance-fairness preferences per client and fuses client hypernetworks preference-by-preference to learn better local and global Pareto fronts.","lead":"A new federated learning method, HetPFL, learns the full trade-off curve between model accuracy and fairness for each participant and for the global model by adapting how preferences are sampled and merging participant models by preference. The paper reports better trade-off curves than seven baselines on four datasets.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lemma 1's exact preference-alignment claim fails for preferences whose rays miss the Pareto front, and Appendix C.4 admits such 'green-area' preferences; the hypernet's front-parameterization and PSA's HVC logic rest on this premise.","rationale":"The central claim has two legs: a geometric front-learning construction and a convergence theorem. I think the geometric leg is the more load-bearing because it is used both by PSA (which optimizes over the preference simplex) and by the way Tables 1-4 measure front quality. The paper states Lemma 1 in a way that implies every sampled preference maps onto the front in the prescribed direction; the appendix then admits preferences with no front intersection, which is an internal contradiction. If Lemma 1 is interpreted only as standard Tchebycheff weak optimality, the exact-alignment passage and the green-area discussion cannot both be correct; if it is interpreted to require exact alignment, the appendix shows it is false. Either way the current text is not coherent. The empirical results might still be salvageable by reinterpreting PSA as learning the support of the front and discarding green-area samples, but that is a substantial revision of the claim. I agree with the reader that this is the weakest assumption; the conditional verdict is appropriate, with the requested revision being to replace Lemma 1 with a correct statement and to prove or remove the green-area Proposition 1. I also noted the nonvanishing constant in Theorem 1's bound, but did not select it as the primary issue because it affects the theoretical headline rather than the front-learning construction itself.","tokens_in":21677,"tokens_out":11417,"duration_ms":126533,"concrete_test":"Use the released code with the SYNTHETIC two-client configuration of Fig. 1/Fig. 9; after training, evaluate hβk on a dense grid of 1000 preferences covering the full simplex. For each λ, record the evaluated loss vector on D_k, compute the angle between that loss vector and the ray spanned by λ, and check whether the Tchebycheff minimizer has both scalarized objectives active. If a non-negligible fraction of the simplex (say more than 5%) shows nonzero angular error, or if the loss-vector ratio deviates from λ_CE/λ_F by more than a small tolerance, Lemma 1's exact-alignment statement is false and the preference-parameterized front claim must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Lemma 1 (Sec. 3, Eq. (3)) is the load-bearing geometric premise of HetPFL: it is used to claim that minimizing the weighted Tchebycheff loss for a sampled preference λ yields a weakly Pareto-optimal model whose loss vector 'aligns exactly with the direction of λ', so that sweeping λ with the hypernet traces the entire local/global Pareto front. That exact-alignment claim is not a consequence of the stated lemma. Exact alignment requires both Tchebycheff objectives to be active at the minimizer, i.e., the ray {t λ} must intersect the front. Appendix C.4 / Fig. 9 explicitly identifies a 'green area' of preferences that have no intersection with the front; for those λ the minimizer is an endpoint/weakly Pareto point whose loss-vector ratio generally differs from λ_CE/λ_F. The hypernet therefore does not encode the requested trade-off on a substantial part of the preference simplex, undermining (i) the interpretation of the learned local/global fronts, (ii) PSA's HVC-based preference-space optimization in Sec. 3.2, and (iii) the front-quality comparisons in Tables 1-4. The appendix's justification for green-area inefficiency cites 'Proposition 1', which does not exist in the paper, so the missing argument cannot repair the contradiction.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes HetPFL, a federated learning framework for learning performance-fairness Pareto fronts. It combines Preference Sampling Adaptation (PSA), which adapts each client's preference sampling distribution via a hypervolume contribution (HVC) criterion, with Preference-aware Hypernet Fusion (PHF), which aggregates client hypernets at the server in a preference-dependent way. The authors claim an error convergence rate of order O(1/t) for the hypernet under weaker assumptions than prior work, and report hypervolume improvements over seven baselines on four datasets.","tokens_in":21927,"tokens_out":7266,"duration_ms":75418,"significance":"The framework is clearly motivated, and the empirical results, if confirmed, would be a useful extension of PraFFL: PSA and PHF are concrete, well-ablated components, and the code release supports reproducibility. The paper also makes a good-faith attempt to connect to bi-level optimization theory. However, the central geometric claim, that sweeping the preference simplex traces the Pareto front, is only valid on a restricted subset of preferences, and the convergence theorem both assumes strong convexity that the objective does not satisfy and analyzes an idealized version of the algorithm. These issues must be resolved before the paper's main claims are supported.","major_comments":[{"comment":"The sentence 'Lemma 1 guarantees that when hβk(λ) is optimal for problem (3), the loss vector of hβk(λ) on dataset Dk aligns exactly with the direction of the preference vector and lies on the Pareto front' is not a consequence of the stated Lemma 1. A minimizer of the weighted Tchebycheff scalarization is weakly Pareto optimal, but its loss vector coincides with the ray {tλ} only when both objectives are active at the minimizer, i.e., when that ray actually intersects the Pareto front. Appendix C.4 and Fig. 9 explicitly identify a 'green area' of preferences that have no intersection with the front; for those preferences the minimizer is an endpoint/weakly Pareto point whose loss-vector ratio generally differs from λ_CE/λ_F. Since PSA's HVC-based adaptation in Sec. 3.2 and the interpretation of the learned front in Sec. 4 rest on this alignment premise, the paper needs either to restrict all front-parameterization claims to the preferences whose rays intersect the front, or to revise the method and claims accordingly. In addition, the appendix's citation of 'Proposition 1' to justify green-area inefficiency is missing from the paper and cannot repair the contradiction.","section":null},{"comment":"Assumption 3 requires g_tch(αk,βk) = E[max_j ℓ_j/λ_j] to be µ1-strongly convex in βk. This is not satisfied by the hypernet parameterization used in the experiments: a max of convex losses is convex but typically not strongly convex, and for neural-network outputs the losses are not convex in βk at all. Strong convexity is used to obtain the contractions in Lemma 3 and Eq. (31), so Theorem 1 does not apply to the actual HetPFL objective. Please either prove convergence under convexity or Polyak-Łojasiewicz-type assumptions that the objective can satisfy, or state clearly that Theorem 1 concerns a regularized or simplified problem that is not the one evaluated.","section":null},{"comment":"The claimed convergence rate is not supported by the stated bound. Equation (18) contains the non-vanishing term (σ1²µ1 + c1²L²q1 + G3²µ1)/µ1³, and the text explicitly says that Δt+1βk converges to this constant as t→∞; the algorithm therefore reaches a neighborhood of the optimum, not the optimum itself. Moreover, with a constant step size (1−ηtζ)^{t/4} is exponentially decaying, while with ηt = 1/t it tends to a constant, so neither case yields an O(1/t) rate from the displayed bound. Please clarify the step-size schedule, correct the rate claim, and distinguish convergence to the optimum from convergence to a biased neighborhood.","section":null},{"comment":"The convergence analysis does not cover the NES estimator used in the implemented algorithm. Equation (13) estimates ∇α E[−HVC] via the score-function gradient, but Assumptions 1-4 and the proof treat ∇α g_hvc as an exact gradient; no bias or variance bound for the NES estimator is given. Additionally, Algorithm 1 samples one preference vector per inner iteration while Eqs. (12)-(13) require a batch of N vectors (N=4 in Appendix C.1), so the pseudocode and the update rules do not match. Please either analyze the estimator actually used, or explicitly state that Theorem 1 applies to an idealized exact-gradient version of HetPFL.","section":null}],"minor_comments":[{"comment":"Table 1 reports only averaged values over three runs, with no standard deviations or significance tests; the abstract's and Sec. 4.2's claim that HetPFL 'significantly outperforms' the baselines is therefore not statistically supported. Please add error bars or statistical tests.","section":"Table 1, Sec. 4.2"},{"comment":"The notation Λαk is used both for the sampling distribution p(αk) and for the set of N sampled preference vectors; please disambiguate these two objects.","section":"Sec. 3.2, Eq. (13)"},{"comment":"The text cites 'Proposition 1' to justify the inefficiency of sampling preferences from the green area, but no Proposition 1 appears anywhere in the paper; please add the missing statement and proof, or remove the citation.","section":"Appendix C.4"},{"comment":"The definitions of z1 and z2 are garbled: z1 is written without closing parentheses and includes a square-root term that is not clearly grouped, and z2 appears to be missing a closing parenthesis. Please rewrite this step of the proof cleanly.","section":"Appendix B, around Eq. (31)"},{"comment":"On BANK, PSA alone decreases global HV from 0.895 to 0.886, while the text says that 'similar patterns can be observed' across datasets; please qualify this claim or explain the exception.","section":"Table 4, Ablation study"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope, and the self-citation to PraFFL is appropriate given that PraFFL is the direct predecessor. The main risk is overclaiming: the geometric lemma and the convergence-rate statement need substantial correction before publication, and the empirical evidence would be strengthened by error bars. I do not see a citation-ethics or novelty-disclosure concern beyond what is stated above."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one if you care about multi-objective fairness in FL. HetPFL extends PraFFL with two new pieces: PSA adapts each client's preference sampling distribution via hypervolume contribution, and PHF aggregates hypernets on the server with a preference-dependent FusionNet. Both are sensible, and the ablations show each adds value. Experiments cover four datasets, seven baselines, with public code; gains are modest but consistent (roughly 2% local, 5.5% global hypervolume). That part is solid.\n\nThe soft spots are in the theory and presentation. Theorem 1 assumes strong convexity of a Tchebycheff max-loss (Assumption 3). Real neural networks are neither strongly convex nor smooth, and the NES estimator used for sampling-distribution gradients is absent from the analysis. So the O(1/t) result should be read as a stylized bound, not a faithful description of what the algorithm does. This is common in FL theory, but the paper sells it as a main contribution.\n\nThe bigger issue is geometric. Lemma 1 is a standard Tchebycheff result, but the paper's next sentence claims the loss vector aligns exactly with the preference direction. That only holds when the ray along λ intersects the Pareto front. Appendix C.4's Figure 9 admits a 'green area' of preferences with no intersection, and the appendix then cites a 'Proposition 1' that does not exist. The PSA's HVC logic and the interpretation of the learned fronts depend on this alignment claim, so the theoretical scaffolding is incomplete. The method might still work empirically—the ablations suggest it does—but the written justification needs repair.\n\nMinor: Table 1 reports three runs but no error bars; Lemma 2's probability expression looks garbled. These are cosmetic but should be fixed.\n\nVerdict: worth a serious referee. The empirical direction is positive, the method is new enough, and the flaws are fixable: state Lemma 1's conditions precisely, add or delete the missing proposition, add error bars, and temper the theory claims. I'd bring it to reading group, though it's not a paradigm shift.","headline":"HetPFL is a useful, incremental extension of Pareto-front learning to fair federated learning with genuine empirical gains, but its theory section overclaims and its geometric Lemma 1 needs repair before the paper's own logic holds together.","tokens_in":22507,"tokens_out":3595,"would_cite":true,"duration_ms":37671,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HetPFL learns both local and global performance-fairness Pareto fronts in federated learning, with a one-over-time convergence rate.","keywords":["federated learning","group fairness","Pareto front","hypernetwork","preference sampling","hypervolume contribution","performance-fairness trade-off"],"falsifier":"Take SYNTHETIC, where the true Pareto front can be computed exhaustively; train HetPFL, sample 1,000 preferences uniformly over the simplex, and measure the cosine alignment between each model's $(\\ell_{\\mathrm{CE}}, \\ell_F)$ loss vector and its preference $\\lambda$, plus the distance to the true front. Lemma 1 predicts perfect alignment and front membership for every preference, so if a substantial fraction of green-area preferences map to points off their ray or off the front, the central discovery claim collapses.","tokens_in":21409,"feed_emoji":"⚖️","tokens_out":9138,"duration_ms":87909,"temperature":0.7,"pith_summary":"This paper attempts to show that the performance-fairness trade-off in federated learning can be represented as a whole family of models, one for each preference between accuracy and group fairness, rather than a single compromise model. HetPFL does this with a hypernetwork on each client that maps a preference vector to a preference-specific model, plus an adaptive preference sampler (PSA) and a server-side preference-aware hypernet fusion (PHF). The authors claim that uniform preference sampling wastes training on heterogeneous local fronts and that ignoring the global dataset leaves the aggregated model suboptimal for many preferences; their two mechanisms are meant to fix exactly those gaps. They prove a hypernetwork error bound with dominant term $(1-\\eta_t\\zeta)^{t/4}$, giving an overall $O(1/t)$ convergence rate, and report hypervolume gains of about 1.75% on local fronts and 5.5% on the global front over the strongest baseline on four datasets.","feed_headline":"One federated model learns every fairness-performance trade-off point","feed_subtitle":"Adaptive preference sampling plus hypernet fusion learns both local and global trade-off fronts on 4 datasets.","key_machinery":"The machinery has three moving parts. First, a hypernetwork $h_{\\beta_k}:\\lambda\\mapsto\\theta_k$ generates a model for each preference vector $\\lambda$ in the two-dimensional simplex; training it minimizes the weighted Tchebycheff scalarization $g_{\\mathrm{tch}}=\\max_{j\\in\\{\\mathrm{CE},\\mathrm{F}\\}} \\ell_j/\\lambda_j$, which by Lemma 1 makes the model weakly Pareto optimal with loss vector aligned to $\\lambda$. Second, PSA lets each client model its preference sampler as a Dirichlet distribution $p(\\alpha_k)$ and updates it using the gradient of negative hypervolume contribution, estimated by Natural Evolution Strategies, inside a bi-level optimization coupled with the hypernet update. Third, PHF introduces FusionNet $W_\\phi(\\lambda)$, which maps a preference to nonnegative fusion weights over the clients' hypernets; the server optimizes $\\phi$ so the combined global hypernet minimizes the same scalarized loss on shared latent features. The convergence proof combines Lemma 2 on FedAvg's linear rate, Lemma 3 on one-round bi-level contraction, and four Lipschitz/strong-convexity/bounded-gradient assumptions, yielding Theorem 1's error bound.","core_discovery":"The paper's central claim is that a hypernetwork trained with the weighted Tchebycheff scalarization can learn the entire performance-fairness Pareto front of each client, provided the preference sampling distribution is adapted to where that client's front actually lies, and that the global front can be improved by learning preference-dependent fusion weights for the clients' hypernetworks. Formally, Theorem 1 bounds the squared distance from the learned hypernetwork to the optimal one at round $t+1$ by $(3/4)^{\\tau_p t} \\Delta^0_{\\beta_k} + z_1(1-\\eta_t\\zeta)^{t/4}\\sqrt{\\Delta^0_{\\psi_k}} + z_2(1-\\eta_t\\zeta)^{t/2}\\Delta^0_{\\psi_k}$ plus an optimization-error term, whose dominant time-dependent term gives $O(1/t)$ convergence. The experimental claim is that on SYNTHETIC, COMPAS, BANK, and ADULT, HetPFL beats seven baselines in hypervolume for both local and global fronts, and that its preference sampling is denser where the local front actually lies.","pith_inferences":["Editorial inference: HVC-based adaptation is not specific to fairness; the same preference-sampling-plus-hypernetwork loop should port to other two-objective federated problems, such as accuracy versus energy or accuracy versus robustness, because Tchebycheff scalarization and hypervolume contribution only see the loss vector.","Editorial inference: the 'green area' discussion in Appendix C.4 raises a risk the paper does not address: an HVC-driven sampler could collapse all mass onto a subset of the Pareto front, leaving other preference rays uncovered; adding a coverage or entropy regularizer on $p(\\alpha_k)$ would be a natural extension to test.","Editorial inference: one could test PHF against a cheaper alternative, a per-preference weighted average of client models rather than client hypernets, to see whether hypernet-space fusion is essential or whether the gain comes merely from preference-dependent weighting."],"forward_implications":["Once trained, HetPFL can answer 'what model do I get if I weight accuracy at 0.8 and fairness at 0.2?' at inference time, without retraining, so a single run sweeps out as many trade-off points as needed.","The per-client sampling adaptation is designed to keep front quality when data are strongly heterogeneous, and the heterogeneity and client-count experiments (10 to 300 clients) support this scalability claim, where the closest prior global front degrades.","The global Pareto front becomes an explicit optimization target through PHF, so the aggregated model is no longer just a FedAvg average; the authors show global hypervolume improves by about 5.5% over the best baseline.","Under the paper's assumptions, front quality improves with communication rounds at order $O(1/t)$, matching the usual FedAvg rate, so adding the fairness trade-off machinery does not change the asymptotic communication cost."],"supporting_citations":[{"why":"Supplies the weighted Tchebycheff scalarization and Lemma 1, which guarantee weak Pareto optimality with loss vectors aligned to the preference.","marker":"[Miettinen, 1999]"},{"why":"Provides the two-timescale bi-level optimization framework and the one-round convergence lemma (Lemma 3) that the hypernet and sampling-distribution updates rely on.","marker":"[Hong et al., 2023]"},{"why":"Gives Lemma 2, the linear convergence of the communicated model under FedAvg, which enters the hypernet error bound.","marker":"[Collins et al., 2021]"},{"why":"Defines the FedAvg aggregation used in Phase I for the communicated model.","marker":"[McMahan et al., 2017]"},{"why":"PraFFL is the closest prior hypernet-based preference Pareto-front method, the main baseline and starting point HetPFL extends.","marker":"[Ye et al., 2025]"},{"why":"Defines the performance and group-fairness losses used in the objectives and provides the SYNTHETIC dataset setting.","marker":"[Zeng et al., 2021]"},{"why":"Defines hypervolume, the quality metric that is also the basis of the HVC indicator used for preference sampling adaptation.","marker":"[Zitzler and Thiele, 1999]"},{"why":"Supplies the evaluation convention (error rate and demographic-parity disparity) and the FairFed baseline.","marker":"[Ezzeldin et al., 2023]"}],"fun_headline_variants":["Adaptive sampling learns client-specific fairness-performance fronts","HetPFL: Adaptive sampling and hypernet fusion for Pareto fronts","Per-client preference sampling beats fixed sampling for fairness fronts","Adaptive hypernet fusion captures local and global fairness-performance fronts","Sampling adapted to each client's Pareto front improves federated fairness"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that every sampled preference produces a model whose loss vector lies exactly on the Pareto front along the preference's ray, yet the paper's Appendix C.4 shows whole 'green area' regions of the preference simplex whose rays never intersect the front, so the alignment premise fails there.","fun_headline_variants_meta":{"raw":{"variants":["Adaptive sampling learns client-specific fairness-performance fronts","HetPFL: Adaptive sampling and hypernet fusion for Pareto fronts","Per-client preference sampling beats fixed sampling for fairness fronts","Adaptive hypernet fusion captures local and global fairness-performance fronts","Sampling adapted to each client's Pareto front improves federated fairness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00101,"raw_usage":{"total_tokens":4288,"prompt_tokens":986,"completion_tokens":3302,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":602,"completion_tokens_details":{"reasoning_tokens":3219}},"tokens_in":602,"tokens_out":3302,"duration_ms":25518,"temperature":1.0,"reasoning_tokens":3219,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:54:39.223933+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take SYNTHETIC, where the true Pareto front can be computed exhaustively; train HetPFL, sample 1,000 preferences uniformly over the simplex, and measure the cosine alignment between each model's $(\\ell_{\\mathrm{CE}}, \\ell_F)$ loss vector and its preference $\\lambda$, plus the distance to the true front. Lemma 1 predicts perfect alignment and front membership for every preference, so if a substantial fraction of green-area preferences map to points off their ray or off the front, the central discovery claim collapses.","supporting_citations":[{"cited_title":"Nonlinear multiobjective optimization, volume","cited_arxiv_id":null,"evidence_quote":"Supplies the weighted Tchebycheff scalarization and Lemma 1, which guarantee weak Pareto optimality with loss vectors aligned to the preference."},{"cited_title":"A two-timescale stochastic al- gorithm framework for bilevel optimization: Complexity analysis and application to actor-critic","cited_arxiv_id":null,"evidence_quote":"Provides the two-timescale bi-level optimization framework and the one-round convergence lemma (Lemma 3) that the hypernet and sampling-distribution updates rely on."},{"cited_title":"Exploiting shared rep- resentations for personalized federated learning","cited_arxiv_id":null,"evidence_quote":"Gives Lemma 2, the linear convergence of the communicated model under FedAvg, which enters the hypernet error bound."},{"cited_title":"Multiobjective evolutionary algorithms: a comparative case study and the strength pareto approach","cited_arxiv_id":null,"evidence_quote":"Defines hypervolume, the quality metric that is also the basis of the HVC indicator used for preference sampling adaptation."},{"cited_title":"Fairfed: Enabling group fairness in federated learning","cited_arxiv_id":null,"evidence_quote":"Supplies the evaluation convention (error rate and demographic-parity disparity) and the FairFed baseline."}],"review_version":1}