{"id":"27a76099-a7a2-4fc6-9634-76ff00ffa377","arxiv_id":"2505.01788","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"On a single Virus-MNIST benchmark, adding homomorphic encryption to the APPLE personalized federated learning algorithm produced the highest reported accuracy among four privacy variants.","lead":"This paper benchmarks privacy-preserving federated personalized learning on the Virus-MNIST dataset with 200 clients. It reports that APPLE+HE gives the best accuracy (99.34%) while APPLE+DP is the fastest, and recommends APPLE+HE for privacy-preserving personalized machine learning.","discovery_kind":"incremental","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The APPLE+HE recommendation cannot be audited: implementation and privacy parameters are undisclosed, no error bars are reported, and the paper contradicts itself (99.34% in Table III vs 99.37% in the text), so the ranking may be an artifact.","rationale":"I read the paper's central claim as an empirical recommendation: APPLE+HE outperforms APPLE+DP, APPLE+SA, and APPLE+SMPC on all metrics. For that claim to hold, the benchmark must be reproducible and the four variants must be compared under equivalent conditions. The manuscript does not provide the material needed to verify this: no code, no HE parameters, no DP privacy budget, no error bars, and no repeated trials. The internal inconsistency between 99.34% (Table III) and 99.37% (text) further undermines confidence in the reported numbers. The observed margin between first and second place is small enough that a configuration difference, rather than a genuine algorithmic advantage, could explain it. The reader's weakest assumption identifies the same load-bearing issue, and the reader's REJECT verdict is consistent with my reading. No additional objection is needed, and no manufactured concern is warranted.","tokens_in":12265,"tokens_out":2874,"duration_ms":31553,"concrete_test":"Run one check: obtain the authors' code and configuration and reproduce Table III on Virus-MNIST with the same 200-client deployment, reporting per-seed mean and standard deviation for all four variants, along with the exact HE scheme, DP epsilon, and SA/SMPC protocol. If the APPLE+HE margin over APPLE+DP and APPLE+SA collapses within two standard deviations, or if the 99.34/99.37 discrepancy reappears, the headline recommendation is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is an empirical ranking, so it depends on a fair, auditable comparison. That condition is not met. In Section IV-B the four variants are specified only at a conceptual level (Eqs. 1-4): no HE scheme or parameter set, no DP noise mechanism or epsilon/delta values, no SA/SMPC protocol or security model, and no number of clients or rounds beyond the mention of 200 clients in Section IV-A. Section V-C then reports a single benchmark in Table III with no variance or repeated seeds. The paper is internally inconsistent: the text in Section V-C states APPLE+HE achieves 99.37% accuracy, while Table III reports 99.34%. It also reports APPLE+DP at 97.48%, slightly above plain APPLE (97.41%, Table II), which would require the DP noise to be negligible or the training budgets to differ; either way the comparison is not controlled. Because the margin of APPLE+HE over APPLE+DP and APPLE+SA is only about 1.9 percentage points, any of these confounds, such as a small privacy budget, a different test split, or different hyperparameters, could reverse the stated ordering. The reader's weakest assumption is exactly right: the ranking presupposes correct, equivalent implementations, and the manuscript supplies no evidence for that.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports an empirical comparison of federated personalized learning (FPL) algorithms on the Virus-MNIST dataset. It first evaluates existing FPL methods (APFL, APPLE, Ditto, FedALA, FedFomo, and others) and then combines the best performer, APPLE, with differential privacy (DP), homomorphic encryption (HE), secure aggregation (SA), and secure multi-party computation (SMPC). It reports that APPLE+HE achieves the highest accuracy, precision, recall, and F1-score (99.34% in Table III, stated as 99.37% in the text), recommends APPLE+HE, and notes that APPLE+DP offers more efficient execution in terms of server clock running time.","tokens_in":12704,"tokens_out":4342,"duration_ms":46094,"significance":"If the comparison were correctly implemented and fully documented, the benchmark could be a useful reference for practitioners choosing among privacy-preserving personalization techniques. The paper's broad coverage of the FPL literature and its clear definition of standard evaluation metrics are strengths. However, the central contribution is a purely empirical ranking, and the manuscript provides no code, no disclosure of privacy or encryption parameters, no uncertainty quantification, and only a single dataset. The ranking is therefore not auditable, the comparison is not controlled, and the internal accuracy inconsistency directly undermines the headline claim. The fact that the paper ships no reproducible artifacts means the claimed contribution cannot be verified from the manuscript alone.","major_comments":[{"comment":"The description of the four privacy-preserving variants is conceptual only, relying on Eqs. (1)-(4). No differential privacy mechanism (Laplace vs. Gaussian), privacy budget epsilon/delta, homomorphic encryption scheme or security parameter, secure aggregation protocol, or SMPC threat model is disclosed. Because the central claim is an empirical ranking of APPLE+HE over APPLE+DP, APPLE+SA, and APPLE+SMPC, these missing implementation details are load-bearing: without them the experiments cannot be reproduced or audited, and the ordering could change under different parameter choices.","section":"IV-B and V-C"},{"comment":"The headline ranking rests on a single accuracy value per algorithm, with no confidence intervals, standard deviations, or repeated-seed results, and no statistical test. The margin between APPLE+HE (99.34%) and APPLE+DP (97.48%) or APPLE+SA (97.44%) is only about 1.9 percentage points, which is well within the range of seed-to-seed variability for deep federated training. In addition, the text in Section V-C reports 99.37% for APPLE+HE while Table III reports 99.34% for the same quantity, so the paper is internally inconsistent.","section":"V-C, Table III"},{"comment":"The comparison is not controlled: APPLE+DP is reported at 97.48% accuracy, which is higher than the plain APPLE model's 97.41% in Table II. For non-negligible differential privacy noise this is unexpected, and the more likely explanation is that the DP noise was negligible or that the training budgets, test splits, or hyperparameters differed across the variants. The manuscript does not state the data partition, the number of communication rounds, or the hyperparameters used for each PPMLFPL variant, so a reader cannot determine whether all four algorithms were evaluated under identical conditions.","section":"IV-A, V-B, V-C"},{"comment":"The execution-time claim is also under-specified. 'Server clock running time' is not defined, no hardware details are given, no repeated measurements or error bars are reported, and the times in Table IV increase almost perfectly linearly with client count for every algorithm, an artifact-like pattern that is not discussed. Furthermore, Section V-C concedes that 'future work will consider other factors, such as further computational complexity, scalability, and the specific requirements of the application, before conclusively determining the best approach,' which is in direct tension with the paper's earlier recommendation of APPLE+HE.","section":"Table IV and V-C"}],"minor_comments":[{"comment":"The heading 'Evaluation Matrices' should read 'Evaluation Metrics.'","section":"V-A"},{"comment":"Equation (2) writes the differential privacy condition as '△GlobalModel ≤ ε,' which is not a correct statement of epsilon-differential privacy; the inequality should be corrected or removed.","section":"IV-B"},{"comment":"Table I is garbled in the provided text, with misaligned columns and entries, making the literature summary difficult to read.","section":"Table I"},{"comment":"There are numerous typos and grammatical errors, including 'approache,' 'Matrices Perspective,' and 'creating it' in the abstract; the manuscript needs careful proofreading.","section":"Abstract and V-C"},{"comment":"Figures 5 and 6 are described as having rows and columns of accuracy or loss values, but the axes and color encoding are not explained, so as printed the figures are not informative.","section":"Figures 5 and 6"},{"comment":"Several references are duplicated or incomplete; for example, references [6] and [8] are the same homomorphic-encryption book, and the Tune reference [44] cites a preprint rather than the standard platform documentation.","section":"References"}],"recommendation":"reject","confidential_remarks":"As submitted, the manuscript is not suitable for publication. The central claim is an empirical ranking that cannot be audited because no implementation, parameters, or code are disclosed, and the reported numbers are internally inconsistent. These are load-bearing problems, not local presentation issues. I see no path to acceptance without a complete rewrite of the experimental section and the addition of reproducibility artifacts."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nShort version: this paper is a parameter scan over existing FPL algorithms and privacy primitives on Virus-MNIST. The only new thing is APPLE+HE, which is just the existing APPLE algorithm wrapped in homomorphic encryption. That is a routine combination, not a new mechanism. The paper's recommendation to use APPLE+HE is built on a benchmark that cannot be checked from the manuscript.\n\nWhat the paper does well: the literature review is broad, and the authors do run 16 FPL algorithms and then compare four privacy wrappers on the same dataset. If the numbers were trustworthy, it would give practitioners a data point. Using Ray tune for hyperparameter search is reasonable.\n\nThe soft spots are load-bearing. The comparison is missing exactly the details that decide the ranking: no DP noise scale or epsilon, no HE scheme or parameters, no SA/SMPC protocol details, no number of rounds, and no indication that all variants got the same training budget. APPLE+HE beats APPLE+DP and APPLE+SA by about 1.9 percentage points, so any of those missing choices could invert the ordering. There are no error bars or repeated runs. The paper contradicts itself: the text says APPLE+HE reached 99.37%, while Table III says 99.34%. Table IV shows near-linear execution times that look too clean (e.g., APPLE+SMPC: 31,123 / 36,128 / 41,133 / 56,143 ms), and the paper never discusses that pattern. Also, APPLE was selected as the best on the same benchmark that later supports APPLE+HE, which is a selection-on-the-test-set issue.\n\nOverall, there is no way to verify the central claim. This is not a serious contribution as it stands; it is a rough draft. It could be a starting point for a careful study, but only if the authors disclose parameters, release code, and run repeated trials.\n\nMy recommendation: desk reject. It does not deserve referee time in its current form, and the novelty is too low to justify a major revision. If the authors can provide code and full experimental details, a different paper could be worth a look.\n\nBest,\n[Your name]","headline":"A benchmark comparison that cannot be audited: no privacy parameters, no error bars, no code, and an internal inconsistency in the headline number.","tokens_in":13075,"tokens_out":3191,"would_cite":false,"duration_ms":31975,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This study reports that APPLE+HE, the homomorphic-encryption variant of the APPLE personalized federated learning algorithm, achieved the highest accuracy, precision, recall, and F1 score among four privacy-preserving variants on the…","keywords":["federated personalized learning","privacy preserving machine learning","model personalization","homomorphic encryption","differential privacy","secure aggregation","secure multi-party computation","Virus-MNIST"],"falsifier":"Re-run the four APPLE variants on the same Virus-MNIST test split with the same 200-client setup and the same disclosed privacy parameters (for example, a specific epsilon and delta for APPLE+DP and a named HE scheme with its parameters); if APPLE+DP achieves accuracy close to or above 99.34% at a comparable or smaller privacy loss, the paper's claim that APPLE+HE is the best-performing approach would not hold.","tokens_in":12117,"feed_emoji":"🔐","tokens_out":3577,"duration_ms":35366,"temperature":0.7,"pith_summary":"The paper evaluates four ways to add privacy protection to a personalized federated learning algorithm called APPLE, and it argues that combining APPLE with homomorphic encryption (APPLE+HE) gives the best predictive performance. On the Virus-MNIST dataset with 200 clients, APPLE+HE is reported to reach 99.34% accuracy, precision, recall, and F1-score, beating the differential privacy, secure aggregation, and secure multi-party computation variants. The authors also report that the differential privacy variant has the lowest server clock running time, while the SMPC variant has the highest. The paper recommends APPLE+HE for privacy-preserving personalized federated learning tasks, while acknowledging that efficiency, scalability, and application-specific requirements still need to be weighed.","feed_headline":"Homomorphic encryption beats DP and SMPC in federated personalization","feed_subtitle":"The homomorphic encryption variant hits 99.34% accuracy on Virus-MNIST; differential privacy runs fastest.","key_machinery":"The central object is the APPLE algorithm (Adaptive Personalized Cross-Silo Federated Learning), a personalized federated learning method that adapts local aggregation for each client. The paper combines APPLE with four privacy-preserving primitives: homomorphic encryption (HE), differential privacy (DP), secure aggregation (SA), and secure multi-party computation (SMPC). Homomorphic encryption allows the server to sum encrypted local model updates directly, so the aggregation step operates on ciphertexts and decryption reveals only the aggregated result, not individual contributions. This property is what the paper credits for APPLE+HE's high accuracy, while the other techniques trade away accuracy or speed through noise, masking, or cryptographic overhead.","core_discovery":"The paper's central claim is that APPLE+HE outperformed the three competing approaches (APPLE+SMPC, APPLE+DP, and APPLE+SA) on every reported performance metric. In particular, APPLE+HE achieved 99.34% accuracy, precision, recall, and F1-score, compared with 97.48%, 97.44%, and 85.38% for APPLE+DP, APPLE+SA, and APPLE+SMPC respectively. The authors interpret this as evidence that homomorphic encryption can protect individual model updates during aggregation without degrading model utility, and they conclude that the results strongly support APPLE+HE as the recommended algorithm for privacy-preserving machine learning in federated personalized settings, with APPLE+DP offering the most efficient execution.","pith_inferences":["A single dataset, Virus-MNIST, carries the entire comparison; the ordering of APPLE+HE over APPLE+DP may not transfer to other data distributions, model architectures, or client counts.","The paper does not disclose the DP noise scale (for example, the epsilon and delta budgets) or the HE scheme and parameter settings, so the reported margins could shift if those privacy parameters were chosen differently.","A natural testable extension would be to vary the privacy budget continuously and plot the accuracy-versus-privacy frontier; the paper's headline ranking might invert at small epsilon values where DP noise is large.","In practice, homomorphic encryption usually imposes substantial communication and computation costs on real workloads, so the paper's recommendation of APPLE+HE should be read as utility-focused rather than efficiency-focused."],"forward_implications":["If APPLE+HE genuinely maintains near-perfect accuracy on Virus-MNIST while preserving privacy, it suggests that homomorphic encryption can be a practical choice for personalized federated learning without a utility penalty.","The large gap between APPLE+HE (99.34%) and APPLE+SMPC (85.38%) implies that the choice of cryptographic protocol materially affects model quality, not just runtime, on this benchmark.","APPLE+DP's lower execution times indicate that differential privacy remains competitive when computational efficiency is the priority, even if its accuracy is slightly below APPLE+HE.","The results position APPLE+HE as a candidate default for privacy-preserving personalized federated learning in settings where accuracy is paramount and clients can tolerate the added latency described in the paper.","The paper's conclusion that APPLE+HE should be recommended is conditional on further evaluation of computational complexity, scalability, and application requirements, which it explicitly defers to future work."],"supporting_citations":[{"why":"Supplies the APPLE algorithm, the personalized federated learning base that all four privacy-preserving variants build on.","marker":"[30]"},{"why":"Defines differential privacy, the noise-adding mechanism underlying the APPLE+DP variant.","marker":"[5]"},{"why":"Introduces homomorphic encryption, the cryptographic primitive used in the recommended APPLE+HE variant.","marker":"[6]"},{"why":"Lays out secure multi-party computation, which forms the basis of the APPLE+SMPC implementation.","marker":"[7]"},{"why":"Describes practical secure aggregation, the protocol behind the APPLE+SA variant.","marker":"[3]"},{"why":"Provides the federated learning setup and the data preprocessing procedures the authors say they followed for the Virus-MNIST experiments.","marker":"[1]"},{"why":"Ray Tune is the hyperparameter optimization tool used to tune the neural network configurations in the experimental setup.","marker":"[44]"}],"fun_headline_variants":["Homomorphic encryption tops accuracy in federated personalization","HE beats DP and SMPC on accuracy, DP runs fastest","Federated learning: HE scores 99.34% accuracy, DP fastest","Privacy-preserving ML: HE wins accuracy, DP wins speed","In federated personalization, HE hits 99.34% while DP is quickest"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported ranking of APPLE+HE over APPLE+DP, APPLE+SA, and APPLE+SMPC assumes all four variants were implemented correctly, given identical training budgets, hyperparameters, and test splits, yet the paper does not disclose the DP noise scale, the HE scheme, or any implementation code.","fun_headline_variants_meta":{"raw":{"variants":["Homomorphic encryption tops accuracy in federated personalization","HE beats DP and SMPC on accuracy, DP runs fastest","Federated learning: HE scores 99.34% accuracy, DP fastest","Privacy-preserving ML: HE wins accuracy, DP wins speed","In federated personalization, HE hits 99.34% while DP is quickest"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000928,"raw_usage":{"total_tokens":3967,"prompt_tokens":928,"completion_tokens":3039,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":2945}},"tokens_in":544,"tokens_out":3039,"duration_ms":17790,"temperature":1.0,"reasoning_tokens":2945,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:09:50.170150+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the four APPLE variants on the same Virus-MNIST test split with the same 200-client setup and the same disclosed privacy parameters (for example, a specific epsilon and delta for APPLE+DP and a named HE scheme with its parameters); if APPLE+DP achieves accuracy close to or above 99.34% at a comparable or smaller privacy loss, the paper's claim that APPLE+HE is the best-performing approach would not hold.","supporting_citations":[{"cited_title":"Adapt to adaptation: Learning personalization for cross-silo federated learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the APPLE algorithm, the personalized federated learning base that all four privacy-preserving variants build on."},{"cited_title":"Differential privacy,","cited_arxiv_id":null,"evidence_quote":"Defines differential privacy, the noise-adding mechanism underlying the APPLE+DP variant."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces homomorphic encryption, the cryptographic primitive used in the recommended APPLE+HE variant."},{"cited_title":"Secure multi -party computation,","cited_arxiv_id":null,"evidence_quote":"Lays out secure multi-party computation, which forms the basis of the APPLE+SMPC implementation."},{"cited_title":"Communication-efficient learning of deep networks from decentralized data,","cited_arxiv_id":null,"evidence_quote":"Provides the federated learning setup and the data preprocessing procedures the authors say they followed for the Virus-MNIST experiments."}],"review_version":1}