REVIEW 4 major objections 5 minor 57 references
fluke: Federated Learning Utility frameworK for Experimentation and research
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Adding a new federated learning algorithm in fluke requires defining only the client and/or server classes.
desk verdict fluke is a clean, genuinely useful FL prototyping package with a solid algorithm zoo, but the 'minimal overhead' claim is unquantified and the paper's own tutorial shows you need more than just client/server classes. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing structure is the class hierarchy in the fluke package: fluke.algorithms.CentralizedFL (or PersonalizedFL) as the entry point, plus fluke.client.Client and fluke.server.Server, whose defaults implement a FedAvg-style protocol. A Channel class simulates the server–client communication link and logs overhead, while YAML configuration files separate experiment settings (dataset, distribution, rounds, seed) from algorithm settings (model, hyperparameters). This combination lets a novel algorithm be expressed as an override of fit, send_model, receive_model, aggregate, and related methods, with the surrounding simulation handled by the base classes.
What would settle it
Run the same algorithm in fluke and in its original implementation under identical seeds, data partitions, and hyperparameters on MNIST and compare the accuracy curves; if the curves or final values diverge substantially, the framework's benchmarking validity collapses. A cheaper check is to re-run the Figure 2 configuration with several seeds and see whether the reported ordering among methods is stable.
Extended reading notes
Core claim
The paper's central claim is that the design of fluke reduces the work of adding a federated learning algorithm to the definition of two kinds of classes. The Client class owns local training and model exchange; the Server class owns client selection, broadcasting, and aggregation. A third class, the algorithm entry point, mostly declares which client and server classes to use. The framework simulates communication through a Channel object, and all other machinery, including data loading, non-IID splits, model definitions, evaluation, and logging, is provided. The central design bet is that FL algorithms are best expressed server-side and client-side, so researchers can focus on the learning components rather than on infrastructure.
Load-bearing premise
The load-bearing premise is that fluke's bundled implementations of published FL algorithms faithfully reproduce the original methods, so that default-settings runs yield valid comparisons; the paper does not check this against original sources.
Editorial extensions
If this is right
- A researcher can benchmark a new idea against several existing FL algorithms by writing one experiment config and swapping algorithm config files, without modifying code between runs.
- Publishing a fluke-based algorithm gives other researchers a runnable implementation plus configuration, which supports the paper's reproducibility goal.
- The built-in datasets and non-IID distribution functions remove the need to implement data partitioning, so different papers can compare under the same data conditions.
- Because only client and server classes change, a new algorithm can be tested while keeping the same evaluation, logging, and channel-overhead tracking as the baselines.
Reading between the lines
- Counting the overridden methods needed to implement a published variant, such as a proximal-term or extrapolation FL method, would provide a quantitative check of the minimal overhead claim.
- A companion validation study that compares each bundled baseline against its original implementation would make fluke a reference benchmark rather than a convenience; the current paper does not include that check.
- The experiment and algorithm configuration split suggests a simple batch-comparison workflow that could be extended to auto-generate benchmark tables for papers.
- The framework's current single-machine, synchronous design limits its use to algorithm-level questions; topology and communication studies would need the multi-device support listed as future work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces fluke, an open-source Python package for simulating centralized federated learning (FL) experiments on a single machine. The author argues that existing FL frameworks are too rigid or have too steep a learning curve for researchers who want to prototype new algorithms, and claims that fluke can be used out of the box for benchmarking and extended with new algorithms by defining client and/or server classes with minimal overhead. The paper describes the package architecture, the command-line interface, a tutorial for adding a new algorithm, and presents an example benchmark of several FL algorithms on MNIST. The manuscript also states that fluke explicitly does not aim to compete with full-featured frameworks and is scoped to a synchronous, centralized, simulated setting.
Significance. If the central usability claims are correct, fluke could be a valuable tool for FL researchers by lowering the barrier to prototyping and improving reproducibility of baseline comparisons. The paper ships open-source code, documentation, tutorials, and a set of pre-implemented algorithms and datasets, which are concrete and inspectable artifacts. However, the significance assessment is conditional because the paper does not quantify the claimed 'minimal overhead,' does not validate the faithfulness of its bundled algorithm implementations, and does not compare against existing frameworks such as Flower. These omissions leave the central value proposition unverified, although the design is plausible and the modular architecture appears sensible. The paper is likely useful as a software announcement, but its scientific claims require additional evidence.
major comments (4)
- [Implementing a new algorithm] The paper's central claim, stated in the introduction and features list, is that implementing a new algorithm 'simply requires the definition of the client and/or the server classes.' This is contradicted by the tutorial in the same section, which requires a third class inheriting from fluke.algorithms.CentralizedFL (or PersonalizedFL), plus a YAML configuration file and a specific file layout for the CLI to discover the algorithm. The paper does not provide any quantitative measure of overhead, such as lines of code, number of methods overridden, or time-to-first-experiment, nor does it compare the effort against reimplementation from scratch or against another framework. The 'minimal overhead' claim is therefore unsupported by the evidence in the manuscript.
- [Benchmarking] The benchmarking use case, illustrated in Figure 2, shows accuracy curves for FedAvg, SCAFFOLD, FedExP, FedDyn, and FedOpt run with default hyperparameters on MNIST. The figure appears to report a single unseeded run with no error bars, and the paper provides no comparison of these curves against the results reported in the original papers. Because the framework's value for benchmarking depends on the bundled implementations being faithful to their published descriptions, the absence of validation undermines the benchmarking claim. The paper should either report aggregate statistics over multiple seeds with hyperparameter details or cite external evidence that the implementations reproduce known results.
- [fluke and Algorithm 1] The framework is explicitly built around a synchronous, centralized FedAvg-like protocol, as shown in Algorithm 1. The tutorial states that the server's fit method 'should not need to be overridden' only 'as long as the protocol follows the usual one in centralized FL.' For algorithms that deviate from this protocol—such as FedOpt's server-side optimizer state, SCAFFOLD's client control variates, personalized local models, or asynchronous updates—the user must reimplement orchestration logic by overriding fit, broadcast_model, receive_model, or finalize. The paper does not scope its 'minimal overhead' claim to FedAvg-conforming algorithms, and it does not discuss the additional effort for these common cases. This is a load-bearing gap because the paper motivates fluke by the need to prototype a wide range of new FL algorithms, not merely FedAvg variants.
- [Use cases] The paper motivates fluke by asserting that existing frameworks are 'not easy to extend' and that Flower, while extensible, has non-trivial configuration and customization. Yet no comparative evaluation against Flower or any other framework is provided. The paper does not include a user study, code complexity metrics, or any direct comparison of the effort required to implement the same algorithm in fluke versus in an existing framework. Without such evidence, the claimed advantage over the status quo remains an assertion rather than a demonstrated property.
minor comments (5)
- [Features list] There are several typographical errors, including 'minimazing' (should be 'minimizing'), 'fluke raises from' (should be 'arises'), 'ant in general' (likely 'and in general'), 'learing' (should be 'learning'), and 'oof-the-shelf' (should be 'off-the-shelf'). These should be corrected throughout.
- [CLI examples] The code block for running a custom algorithm contains '----config' (four dashes) instead of '--config', and the word 'console' appears in multiple code blocks as if it were part of the code. The wording 'The command must be run in the directory where my algorithm.py is located' should clarify whether all custom modules must reside in the current working directory or whether Python path handling is allowed.
- [Figure 1] Figure 1 is not referenced in the text before the 'Architecture overview' section, and the module names in the figure use abbreviations ('comm','algo','eval') that are only defined later in the prose. Adding an explicit reference to Figure 1 in the text and expanding the abbreviations would improve readability.
- [Documentation and tutorials] The paper states that the tutorial for adding a custom algorithm is available online, but the manuscript itself says 'for space reasons, the following steps do not include all the necessary details.' It would be helpful to include at least a minimal code listing or pseudocode for a non-trivial custom algorithm (e.g., FedOpt) to demonstrate the claimed low overhead, rather than deferring entirely to external documentation.
- [References] The reference list contains formatting inconsistencies, such as 'Https://github.com/zalandoresearch/fashion-mnist' with a capital H, and several repository citations that list only 'et al.' without full author names. These should be normalized to the journal's style.
Circularity Check
No significant circularity: the paper is a software/utility description with no derived quantities, fitted parameters, or load-bearing self-citation chain.
full rationale
This is a systems/software paper describing the fluke federated learning package. There is no derivation chain, no fitted parameter, and no predicted quantity whose value is forced by construction. The central claim is that new FL algorithms can be added by defining client and/or server classes with 'minimal overhead'; this is an engineering usability claim, not a mathematical or empirical prediction. The paper's own tutorial does show that a third class ('The final class to define is the one representing the whole algorithm') is also needed, which weakens the literal 'client and/or server' phrasing, but this is a precision/verifiability concern rather than circularity: the claim is not defined in terms of its own output. Figure 2 shows accuracy curves for fluke's bundled implementations under default settings; those implementations are the objects being demonstrated, not results derived from the framework, and the figure is explicitly an 'example of algorithms' performance with fluke' rather than evidence that the framework's own extension mechanism predicts anything. Self-references to the fluke documentation and GitHub repository point to independently checkable open-source artifacts, so they are not load-bearing self-citations that smuggle in an assumption. No uniqueness theorem is invoked, no ansatz is imported via citation, and no known result is renamed. Accordingly, there is no circular step to exhibit, and the honest finding is score 0.
Assumptions & free parameters
assumptions (2)
- domain assumption The centralized simulation architecture faithfully represents real federated learning conditions, including stable, synchronous communication.
- domain assumption The bundled implementations of published FL algorithms are correct re-implementations of the original methods.
Cite this review
Pith. "Pith review of fluke: Federated Learning Utility frameworK for Experimentation and research." pith.science (2026). https://pith.science/paper/EH4CPA2X
@misc{pith2026241215728,
author = {Pith},
title = {Pith review of: fluke: Federated Learning Utility frameworK for Experimentation and research},
year = {2026},
howpublished = {\url{https://pith.science/paper/EH4CPA2X}},
note = {Machine review of arXiv:2412.15728}
}
read the original abstract
Since its inception in 2016, Federated Learning (FL) has been gaining tremendous popularity in the machine learning community. Several frameworks have been proposed to facilitate the development of FL algorithms, but researchers often resort to implementing their algorithms from scratch, including all baselines and experiments. This is because existing frameworks are not flexible enough to support their needs or the learning curve to extend them is too steep. In this paper, we present \fluke, a Python package designed to simplify the development of new FL algorithms. fluke is specifically designed for prototyping purposes and is meant for researchers or practitioners focusing on the learning components of a federated system. fluke is open-source, and it can be either used out of the box or extended with new algorithms with minimal overhead.
Figures
Reference graph
Works this paper leans on
-
[1]
Acar, D. A. E.; Zhao, Y .; Matas, R.; Mattina, M.; What- mough, P .; and Saligrama, V . 2021. Federated Learning Based on Dynamic Regularization. In International Conference on Learning Representations
work page 2021
-
[2]
Acar, D. A. E.; et al. 2021. FedDyn. https://github. com/alpemreacar/FedDyn. Repository
work page 2021
-
[3]
O.; Beguier, C.; and Tramel, E
Andreux, M.; du Terrail, J. O.; Beguier, C.; and Tramel, E. W . 2020. Siloed Federated Learning for Multi- centric Histopathology Datasets. In Domain Adapta- tion and Representation Transfer , and Distributed and Collaborative Learning, 129–139. Cham: Springer In- ternational Publishing. ISBN 978-3-030-60548-3
work page 2020
-
[4]
Arivazhagan, M. G.; Aggarwal, V .; Singh, A. K.; and Choudhary, S. 2019. Federated Learning with Person- alization Layers. arXiv:1912.00818
arXiv 2019
-
[5]
Authors, T. T. F. 2018. TensorFlow Federated
work page 2018
-
[6]
J.; Topal, T.; Mathur, A.; Qiu, X.; Fernandez-Marques, J.; Gao, Y .; Sani, L.; Li, K
Beutel, D. J.; Topal, T.; Mathur, A.; Qiu, X.; Fernandez-Marques, J.; Gao, Y .; Sani, L.; Li, K. H.; Parcollet, T.; de Gusm˜ ao, P . P . B.; and Lane, N. D
-
[7]
Biewald, L. 2020. Experiment Tracking with Weights and Biases. Software available from wandb.com
work page 2020
-
[8]
Caldas, S.; Duddu, S. M. K.; Wu, P .; Li, T.; Koneˇ cn´ y, J.; McMahan, H. B.; Smith, V .; and Talwalkar, A
Show all 57 references
-
[9]
Cohen, G.; Afshar, S.; Tapson, J.; and van Schaik, A
-
[10]
Collins, L.; Hassani, H.; Mokhtari, A.; and Shakkot- tai, S. 2021. Exploiting Shared Representations for Personalized Federated Learning. In Meila, M.; and Zhang, T., eds., Proceedings of the 38th International Conference on Machine Learning, volume 139 of Pro- ceedings of Mac...
2021
-
[11]
Dai, Y .; Chen, Z.; Li, J.; Heinecke, S.; Sun, L.; and Xu, R. 2023. Tackling data heterogeneity in feder- ated learning with class prototypes. In Proceedings of the Thirty-Seventh AAAI Conference on Artificial In- telligence and Thirty-Fifth Conference on Innovative Application...
2023
-
[12]
Dai, Y .; et al. 2023. FedNH. https://github.com/ Y utong-Dai/FedNH/. Repository
2023
-
[13]
N.; Crowley, E
Darlow, L. N.; Crowley, E. J.; Antoniou, A.; and Storkey, A. J. 2018. CINIC-10 is not ImageNet or CIFAR-10. arXiv:1810.03505
2018 arXiv
-
[14]
M.; and Mahdavi, M
Deng, Y .; Kamani, M. M.; and Mahdavi, M
-
[15]
Diao, Y .; et al. 2023. NIID-bench. https://github.com / Xtra-Computing/NIID-Bench
2023
-
[16]
T.; Tran, N
Dinh, C. T.; Tran, N. H.; and Nguyen, T. D. 2020. Per- sonalized federated learning with moreau envelopes. In Proceedings of the 34th International Conference on Neural Information Processing Systems , NIPS ’20. Red Hook, NY , USA: Curran Associates Inc. ISBN 9781713829546
2020
-
[17]
Fallah, A.; Mokhtari, A.; and Ozdaglar, A. 2020. Per- sonalized Federated Learning with Theoretical Guar- antees: A Model-Agnostic Meta-Learning Approach. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.; and Lin, H., eds., Advances in Neural Informa- tion Processing Sy...
2020
-
[18]
J.; Edwards, B.; Pati, S.; Riv- iera, W .; Sharma, M.; Moorthy, P
Foley, P .; Sheller, M. J.; Edwards, B.; Pati, S.; Riv- iera, W .; Sharma, M.; Moorthy, P . N.; Wang, S.-h.; Martin, J.; Mirhaji, P .; Shah, P .; and Bakas, S. 2022. OpenFL: the open federated learning library. Physics in Medicine & Biology
2022
-
[19]
Hahn, S.-J.; Jeong, M.; and Lee, J. 2022. Connecting Low-Loss Subspace for Personalized Federated Learn- ing. In Proceedings of the 28th ACM SIGKDD Confer- ence on Knowledge Discovery and Data Mining , KDD ’22. ACM
2022
-
[20]
H.; Qi; and Brown, M
Hsu, T.-M. H.; Qi; and Brown, M. 2019. Measuring the Effects of Non-Identical Data Distribution for Fed- erated Visual Classification. ArXiv, abs/1909.06335
2019 arXiv
-
[21]
Huang, Y .; Chu, L.; Zhou, Z.; Wang, L.; Liu, J.; Pei, J.; and Zhang, Y . 2021. Personalized Cross-Silo Fed- erated Learning on Non-IID Data. Proceedings of the AAAI Conference on Artificial Intelligence , 35(9): 7865–7873
2021
-
[22]
Jhunjhunwala, D.; Wang, S.; and Joshi, G. 2023. Fed- ExP: Speeding Up Federated Averaging via Extrapo- lation. In The Eleventh International Conference on Learning Representations
2023
-
[23]
Jhunjhunwala, D.; et al. 2023. FedExP . https://github . com/Divyansh03/FedExP. Repository
2023
-
[24]
P .; Kale, S.; Mohri, M.; Reddi, S.; Stich, S.; and Suresh, A
Karimireddy, S. P .; Kale, S.; Mohri, M.; Reddi, S.; Stich, S.; and Suresh, A. T. 2020. SCAFFOLD: Stochastic Controlled Averaging for Federated Learn- ing. In III, H. D.; and Singh, A., eds., Proceedings of the 37th International Conference on Machine Learn- ing, volume 119 of...
2020
-
[25]
Krizhevsky, A. 2009. Learning multiple layers of fea- tures from tiny images. Technical report
2009
-
[26]
Le, Y .; and Y ang, X. 2024. Tiny ImageNet
2024
-
[27]
LeCun, Y .; and Cortes, C. 2010. MNIST handwritten digit database
2010
-
[28]
Lee, G.; et al. 2022. FedNTD. https://github.com/Lee- Gihun/FedNTD. Repository
2022
-
[29]
Li, Q.; Diao, Y .; Chen, Q.; and He, B. 2022. Feder- ated Learning on Non-IID Data Silos: An Experimen- tal Study. In IEEE International Conference on Data Engineering
2022
-
[30]
Li, Q.; He, B.; and Song, D. 2021. Model-Contrastive Federated Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion
2021
-
[31]
Li, T.; Hu, S.; Beirami, A.; and Smith, V . 2020. Ditto: Fair and Robust Federated Learning Through Person- alization. In International Conference on Machine Learning
2020
-
[32]
Li, T.; et al. 2020. Ditto. https://github.com/litian 96/ ditto. Repository
2020
-
[33]
Li, X.; Jiang, M.; Zhang, X.; Kamp, M.; and Dou, Q
-
[34]
Li, X.; et al. 2021. FedBN. https://github.com/med- air/FedBN. Repository
2021
-
[35]
P .; Liu, T.; Ziyin, L.; Salakhutdinov, R.; and Morency, L.-P
Liang, P . P .; Liu, T.; Ziyin, L.; Salakhutdinov, R.; and Morency, L.-P . 2020. Think locally, act globally: Fed- erated learning with local and global representations. arXiv preprint arXiv:2001.01523
2020 arXiv
-
[36]
Liu, Y .; Fan, T.; Chen, T.; Xu, Q.; and Y ang, Q. 2021. FA TE: an industrial grade platform for collaborative learning with data protection. J. Mach. Learn. Res. , 22(1)
2021
-
[37]
Luo, M.; Chen, F.; Hu, D.; Zhang, Y .; Liang, J.; and Feng, J. 2021. No Fear of Heterogeneity: Classi- fier Calibration for Federated Learning with Non-IID Data. In Beygelzimer, A.; Dauphin, Y .; Liang, P .; and V aughan, J. W ., eds.,Advances in Neural Information Processing Systems
2021
-
[38]
McMahan, B.; Moore, E.; Ramage, D.; Hampson, S.; and Arcas, B. A. y. 2017. Communication-Efficient Learning of Deep Networks from Decentralized Data. In Singh, A.; and Zhu, J., eds., Proceedings of the 20th International Conference on Artificial Intelli- gence and Statistics, vo...
2017
-
[39]
Mengdi, W .; et al. 2024. TurboSVM-FL. https://github. com/Kasneci-Lab/TurboSVM-FL. Repository
2024
-
[40]
Netzer, Y .; Wang, T.; Coates, A.; Bissacco, A.; Wu, B.; and Ng, A. Y . 2011. Reading digits in natural images with unsupervised feature learning
2011
-
[41]
Oh, J.; Kim, S.; and Y un, S.-Y . 2022. FedBABU: To- ward Enhanced Representation for Federated Image Classification. In International Conference on Learn- ing Representations
2022
-
[42]
Qi, P .; Chiaro, D.; Giampaolo, F.; and Piccialli, F
-
[43]
J.; Charles, Z.; Zaheer, M.; Garrett, Z.; Rush , K.; Koneˇ cn´ y, J.; Kumar, S.; and McMahan, H
Reddi, S. J.; Charles, Z.; Zaheer, M.; Garrett, Z.; Rush , K.; Koneˇ cn´ y, J.; Kumar, S.; and McMahan, H. B. 2021. Adaptive Federated Optimization. In International Conference on Learning Representations
2021
-
[44]
Roth, H. R.; Cheng, Y .; Wen, Y .; Y ang, I.; Xu, Z.; Hsieh, Y .-T.; Kersten, K.; Harouni, A.; Zhao, C.; Lu, K.; Zhang, Z.; Li, W .; Myronenko, A.; Y ang, D.; Y ang, S.; Rieke, N.; Quraini, A.; Chen, C.; Xu, D.; Ma, N.; Dogra, P .; Flores, M.; and Feng, A. 2023. NVIDIA FLARE: ...
2023
-
[45]
K.; Li, T.; Sanjabi, M.; Zaheer, M.; Talwalkar, A.; and Smith, V
Sahu, A. K.; Li, T.; Sanjabi, M.; Zaheer, M.; Talwalkar, A.; and Smith, V . 2018. Federated Optimization in Het- erogeneous Networks. arXiv: Learning
2018
-
[46]
Tan, Y .; Long, G.; Liu, L.; Zhou, T.; Lu, Q.; Jiang, J.; and Zhang, C. 2022. FedProto: Federated Prototype Learning across Heterogeneous Clients. In AAAI Con- ference on Artificial Intelligence
2022
-
[47]
Tan, Y .; et al. 2021. FedProto. https://github.com/ yuetan031/FedProto
2021
-
[48]
Wang, J.; Liu, Q.; Liang, H.; Joshi, G.; and Poor, H. V
-
[49]
Xiao, H.; Rasul, K.; and V ollgraf, R. 2017. Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms. Https://github.com/zalandoresearch/fashion-mnist
2017
-
[50]
Zhang, J.; Li, Z.; Li, B.; Xu, J.; Wu, S.; Ding, S.; and Wu, C. 2022. Federated Learning with Label Distribution Skew via Logits Calibration. In Chaud- huri, K.; Jegelka, S.; Song, L.; Szepesvari, C.; Niu, G.; and Sabato, S., eds., Proceedings of the 39th In- ternational Confe...
2022
-
[55]
In Proceedings of the 34th International Conference on Neural Infor- mation Processing Systems, NIPS ’20
Tackling the objective inconsistency problem in heterogeneous federated optimization. In Proceedings of the 34th International Conference on Neural Infor- mation Processing Systems, NIPS ’20. Red Hook, NY , USA: Curran Associates Inc. ISBN 9781713829546
- [2017]
- [2019]
- [2020]
-
[2021]
In Interna- tional Conference on Learning Representations
Fed {BN}: Federated Learning on Non- {IID} Features via Local Batch Normalization. In Interna- tional Conference on Learning Representations
-
[2022]
arXiv:2007.14390
Flower: A Friendly Federated Learning Re- search Framework. arXiv:2007.14390
2007 arXiv
-
[2024]
In Machine Learning and Knowledge Discovery in Databases
KAF `E: Kernel Aggregation for FEderated. In Machine Learning and Knowledge Discovery in Databases. Research Track , 56–71. Cham: Springer Nature Switzerland. ISBN 978-3-031-70359-1
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.