Pith. sign in

REVIEW 3 major objections 2 minor 90 references

This paper claims that AdaptFED, a hypernetwork-based federated learning method with task-aware client embeddings and low-rank conditioning, beats state-of-the-art baselines across eight datasets—but the supplied body is an unrelated graph

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

The abstract describes AdaptFED, a claimed federated learning method, but the full text is an unrelated graph theory paper, so the claimed results are absent from the submission.

T0 review reviewed 2026-08-05 challenge →

load-bearing objection Two unrelated documents are glued together under one arXiv ID: the abstract advertises a federated learning extension of TransFed, the full text is a graph theory paper by different authors—so the central claim has zero in-text support. the 3 major comments →

arxiv 2508.10840 v1 pith:ILPMB5A6 submitted 2025-08-14 cs.CV

Generalizable Federated Learning using Client Adaptive Focal Modulation

classification cs.CV
keywords federated learningfocal modulationhypernetworktask-aware embeddingsnon-IIDcross-domainsource-freelow-rank conditioning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The abstract describes AdaptFED, an extension of a prior transformer-based federated learning framework, in which a learn-to-adapt hypernetwork generates personalized focal modulation layers for each client, refined with task-aware client embeddings and a low-rank conditioning variant to reduce communication cost. It asserts enhanced theoretical bounds on adaptation performance and validation across eight datasets spanning images, time-series, and multilingual text, with state-of-the-art results in source-free and cross-task setups. The full text supplied, however, is a different manuscript about quadratic embedding constants of graph joins and Cartesian products; it contains no description of AdaptFED, no bounds, and no experiments. The abstract itself promises these contents, but they are absent from the body. A sympathetic reader can only take the abstract's claims as an unverified assertion of a plausible idea.

Core claim

On its own terms, the paper claims to establish that AdaptFED outperforms state-of-the-art baselines in federated learning, particularly in source-free and cross-task settings, through per-client focal modulation layers generated by a hypernetwork with task-aware embeddings and low-rank conditioning, supported by theoretical bounds and eight-dataset experiments. The supplied full text does not support this: it is a mathematics manuscript on the quadratic embedding constant, with no mention of federated learning, focal modulation, or experiments. The discovery therefore exists only at the level of the abstract's assertion.

What carries the argument

The named mechanism is the learn-to-adapt hypernetwork that generates personalized focal modulation layers per client, together with task-aware client embeddings and low-rank hypernetwork conditioning to reduce server-client communication overhead. In the abstract this is the object claimed to carry generalization across non-IID and cross-domain clients; in the supplied text it does not appear.

Load-bearing premise

The load-bearing premise is that the supplied body text is the AdaptFED manuscript and that a hypernetwork can actually generate personalized focal modulation layers that generalize across non-IID and cross-domain clients.

What would settle it

Search the full text for any occurrence of 'AdaptFED', 'focal modulation', or 'hypernetwork' and for any experimental table; their complete absence would confirm that the promised method and eight-dataset validation are not in this submission. Alternatively, run the stated eight-dataset benchmark comparing AdaptFED against the baselines in source-free and cross-task settings; the superiority claim fails if AdaptFED does not outperform them.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • If correct, a single global transformer could specialize to each client's data distribution via per-client focal modulation without exchanging raw samples.
  • If correct, task-aware embeddings would let the same adaptation mechanism transfer across modalities, from images to time-series to multilingual text.
  • If correct, low-rank conditioning would scale communication savings with the rank of the conditioning subspace, making deployment feasible on resource-constrained devices.
  • If correct, the announced bounds would let practitioners estimate adaptation error from client embeddings before deployment.
  • If correct, the eight-dataset results would place AdaptFED ahead of existing baselines in source-free and cross-task federated benchmarks.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Because the body text contains none of the promised FL content, the abstract's central claim is currently unfalsifiable in this document; the falsifier is simply locating the missing method and experiments.
  • A testable consequence of the hypernetwork design is that communication cost should fall as the conditioning rank is reduced; measuring exchanged bytes at several ranks would check that claim independently of the datasets.
  • The graph-theoretic body suggests a possible submission mix-up; if so, the practical issue is that readers have no way to verify the abstract without a separate, complete version of the FL paper.
  • If the eight-dataset claim is real, exact protocol transparency (splits, baselines, metrics) is a precondition for reproducing the comparisons; its absence would make the state-of-the-art claim untestable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The submission presents an abstract claiming an extended federated learning (FL) method, AdaptFED, built on the authors' prior TransFed framework. The abstract asserts a refined adaptation strategy with task-aware client embeddings, enhanced theoretical bounds on adaptation performance, and extensive experiments on eight datasets demonstrating superiority over state-of-the-art baselines, plus a low-rank hypernetwork conditioning variant for communication efficiency. However, the full text provided is an entirely different manuscript: a graph theory paper titled 'Quadratic Embedding Constants of Cartesian Products and Joins of Graphs' by P.N. Choudhury and R. Nandi (arXiv:2508.10834v2). This body contains no federated learning, no hypernetwork, no focal modulation, no client adaptation, no theoretical bounds on adaptation, and no experimental evaluation of any kind.

Significance. If the claims in the abstract were substantiated, AdaptFED could represent a meaningful advance in generalizable federated learning, particularly for source-free and cross-task scenarios. The paper, however, provides no method description, no derivations, no datasets, and no results. The only concrete artifact is the abstract's assertion, which is not supported by any content in the body. The graph-theoretic material is unrelated to the claimed contribution, so the significance of the FL claims cannot be assessed. The paper therefore makes a significant claim without providing any of the evidence that would be required to evaluate it.

major comments (3)
  1. [Full Text (entire manuscript)] The body of the submission is a graph theory paper on quadratic embedding constants, not the federated learning paper described in the abstract. There is no description of AdaptFED, TransFed, hypernetworks, focal modulation, client embeddings, low-rank conditioning, or any federated learning experiment. The central claim in the abstract—'Extensive experiments on eight diverse datasets reaffirm the superiority of our method'—is completely unsupported because the experiments are absent. This is not a fixable local gap but a wholesale mismatch between the claimed contribution and the submitted text.
  2. [Abstract, items (1)–(3)] The abstract promises (1) a refined adaptation strategy with task-aware client embeddings, (2) enhanced theoretical bounds on adaptation performance, and (3) broader empirical validation across eight datasets including time-series and multilingual data. None of these appear in the full text. In particular, there is no equation, theorem, or proof for the claimed 'enhanced theoretical bounds,' and no table or figure reporting experimental results. The reader cannot verify any of the abstract's assertions.
  3. [Introduction (graph theory paper)] The body's Introduction defines graph-theoretic concepts such as quadratic embedding constants and discusses Schoenberg's work. This content is unrelated to federated learning. Even taken on its own merits, it cannot serve as support for the abstract's federated learning claims. The absence of any connection between the body and the abstract means the manuscript is internally inconsistent.
minor comments (2)
  1. [Title and formatting] The title in the full text is 'QUADRA TIC EMBEDDING CONST ANTS OF CAR TESIAN PRODUCTS AND JOINS OF GRAPHS' with unusual spacing; this appears to be a rendering artifact, but it further signals that the manuscript is not the intended AdaptFED paper.
  2. [References] The reference list is entirely devoted to graph theory and distance geometry, with no citations to federated learning, hypernetworks, or focal modulation. This is consistent with the mismatch but also means the paper cannot be used to situate the claimed contribution in the FL literature.

Circularity Check

2 steps flagged

Central AdaptFED claim is asserted without method or experiments; only self-citation to TransFed supports it.

specific steps
  1. self citation load bearing [Abstract, first paragraph]
    "Our prior work, TransFed, introduced a robust transformer-based FL framework that leverages a learn-to-adapt hypernetwork to generate personalized focal modulation layers per client, outperforming traditional methods in non-IID and cross-domain settings. In this extended version, we propose AdaptFED, where we deepen the investigation of focal modulation in generalizable FL by incorporating: (1) a refined adaptation strategy ... (3) broader empirical validation ..."

    The abstract's rationale for AdaptFED's existence and expected performance is anchored entirely in the authors' own prior TransFed. The full text contains no AdaptFED method or evaluation; thus the only evidence that the extension is viable is the self-citation to TransFed's claimed success. The central claim (AdaptFED superiority) is thereby supported by a self-citation chain rather than by independent derivation or external benchmarks.

  2. other [Abstract, first paragraph, item (3)]
    "Extensive experiments on eight diverse datasets reaffirm the superiority of our method over state-of-the-art baselines, particularly in source-free and cross-task federated setups."

    This assertion of empirical superiority is the sole basis for the paper's conclusion, yet the full text provides no experiments, no datasets, and no results. The claim is therefore self-referential: the paper's evidence for its conclusion is the conclusion restated as an assertion. No external or internal data is available to check the claim, making it equivalent to an unsupported assertion rather than a derived result.

full rationale

The abstract describes an FL extension of the authors' prior TransFed, claiming enhanced theoretical bounds and validation on eight datasets. However, the full text is an unrelated graph theory paper ('Quadratic Embedding Constants of Cartesian Products and Joins of Graphs') by different authors and contains no federated learning, no hypernetwork, no focal modulation, no experiments, and no bounds. The only support for the central claim of AdaptFED's superiority is the abstract's own assertion, which in turn relies on the self-cited prior work TransFed. Since the method, derivations, and experiments are entirely absent, the claim cannot be validated against any external or independent content. This is a load-bearing self-citation chain and an unsupported assertion that functions as its own evidence, warranting a circularity score of 6.

Axiom & Free-Parameter Ledger

0 free parameters · 4 axioms · 0 invented entities

The claimed AdaptFED method contributes no auditable free parameters because the method is absent from the text; the abstract mentions a low-rank hypernetwork variant but gives no rank, dimension, or loss details. The graph theory body relies on standard background: finite simple connected unweighted graphs, Schoenberg's conditional negative definiteness characterization, and the QEC definition via the distance matrix spectrum. The only ad hoc assumption attached to the stated claim is that a hypernetwork can generate transferable personalized focal modulation layers per client, asserted without demonstration in the abstract. No invented entities: the body introduces no new particles, forces, or dimensions; the abstract's task-aware client embeddings and low-rank conditioning are design components of a claimed method, not new entities, and they appear nowhere in the provided text.

axioms (4)
  • domain assumption All graphs are finite, simple, connected, and unweighted
    Stated at the start of §1 of the body; the QEC definitions and all derived formulas apply only within this class.
  • standard math QE class is characterized by conditional negative definiteness of the distance matrix (Schoenberg 1935)
    The body relies on this classical theorem, cited as [21] in §1, to define and compute the quadratic embedding constant.
  • standard math QEC equals the least eigenvalue of the distance matrix restricted to the subspace orthogonal to the all-ones vector
    Definition of QEC in §1 following Obata-Zakiyyah [19]; all results in the body are expressed through this spectral quantity.
  • ad hoc to paper A hypernetwork can generate personalized focal modulation layers that transfer across non-IID and cross-domain clients
    This is the load-bearing premise of the abstract's claimed method (abstract, items 1 and 3); it is asserted without derivation or demonstration in the provided text.

reviewed 2026-08-05 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Generalizable Federated Learning using Client Adaptive Focal Modulation." pith.science (2026). https://pith.science/paper/ILPMB5A6

@misc{pith2026250810840,
  author       = {Pith},
  title        = {Pith review of: Generalizable Federated Learning using Client Adaptive Focal Modulation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ILPMB5A6}},
  note         = {Machine review of arXiv:2508.10840}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Federated learning (FL) has proven essential for privacy-preserving, collaborative training across distributed clients. Our prior work, TransFed, introduced a robust transformer-based FL framework that leverages a learn-to-adapt hypernetwork to generate personalized focal modulation layers per client, outperforming traditional methods in non-IID and cross-domain settings. In this extended version, we propose AdaptFED, where we deepen the investigation of focal modulation in generalizable FL by incorporating: (1) a refined adaptation strategy that integrates task-aware client embeddings to personalize modulation dynamics further, (2) enhanced theoretical bounds on adaptation performance, and (3) broader empirical validation across additional modalities, including time-series and multilingual data. We also introduce an efficient variant of TransFed that reduces server-client communication overhead via low-rank hypernetwork conditioning, enabling scalable deployment in resource-constrained environments. Extensive experiments on eight diverse datasets reaffirm the superiority of our method over state-of-the-art baselines, particularly in source-free and cross-task federated setups. Our findings not only extend the capabilities of focal modulation in FL but also pave the way for more adaptive, scalable, and generalizable transformer-based federated systems. The code is available at http://github.com/Tajamul21/TransFed

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

90 extracted references · 59 canonical work pages · 1 internal anchor

  1. [1]

    Abnar and W

    S. Abnar and W. Zuidema. Quantifying attention flow in transformers. arXiv preprint arXiv:2005.00928, 2020

  2. [2]

    D. A. E. Acar, Y. Zhao, R. M. Navarro, M. Mattina, P. N. Whatmough, and V. Saligrama. Federated learning based on dynamic regularization. arXiv preprint arXiv:2111.04263, 2021

  3. [3]

    Achituve, A

    I. Achituve, A. Shamsian, A. Navon, G. Chechik, and E. Fetaya. Personalized federated learning with gaussian processes. Advances in NeurIPS, 34: 0 8392--8406, 2021

  4. [4]

    Alberti, A

    E. Alberti, A. Tavera, C. Masone, and B. Caputo. Idda: A large-scale multi-domain dataset for autonomous driving. IEEE Robotics and Automation Letters, 5 0 (4): 0 5526--5533, 2020

  5. [5]

    M. G. Arivazhagan, V. Aggarwal, A. K. Singh, and S. Choudhary. Federated learning with personalization layers. arXiv preprint arXiv:1912.00818, 2019

  6. [6]

    Ashraf, F

    T. Ashraf, F. Bin Afzal Mir, and I. A. Gillani. Transfed: A way to epitomize focal modulation using transformer-based federated learning. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 554--563, 2024

  7. [7]

    Caron, H

    M. Caron, H. Touvron, I. Misra, H. J \'e gou, J. Mairal, P. Bojanowski, and A. Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF ICCV, pages 9650--9660, 2021

  8. [8]

    Chen and W.-L

    H.-Y. Chen and W.-L. Chao. On bridging generic and personalized federated learning for image classification. arXiv preprint arXiv:2107.00778, 2021

  9. [9]

    Chen, C.-H

    H.-Y. Chen, C.-H. Tu, Z. Li, H.-W. Shen, and W.-L. Chao. On the importance and applicability of pre-training for federated learning. arXiv preprint arXiv:2206.11488, 2022

  10. [10]

    Collins, H

    L. Collins, H. Hassani, A. Mokhtari, and S. Shakkottai. Exploiting shared representations for personalized federated learning. In ICML, pages 2089--2099. PMLR, 2021

  11. [11]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020

  12. [12]

    Fallah, A

    A. Fallah, A. Mokhtari, and A. Ozdaglar. Personalized federated learning with theoretical guarantees: A model-agnostic meta-learning approach. Advances in NeurIPS, 33: 0 3557--3568, 2020

  13. [13]

    Z. Feng, Y. Wang, J. Li, F. Yang, J. Lou, T. Mi, R. Qiu, Z. Liao, et al. Robust and communication-efficient federated domain adaptation via random features. arXiv preprint arXiv:2311.04686, 2023

  14. [14]

    Ganin and V

    Y. Ganin and V. Lempitsky. Unsupervised domain adaptation by backpropagation. In ICML, pages 1180--1189. PMLR, 2015

  15. [15]

    Geirhos, P

    R. Geirhos, P. Rubisch, C. Michaelis, M. Bethge, F. A. Wichmann, and W. Brendel. Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness. arXiv preprint arXiv:1811.12231, 2018

  16. [16]

    Ghosh, J

    A. Ghosh, J. Chung, D. Yin, and K. Ramchandran. An efficient framework for clustered federated learning. Advances in NeurIPS, 33: 0 19586--19597, 2020

  17. [17]

    B. Gong, Y. Shi, F. Sha, and K. Grauman. Geodesic flow kernel for unsupervised domain adaptation. In 2012 IEEE CVPR. IEEE, 2012

  18. [18]

    Griffin, A

    G. Griffin, A. Holub, and P. Perona. Caltech-256 object category dataset. 2007

  19. [19]

    D. Ha, A. Dai, and Q. V. Le. Hypernetworks. arXiv preprint arXiv:1609.09106, 2016

  20. [20]

    Hanzely, S

    F. Hanzely, S. Hanzely, S. Horv \'a th, and P. Richt \'a rik. Lower bounds and optimal algorithms for personalized federated learning. Advances in NeurIPS, 33: 0 2304--2315, 2020

  21. [21]

    K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. In Proceedings of the IEEE CVPR, pages 770--778, 2016

  22. [22]

    Hoyer, D

    L. Hoyer, D. Dai, and L. Van Gool. Daformer: Improving network architectures and training strategies for domain-adaptive semantic segmentation. In Proceedings of the IEEE/CVF CVPR, 2022

  23. [23]

    Hsieh, A

    K. Hsieh, A. Phanishayee, O. Mutlu, and P. Gibbons. The non-iid data quagmire of decentralized machine learning. In ICML. PMLR, 2020

  24. [24]

    T.-M. H. Hsu, H. Qi, and M. Brown. Federated visual classification with real-world data distribution. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part X 16, pages 76--92. Springer, 2020

  25. [25]

    Huang, L

    Y. Huang, L. Chu, Z. Zhou, L. Wang, J. Liu, J. Pei, and Y. Zhang. Personalized cross-silo federated learning on non-iid data. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, 2021

  26. [26]

    Johnson, M

    J. Johnson, M. Douze, and H. J \'e gou. Billion-scale similarity search with gpus. IEEE Transactions on Big Data, 7 0 (3), 2019

  27. [27]

    G. A. Kaissis, M. R. Makowski, D. R \"u ckert, and R. F. Braren. Secure, privacy-preserving and federated machine learning in medical imaging. Nature Machine Intelligence, 2 0 (6): 0 305--311, 2020

  28. [28]

    S. P. Karimireddy, M. Jaggi, S. Kale, M. Mohri, S. J. Reddi, S. U. Stich, and A. T. Suresh. Mime: Mimicking centralized stochastic algorithms in federated learning. arXiv preprint arXiv:2008.03606, 2020 a

  29. [29]

    S. P. Karimireddy, S. Kale, M. Mohri, S. Reddi, S. Stich, and A. T. Suresh. Scaffold: Stochastic controlled averaging for federated learning. In ICML, pages 5132--5143. PMLR, 2020 b

  30. [30]

    Kermany, K

    D. Kermany, K. Zhang, M. Goldbaum, et al. Labeled optical coherence tomography (oct) and chest x-ray images for classification. Mendeley data, 2 0 (2): 0 651, 2018

  31. [31]

    G. Kim, J. Kim, and B. Han. Communication-efficient federated learning with accelerated client gradient. In Proceedings of the IEEE/CVF CVPR, pages 12385--12394, 2024

  32. [32]

    Krizhevsky, G

    A. Krizhevsky, G. Hinton, et al. Learning multiple layers of features from tiny images. 2009

  33. [33]

    J. N. Kundu, A. Kulkarni, A. Singh, V. Jampani, and R. V. Babu. Generalize then adapt: Source-free domain adaptive semantic segmentation. In Proceedings of the IEEE/CVF ICCV, pages 7046--7056, 2021

  34. [34]

    Li and J

    D. Li and J. Wang. Fedmd: Heterogenous federated learning via model distillation. arXiv preprint arXiv:1910.03581, 2019

  35. [35]

    H. Li, Z. Cai, J. Wang, J. Tang, W. Ding, C.-T. Lin, and Y. Shi. Fedtp: Federated learning by transformer personalization. IEEE transactions on neural networks and learning systems, 2023 a

  36. [36]

    T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith. Federated optimization in heterogeneous networks. Proceedings of Machine learning and systems, 2: 0 429--450, 2020

  37. [37]

    T. Li, S. Hu, A. Beirami, and V. Smith. Ditto: Fair and robust federated learning through personalization. In ICML. PMLR, 2021 a

  38. [38]

    X. Li, M. Jiang, X. Zhang, M. Kamp, and Q. Dou. Fedbn: Federated learning on non-iid features via local batch normalization. arXiv preprint arXiv:2102.07623, 2021 b

  39. [39]

    Y. Li, N. Wang, J. Shi, and X. Hou. Revisiting batch normalization for practical domain adaptation. arXiv preprint arXiv:1603.04779, 2016

  40. [40]

    Z. Li, Q. Li, Y. Zhou, W. Zhong, G. Zhang, and C. Wu. Edge-cloud collaborative learning with federated and centralized features. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in IR, pages 1949--1953, 2023 b

  41. [41]

    Z. Li, T. Lin, X. Shang, and C. Wu. Revisiting weighted aggregation in federated learning with neural networks. In ICML. PMLR, 2023 c

  42. [42]

    Q. Lian, F. Lv, L. Duan, and B. Gong. Constructing self-motivated pyramid curriculums for cross-domain semantic segmentation: A non-adversarial approach. In Proceedings of the IEEE/CVF ICCV, 2019

  43. [43]

    P. P. Liang, T. Liu, L. Ziyin, N. B. Allen, R. P. Auerbach, D. Brent, R. Salakhutdinov, and L.-P. Morency. Think locally, act globally: Federated learning with local and global representations. arXiv preprint arXiv:2001.01523, 2020

  44. [44]

    B. Liu, Y. Guo, and X. Chen. Pfa: Privacy-preserving federated adaptation for effective model personalization. In Proceedings of the Web Conference 2021, pages 923--934, 2021 a

  45. [45]

    Y. Liu, W. Zhang, and J. Wang. Source-free domain adaptation for semantic segmentation. In Proceedings of the IEEE/CVF CVPR, 2021 b

  46. [46]

    M. Long, Y. Cao, J. Wang, and M. Jordan. Learning transferable features with deep adaptation networks. In ICML. PMLR, 2015

  47. [47]

    Y. Luo, L. Zheng, T. Guan, J. Yu, and Y. Yang. Taking a closer look at domain shift: Category-level adversaries for semantics consistent domain adaptation. In Proceedings of the IEEE/CVF CVPR, 2019

  48. [48]

    X. Ma, J. Zhang, S. Guo, and W. Xu. Layer-wised model aggregation for personalized federated learning. In Proceedings of the IEEE/CVF CVPR, pages 10092--10101, 2022

  49. [49]

    Mansour, M

    Y. Mansour, M. Mohri, J. Ro, and A. T. Suresh. Three approaches for personalization with applications to federated learning. arXiv preprint arXiv:2002.10619, 2020

  50. [50]

    Marfoq, G

    O. Marfoq, G. Neglia, R. Vidal, and L. Kameni. Personalized federated learning through local memorization. In ICML. PMLR, 2022

  51. [51]

    Maria Carlucci, L

    F. Maria Carlucci, L. Porzi, B. Caputo, E. Ricci, and S. Rota Bulo. Autodial: Automatic domain alignment layers. In Proceedings of the IEEE ICCV, pages 5067--5075, 2017

  52. [52]

    Communication-efficient learning of deep networks from decentralized data

    McMahan et al. Communication-efficient learning of deep networks from decentralized data. In AIS, pages 1273--1282. PMLR, 2017

  53. [53]

    Mendieta, T

    M. Mendieta, T. Yang, P. Wang, M. Lee, Z. Ding, and C. Chen. Local learning matters: Rethinking data heterogeneity in federated learning. In Proceedings of the IEEE/CVF CVPR, pages 8397--8406, 2022

  54. [54]

    Nguyen, J

    J. Nguyen, J. Wang, K. Malik, M. Sanjabi, and M. Rabbat. Where to begin? on the impact of pre-training and initialization in federated learning. arXiv preprint arXiv:2206.15387, 2022

  55. [55]

    X. Peng, Z. Huang, Y. Zhu, and K. Saenko. Federated adversarial domain adaptation. arXiv preprint arXiv:1911.02054, 2019

  56. [56]

    L. Qu, Y. Zhou, P. P. Liang, Y. Xia, F. Wang, E. Adeli, L. Fei-Fei, and D. Rubin. Rethinking architecture design for tackling data heterogeneity in federated learning. In Proceedings of the IEEE/CVF CVPR, pages 10061--10071, 2022

  57. [57]

    Ramachandran, N

    P. Ramachandran, N. Parmar, A. Vaswani, I. Bello, A. Levskaya, and J. Shlens. Stand-alone self-attention in vision models. Advances in NeurIPS, 32, 2019

  58. [58]

    S. R. Richter, V. Vineet, S. Roth, and V. Koltun. Playing for data: Ground truth from computer games. In Computer Vision--ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part II 14, pages 102--118. Springer, 2016

  59. [59]

    Saenko, B

    K. Saenko, B. Kulis, M. Fritz, and T. Darrell. Adapting visual category models to new domains. In Computer Vision--ECCV 2010: 11th ECCV Proceedings, Part IV 11. Springer, 2010

  60. [60]

    Saito, K

    K. Saito, K. Watanabe, Y. Ushiku, and T. Harada. Maximum classifier discrepancy for unsupervised domain adaptation. In Proceedings of the IEEE CVPR, pages 3723--3732, 2018

  61. [61]

    Sattler, S

    F. Sattler, S. Wiedemann, K.-R. M \"u ller, and W. Samek. Robust and communication-efficient federated learning from non-iid data. IEEE transactions on neural networks and learning systems, 31 0 (9): 0 3400--3413, 2019

  62. [62]

    Sattler, K.-R

    F. Sattler, K.-R. M \"u ller, and W. Samek. Clustered federated learning: Model-agnostic distributed multitask optimization under privacy constraints. IEEE transactions on neural networks and learning systems, 32 0 (8): 0 3710--3722, 2020

  63. [63]

    Shakespeare

    W. Shakespeare. The complete works of william shakespeare. http://www. gutenberg.org/files/100/old/1994-01-100.zip., 1994. Accessed on 21-03, 2024

  64. [64]

    G. Sun, M. Mendieta, J. Luo, S. Wu, and C. Chen. Fedperfix: Towards partial model personalization of vision transformers in federated learning. In Proceedings of the IEEE/CVF ICCV, pages 4988--4998, 2023

  65. [65]

    T Dinh, N

    C. T Dinh, N. Tran, and J. Nguyen. Personalized federated learning with moreau envelopes. Advances in NeurIPS, 33, 2020

  66. [66]

    Towards personalized federated learning

    Tan et al. Towards personalized federated learning. IEEE Transactions on NNLS, 2022

  67. [67]

    Testolina, F

    P. Testolina, F. Barbato, U. Michieli, M. Giordani, P. Zanuttigh, and M. Zorzi. Selma: Semantic large-scale multimodal acquisitions in variable weather, daytime and viewpoints. IEEE Transactions on Intelligent Transportation Systems, 2023

  68. [68]

    Toldo, A

    M. Toldo, A. Maracani, U. Michieli, and P. Zanuttigh. Unsupervised domain adaptation in semantic segmentation: a review. Technologies, 8 0 (2): 0 35, 2020

  69. [69]

    N. H. Tran, W. Bao, A. Zomaya, M. N. Nguyen, and C. S. Hong. Federated learning over wireless networks: Optimization model design and analysis. In IEEE INFOCOM 2019-IEEE conference on computer communications, pages 1387--1395. IEEE, 2019

  70. [70]

    Tsai, W.-C

    Y.-H. Tsai, W.-C. Hung, S. Schulter, K. Sohn, M.-H. Yang, and M. Chandraker. Learning to adapt structured output space for semantic segmentation. In Proceedings of the IEEE CVPR, pages 7472--7481, 2018

  71. [71]

    Varno, M

    F. Varno, M. Saghayi, L. Rafiee Sevyeri, S. Gupta, S. Matwin, and M. Havaei. Adabest: Minimizing client drift in federated learning via adaptive bias estimation. In European Conference on Computer Vision, pages 710--726. Springer, 2022

  72. [72]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, . Kaiser, and I. Polosukhin. Attention is all you need. Advances in NeurIPS, 30, 2017

  73. [73]

    K. Wang, R. Mathews, C. Kiddon, H. Eichner, F. Beaufays, and D. Ramage. Federated evaluation of on-device personalization. arXiv preprint arXiv:1910.10252, 2019

  74. [74]

    W. Wang, J. Dai, Z. Chen, Z. Huang, Z. Li, X. Zhu, X. Hu, T. Lu, L. Lu, H. Li, et al. Internimage: Exploring large-scale vision foundation models with deformable convolutions. In Proceedings of the IEEE/CVF CVPR, pages 14408--14419, 2023

  75. [75]

    X. Wang, Y. Peng, L. Lu, Z. Lu, M. Bagheri, and R. M. Summers. Chestx-ray8: Hospital-scale chest x-ray database and benchmarks on weakly-supervised classification and localization of common thorax diseases. In Proceedings of the IEEE CVPR, pages 2097--2106, 2017

  76. [76]

    J. Xu, S. Wang, L. Wang, and A. C.-C. Yao. Fedcm: Federated learning with client-level momentum. arXiv preprint arXiv:2106.10874, 2021

  77. [77]

    J. Yang, C. Li, X. Dai, and J. Gao. Focal modulation networks. Advances in NeurIPS, 35: 0 4203--4217, 2022

  78. [78]

    Yang and S

    Y. Yang and S. Soatto. Fda: Fourier domain adaptation for semantic segmentation. In Proceedings of the IEEE/CVF CVPR, 2020

  79. [79]

    C.-H. Yao, B. Gong, H. Qi, Y. Cui, Y. Zhu, and M.-H. Yang. Federated multi-target domain adaptation. In Proceedings of the IEEE/CVF WACV, pages 1424--1433, 2022

  80. [80]

    Zhang, S

    J. Zhang, S. Guo, Z. Qu, D. Zeng, Y. Zhan, Q. Liu, and R. Akerkar. Adaptive federated learning on non-iid data with resource constraint. IEEE Transactions on Computers, 71 0 (7): 0 1655--1667, 2021 a

Showing first 80 references.

This paper was first reviewed by deepseek-v4-flash on August 5, 2026.