Pith. sign in

REVIEW 4 cited by

Non-IID data in Federated Learning: A Survey with Taxonomy, Metrics, Methods, Frameworks and Future Directions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.12377 v2 pith:T2NQAV3N submitted 2024-11-19 cs.LG

classification cs.LG
keywords datanon-iidlearningsurveyclientsdirectionsdistributedfederated
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in machine learning have highlighted Federated Learning (FL) as a promising approach that enables multiple distributed users (so-called clients) to collectively train ML models without sharing their private data. While this privacy-preserving method shows potential, it struggles when data across clients is not independent and identically distributed (non-IID) data. The latter remains an unsolved challenge that can result in poorer model performance and slower training times. Despite the significance of non-IID data in FL, there is a lack of consensus among researchers about its classification and quantification. This technical survey aims to fill that gap by providing a detailed taxonomy for non-IID data, partition protocols, and metrics to quantify data heterogeneity. Additionally, we describe popular solutions to address non-IID data and standardized frameworks employed in FL with heterogeneous data. Based on our state-of-the-art survey, we present key lessons learned and suggest promising future research directions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Privacy-Preserving Federated Averaging with Byzantine Aggregators in Asynchronous Networks

    cs.DC 2026-01 conditional novelty 7.0 of 10

    A new protocol enables differentially private federated averaging in asynchronous networks with fully Byzantine aggregators, using replicated servers, LWE masking, and verifiable cluster shuffling.

  2. Model Fusion via Retrofitting

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A neuron-centric fusion method that clusters intermediate activations of independently trained models into importance-weighted centroids and fits the fused network to them, outperforming baselines in zero-shot non-IID...

  3. On the Effectiveness of Adaptation Strategies for VLM-Based Federated Learning in Remote Sensing

    cs.CV 2026-08 conditional novelty 5.0 of 10

    In federated remote sensing, LoRA tuning of a frozen CLIP model achieves the best accuracy-to-communication trade-off, while full fine-tuning causes severe catastrophic forgetting of pretrained knowledge.

  4. PIcsC: Partitioning-Induced Covariate Shift Correction

    cs.LG 2026-07 reject novelty 3.0 of 10

    A Fisher-information regularizer is proposed to correct partition-induced covariate shift in cross-validation and federated learning, with reported gains of 3-5 points over FedAvg-class baselines.

Pith tools