{"id":"b838e2ea-2f15-4a2f-80bf-e4fabdcb5ac4","arxiv_id":"2509.00799","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of fairness in federated learning: bias sources, mitigation algorithms, evaluation metrics, and open problems.","lead":"This paper reviews how fairness is handled in federated learning, where many devices train a shared model without sharing raw data. It maps the sources of bias, the algorithms proposed to fix them, and the metrics used to measure fairness.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Survey selection methodology cannot support the claimed comprehensiveness; §1.4 is non-reproducible and the reported 'trends' may be artifacts of an undocumented corpus.","rationale":"I read the paper in good faith as a literature survey whose value is organizational rather than novel. The reader's CONDITIONAL verdict and my own analysis converge: the paper is broadly competent but has a weak empirical basis. The central claim of comprehensiveness is not falsifiable through experiments; it depends on the selected literature being representative. The weakest point is precisely the paper-selection methodology in §1.4, which is not reproducible and gives no way to assess whether the chosen 14 algorithms and the cited metrics literature cover the state of the art. This is a load-bearing concern because the survey's contributions—taxonomy, trend identification, and research-gap analysis—are all derived from that corpus. I also verified the reader's noted metric errors: Eq. (3) defines D_C = 1 - cos(φ*,φ) but then gives the expression for cos(φ*,φ), and §6.5's sign interpretation conflicts with the absolute value in Eq. (6). These are concrete correctness defects, but they are localized and correctable. My recommendation is UNCHANGED: the reader already set CONDITIONAL, and my concern supports keeping that verdict rather than rejecting the paper outright, since a revised selection protocol and corrected metric equations could address the issue. I do not see evidence of intent or misconduct, and I am not treating the absence of a documented PRISMA-style flow as fraud—only as a methodological weakness that should be fixed before the survey is treated as authoritative.","tokens_in":23183,"tokens_out":3387,"duration_ms":46901,"concrete_test":"Reproduce the selection protocol from §1.4: run the specified search strings in the five named databases over the years covered by the survey, with two independent reviewers applying the stated 'quality, novelty, relevance' criteria and recording excluded papers. Then recompute the relative prevalence of mitigation strategies (Table 1 categories) and evaluation metrics (Section 6). If the ranking or relative frequencies change by more than 10% when compared with the original selection, the survey's 'trends' conclusions are not robust to corpus choice.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the survey provides a comprehensive, structured overview of fairness in FL, including trends in mitigation strategies and evaluation metrics. For that claim to hold, the surveyed literature must be representative of the state of the art. Section 1.4 describes the selection process only vaguely: search strings \"fair/fairness-aware + Federated Learning/Decentralized Learning\" across ACM, IEEE, Springer, Google Scholar, and ScienceDirect, followed by \"quality, novelty, and relevance\" filtering. It does not report the number of initially retrieved articles, inclusion/exclusion criteria, screening steps, inter-reviewer agreement, a time window, or the list of papers considered but excluded. Without those, the 14 techniques in Table 1 and the metric-frequency observations in Section 6 are not demonstrably representative. Additionally, the metrics chapter contains internal inconsistencies (e.g., Eq. (3) labels a cosine-similarity expression as cosine distance, and the sign interpretation in §6.5 contradicts the absolute value in Eq. (6)). These errors do not by themselves invalidate the entire survey, but they are concrete symptoms of insufficiently rigorous vetting, and the selection method is the load-bearing point because every trend and gap claim inherits its reliability from the corpus chosen.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews fairness in federated learning (FL). It proposes a taxonomy of bias sources (data, client, and model biases), describes fairness notions and their FL-specific adaptations, categorizes mitigation strategies into pre/in/post-processing and into five technique families (optimization formulation, fair resource allocation, reputation/regret-based selection, game-theoretic methods, and gradient-based selection), surveys cross-domain applications, and reviews evaluation metrics used to quantify fairness. The paper's stated contribution is a structured, comprehensive overview of the state of the art, including trends, strengths/limitations of existing methods, and open research directions.","tokens_in":23451,"tokens_out":4873,"duration_ms":58949,"significance":"If the survey's coverage is accepted as representative, it offers a useful structured entry point to fairness-aware FL: the bias taxonomy (Figure 4), the algorithm summary (Table 1), and the catalog of evaluation metrics (Section 6) are potentially valuable for practitioners and newcomers. The manuscript includes no new algorithms or derivations; its value is organizational and critical. The main strengths are the breadth of the literature discussed, the explicit discussion of trade-offs (accuracy, privacy, generalization, utility), and the multi-domain perspective. However, the reliability of the 'comprehensive' and 'trend' claims depends on a literature selection process that is currently underdocumented, and the evaluation-metrics section contains concrete technical errors that need correction before the survey can be used as a dependable reference.","major_comments":[{"comment":"The paper-selection methodology is not reproducible: it does not report search dates, exact query strings, numbers of initially retrieved records, inclusion/exclusion criteria, screening steps, or the list of papers considered but excluded. Since the paper's central claim is a 'comprehensive' overview and since later observations about trends (e.g., the 14 techniques in Table 1 and the metric-adoption statements in Section 6) inherit the representativeness of the undocumented corpus, this is a load-bearing weakness. Please add a full protocol (databases, dates, queries, screening phases, PRISMA-style flow) or substantially temper the comprehensiveness claims.","section":"§1.4"},{"comment":"Equation (3) is internally inconsistent. The text calls cosine similarity 'cosine distance' and writes D_C = 1 − cos(φ*, φ), but the displayed right-hand side is the standard formula for cos(φ*, φ), i.e., the normalized dot product, not 1 − cos(φ*, φ). As printed, the equation implies 1 − cos = cos, which is generally false. Please correct the definition, explicitly state that cosine distance = 1 − cosine similarity, and align the displayed formula with the notation.","section":"§6.2, Eq. (3)"},{"comment":"The sign interpretation in §6.5 contradicts the definition of SPD in Eq. (6). SPD is defined with an absolute value, |P(Ŷ=1|A=0) − P(Ŷ=1|A=1)|, so it cannot take positive or negative values that indicate which group outperforms; the statement that 'positive values in these metrics indicate that the unprivileged group outperforms the privileged group' applies at most to EOD in Eq. (7), which has no absolute value, and not to SPD. Please clarify whether absolute values are intended and restate the interpretation separately for each metric.","section":"§6.4–§6.5, Eqs. (6)–(7)"}],"minor_comments":[{"comment":"The heading '1.1 Federated Learning – Fundamentals and Variants' appears twice; the second occurrence should be renumbered and the subsection hierarchy adjusted accordingly.","section":"§1.1"},{"comment":"Typo: 'performance od training framework' should read 'performance of the training framework'.","section":"§6.1"},{"comment":"In the CGD row, 'Priavte' should be 'Private'. Also, the tick-mark columns (F, A, U, MP, MCT, E) are not defined in the main text; please add a sentence explaining what a tick means.","section":"Table 1"},{"comment":"The notation S_{φ*_i} and S_{φ_i} is confusing: standard deviations are single numbers for the whole vector, not indexed by i. Please define S_{φ*} and S_{φ} as the sample standard deviations and remove the i subscripts.","section":"§6.6, Eq. (8)"},{"comment":"Minor typographical issues: 'Jain s Fairness Index' should be 'Jain's Fairness Index'; the phrase 'a.k.a. cosine distance' in §6.2 should be corrected as noted in the major comments.","section":"§6.5, §6.7"},{"comment":"Grammar: 'However, involves high computational complexity' should be 'However, these approaches involve high computational complexity'.","section":"§4.2.4"},{"comment":"The open-research-direction boxes in Figure 7 are not discussed individually in the text; adding one sentence per direction in Section 7 would improve readability and connect the figure to the prose.","section":"Figure 7"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful survey, but its comprehensiveness claims rest on a non-reproducible literature-selection process. The metric errors in Section 6 are concrete and fixable, and the selection protocol can be added without changing the paper's fundamental scope. I recommend major revision rather than rejection, since the issues are addressable within the manuscript's current structure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a serviceable survey of fairness in FL, not a breakthrough. What it does well: the data/client/model bias trichotomy gives newcomers a clean map, and Table 1 (14 recent algorithms with their objectives, outcomes, and datasets) is genuinely handy. Section 7's discussion of tradeoffs (fairness-accuracy, fairness-privacy, fairness-generalization) is a reasonable synthesis of existing work.\n\nThe soft spots are real but local. Eq (3) defines cosine distance as 1 − cos but then equates it to the dot-product ratio without subtracting from 1; that's mathematically inconsistent. The SPD/EOD section says positive values mean the unprivileged group outperforms the privileged group, but Eq (6) takes the absolute value, so the metric cannot carry a direction. These are typos or conceptual slips that matter in a chapter about metrics. Fixable, but they undermine confidence in that section.\n\nThe larger concern is the selection methodology in §1.4. It reports search strings and sources but no number of initially retrieved papers, inclusion/exclusion criteria, screening steps, or a time window. The final selection is by 'quality, novelty, and relevance,' which is not reproducible. Since the abstract claims a 'comprehensive' overview and Section 6 reports metric-frequency trends, the corpus is load-bearing for those claims. I don't think this invalidates the survey—the taxonomy and technique overview stand on their own—but the trend statements need a transparent corpus or softer language.\n\nOn balance: the paper is a competent consolidation, not a major intellectual advance. It deserves a serious referee—the errors are correctable and the field would benefit from a cleaner version. I'd send it out, with instructions to fix the metric math and add a proper methodology section (or drop the comprehensiveness claims).","headline":"Competent but flawed survey: useful taxonomy and table, but metric errors and a non-reproducible selection method keep it from being fully reliable.","tokens_in":23791,"tokens_out":3954,"would_cite":false,"duration_ms":41834,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey claims that unfairness in federated learning has three root causes—data, client, and model bias—and that existing remedies sort into five strategy families.","keywords":["federated learning","fairness","bias taxonomy","bias mitigation","client selection","fairness metrics","survey"],"falsifier":"Run the survey's own taxonomy as a coding scheme over a complete recent corpus of fairness-aware FL papers and count how many bias sources and mitigation methods fit none of the three or five categories; if a substantial share (more than a few percent) does not fit, the claimed structure is incomplete. A reader could also check whether any major fairness-aware FL algorithm from the same venues is missing from the paper's tables.","tokens_in":23132,"feed_emoji":"⚖️","tokens_out":8269,"duration_ms":92793,"temperature":0.7,"pith_summary":"This survey aims to organize the fairness problem in federated learning (FL) into a single structured picture: where bias enters the training process, what techniques exist to counteract it, and how fairness is measured. It argues that the root causes of unfairness fall into three families—data bias, client bias, and model bias—and that state-of-the-art mitigation strategies can be grouped into optimization-based, resource allocation, reputation/regret-based, game-theoretic, and gradient-based approaches. It also catalogs fairness evaluation metrics and maps fairness applications across edge computing, healthcare, industrial IoT, transport, and wireless networks. A reader would care because fairness failures in FL are what keep the technology from being deployed inclusively: when client selection or aggregation favors some participants, the resulting model skews predictions and reduces trust.","feed_headline":"Bias in federated learning has three roots and five remedy families","feed_subtitle":"New review maps where unfairness enters training and which metrics actually measure it.","key_machinery":"The carrying object is the paper's two-level taxonomy: a three-way split of bias roots (data, client, model) and a five-way split of mitigation strategies (optimization, resource allocation, reputation/regret, game theory, gradients). The taxonomy does the argumentative work by mapping each fairness-aware algorithm to a bias source and a fairness notion—client-level, group, accuracy parity, good-intent, contribution, regret distribution, or expectation—and by making trade-offs visible as structural tensions between fairness and accuracy, privacy, generalization, and utility. The evaluation-metric catalog is the auxiliary mechanism that shows why 'fairness' must be quantified differently depe","core_discovery":"The paper's central claim is that fairness in federated learning is not a single problem but a structured family of problems. It organizes the sources of unfairness into three roots—data bias (collection, distribution, labels, feature skew), client bias (selection, participation, device and communication heterogeneity), and model bias (biased representations, aggregation, algorithmic decisions)—and it organizes the mitigation literature into five families: optimization-problem formulation, fair resource allocation, reputation/regret-based client selection, game-theoretic mechanisms, and gradient-based client selection. The survey further claims that evaluation remains fragmented: fairness is","pith_inferences":["A diagnostic workflow follows directly from the taxonomy but is not spelled out: an FL deployment could classify its fairness failure by root cause and select the matching family; that workflow is testable on real systems.","The paper catalogs trade-offs but not costs; one extension is to plot a cost-fairness frontier for the five families under realistic client heterogeneity, giving practitioners a resource-aware selection rule.","Since the metrics section shows different metrics measure different notions, a standardized fairness report that always includes both a group-parity and a variance-based metric would make results across FL studies comparable; the paper calls for standardization but does not propose concrete report contents.","The noted conflict among fairness notions suggests hybrid strategies—for instance combining game-theoretic contribution rewards with gradient-based reweighting—as a natural next step; the paper leaves this combination unexplored."],"forward_implications":["A practitioner diagnosing an unfair FL system can locate the likely cause by checking which of the three bias roots is active—data distribution, client participation, or model aggregation—and then pick a mitigation family accordingly.","Client selection fairness requires treating two moments separately: the choice of which clients are eligible and the choice of which clients participate each round; a fair algorithm must address both.","Because the five mitigation families have different strengths, no single algorithm dominates; optimal choice depends on whether accuracy, convergence speed, or equity is the priority.","Fairness evaluation should combine group-parity metrics with distributional metrics because each captures a different notion of fairness.","Fairness interventions conflict with accuracy, privacy, generalization, and utility, so fairness-aware FL needs explicit multi-objective design rather than a one-shot constraint."],"supporting_citations":[{"why":"Prior survey of fairness challenges in FL that this paper positions itself as extending.","marker":"[3]"},{"why":"Broad exploration of fairness-aware FL that supplies several fairness notions including client-level, group, and accuracy parity.","marker":"[4]"},{"why":"Examines fairness in relation to privacy, the comparative anchor for the fairness-vs-privacy discussion.","marker":"[5]"},{"why":"FedAvg weighted averaging is identified as a source of aggregation bias and used as a baseline benchmark.","marker":"[17]"},{"why":"q-Fair FL and q-FedAvg objective, a core representative of optimization-based fair resource allocation.","marker":"[33]"},{"why":"Reputation-based client selection (RBCS-F) based on C2MAB and Lyapunov optimization, load-bearing for the reputation/regret strategy family.","marker":"[45]"},{"why":"FairFedCS Lyapunov-based selection that uses Jain's fairness index, connecting client-selection strategy to metrics.","marker":"[51]"},{"why":"Confined gradient descent with fairness-preserving constraints, supporting the optimization family and fairness-utility trade-off.","marker":"[55]"},{"why":"Shapley-value-based contribution fairness (COMFedSV), a representative of the game-theoretic family.","marker":"[59]"},{"why":"FedHEAL parameter-update consistency and fair aggregation, supporting gradient-based client selection.","marker":"[61]"}],"fun_headline_variants":["FL fairness: not one problem but three biases, five fixes","Three bias roots, five remedy families: FL fairness map","Survey: Unfair FL comes from 3 sources, gets 5 fixes","FL fairness: three bias roots, five fixes, but metrics are fragmented","FL fairness isn't one fix—it's three biases and five remedies"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The taxonomy holds only if every important kind of bias falls into one of the three named buckets and if the papers the authors happened to select really represent the whole field.","fun_headline_variants_meta":{"raw":{"variants":["FL fairness: not one problem but three biases, five fixes","Three bias roots, five remedy families: FL fairness map","Survey: Unfair FL comes from 3 sources, gets 5 fixes","FL fairness: three bias roots, five fixes, but metrics are fragmented","FL fairness isn't one fix—it's three biases and five remedies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000972,"raw_usage":{"total_tokens":3946,"prompt_tokens":695,"completion_tokens":3251,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":439,"completion_tokens_details":{"reasoning_tokens":3159}},"tokens_in":439,"tokens_out":3251,"duration_ms":27114,"temperature":1.0,"reasoning_tokens":3159,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T13:10:57.303804+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the survey's own taxonomy as a coding scheme over a complete recent corpus of fairness-aware FL papers and count how many bias sources and mitigation methods fit none of the three or five categories; if a substantial share (more than a few percent) does not fit, the claimed structure is incomplete. A reader could also check whether any major fairness-aware FL algorithm from the same venues is missing from the paper's tables.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"q-Fair FL and q-FedAvg objective, a core representative of optimization-based fair resource allocation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"FairFedCS Lyapunov-based selection that uses Jain's fairness index, connecting client-selection strategy to metrics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Confined gradient descent with fairness-preserving constraints, supporting the optimization family and fairness-utility trade-off."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shapley-value-based contribution fairness (COMFedSV), a representative of the game-theoretic family."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"FedHEAL parameter-update consistency and fair aggregation, supporting gradient-based client selection."}],"review_version":1}