{"id":"e699ab1b-512d-4327-b91e-15a36dd8af65","arxiv_id":"2505.08830","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A review of federated large language models that organizes current methods into feasibility, robustness, security, and future research directions.","lead":"This paper surveys research that combines large language models with federated learning, a distributed training setup that keeps private data on users' devices. It sorts existing work into feasibility, robustness, security, and future directions, and points to areas that need more study.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The DLG attack is misattributed to 'Mu et al.' and FeS is cited twice as [13]/[14]; a citation audit is needed before this survey's taxonomy can be treated as reliable.","rationale":"The reader's weakest assumption is that the selected references are comprehensive and representative. That is a valid concern, and no search protocol is reported, so completeness cannot be independently assessed. My stress-test goes one step further: even within the cited set, there are concrete, verifiable reporting errors. The DLG misattribution and the duplicate FeS reference indicate that the paper's source handling is not reliable enough to support the 'accurate, comprehensive review' claim without correction. This reinforces the reader's CONDITIONAL verdict rather than overturning it: these are fixable issues, and the survey's structure and breadth still have value. I do not see an internally inconsistent central argument that would justify rejection, and the survey does not make novel empirical claims requiring formal verification. The main risk is that readers will cite the taxonomy and gap analysis as authoritative while some underlying references are misreported or double-counted. A full citation audit is the direct check that would settle whether these are isolated slips or symptomatic of broader unreliability.","tokens_in":36100,"tokens_out":3713,"duration_ms":39292,"concrete_test":"Perform a citation audit: for every claim tied to a numbered reference in Sections 2 through 5, verify author names, year, venue, and that the claim matches the cited paper. Specifically check [157] (should be Zhu et al., NeurIPS 2019, not 'Mu et al.') and [13]/[14] (should be one shared MobiCom 2023 FeS entry). If the error rate among sampled citations exceeds roughly 5%, the survey needs revision and re-review before its taxonomy and research-gap conclusions are used as authoritative.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this is an accurate, comprehensive review of FLLM. For a survey, the load-bearing condition is that the cited literature is selected systematically and reported faithfully. That condition is already visibly violated. In Section 4.1.2, the text attributes the deep leakage from gradient (DLG) algorithm to 'Mu et al. [157]', but reference [157] is Zhu, Liu, and Han, 'Deep Leakage from Gradients', NeurIPS 2019. This is a clear author and citation misattribution in a core security discussion. In addition, the FeS framework is cited as [13] in Table 2 (PEFT methods) and as [14] in Section 5.1 (few-shot learning), but both references are the same Cai et al. MobiCom 2023 paper. This duplication inflates the apparent number of distinct FLLM contributions. These are not cosmetic problems: they show that the reference base has not been systematically verified, so the taxonomy and the claim that robustness and security are underdeveloped relative to feasibility may rest on unreliable evidence. The absence of any reported search protocol, inclusion criteria, or date range further prevents a reader from assessing whether omitted work would change the gap analysis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of federated large language models (FLLM). It organizes the literature into four perspectives: feasibility (fine-tuning methods, including full-parameter, PEFT, prompt tuning, and other techniques), robustness (resource, data, and task heterogeneity), security (privacy leakage, poisoning, backdoor attacks, and defenses), and future directions (few-shot learning, unlearning, and IP protection). The paper claims to be an exhaustive survey of recent FLLM research, and its main thesis is that feasibility has received the most attention while robustness and security remain underdeveloped.","tokens_in":36353,"tokens_out":1908,"duration_ms":19194,"significance":"If the reference base were reliable, this survey would provide a useful structured entry point to a fast-growing field. The four-perspective taxonomy is sensible and the paper covers a broad set of recent methods, including many 2023–2024 papers. The authors also make a defensible high-level observation that heterogeneity and security issues are less thoroughly studied than feasibility. The paper does not present new algorithms or experiments, but surveys can be valuable contributions when they organize and assess the literature faithfully. That value is currently compromised by citation inaccuracies and the absence of a documented selection methodology, as detailed below.","major_comments":[{"comment":"The text states: 'Mu et al. [157] studied a gradient-based reconstruction attack algorithm ... they proposed an algorithm called deep leakage from gradient (DLG)'. However, reference [157] is Zhu, Liu, and Han, 'Deep Leakage from Gradients', NeurIPS 2019. The DLG algorithm is authored by Zhu et al., not 'Mu et al.' This is not a cosmetic typo: it misattributes a foundational attack in the core security discussion of the survey, and it indicates that the reference list has not been systematically verified against the cited content.","section":"Section 4.1.2"},{"comment":"The same paper, Cai et al., 'Federated Few-Shot Learning for Mobile NLP' (MobiCom 2023), is cited as [13] in Table 2 and as [14] in Section 5.1. This duplication inflates the apparent number of distinct FLLM contributions and confuses the taxonomy, since the reader cannot tell whether FeS is being listed twice or two different works are being referenced. The reference list must be de-duplicated and all such occurrences reconciled.","section":"Table 2 and Section 5.1"},{"comment":"The abstract promises 'an exhaustive survey' and Section 1 states that 'our work surveys the latest FLLM research,' but the manuscript does not report any literature search protocol, database choices, inclusion criteria, or date range. Without this information, the representativeness of the cited literature cannot be assessed, and the paper's central gap analysis (that robustness and security are underdeveloped relative to feasibility) may reflect the authors' selection rather than the actual state of the field. The authors should add a methodology subsection describing how the literature was collected and filtered.","section":"Section 1 and Abstract"}],"minor_comments":[{"comment":"The caption contains a typo: 'Full-paramete' should be 'Full-parameter'.","section":"Figure 1"},{"comment":"The symbol '✔–' is used in the comparison table but is not defined in the table footnote; please define it explicitly to avoid ambiguity with '✓'.","section":"Table 1"},{"comment":"The sentence 'This generalization ability enables the model to perform better when facing new tasks' is a reasonable claim but no citation is provided; adding a reference to multi-task learning literature would strengthen the statement.","section":"Section 3.3"},{"comment":"The description of the 'analysis-based attacks' category cites [82] and [155], but the text says these methods 'solve a system of linear equations' without elaborating on the specific mechanism; a brief explanation or example would improve readability.","section":"Section 4.1.2"}],"recommendation":"major_revision","confidential_remarks":"The citation problems uncovered here are not isolated. The DLG misattribution and the [13]/[14] duplication suggest that the reference base has not been checked uniformly, and the absence of a search protocol further weakens the paper's reliability as a survey. I recommend a full citation audit and the addition of a methodology section before the paper can be accepted. The taxonomy itself is potentially valuable, so I do not recommend rejection, but the current version is not ready for publication in a serious venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a useful survey, not a research result. It maps the FLLM literature into feasibility, robustness, security, and future directions, and it does a service by pulling together 2024 work and by calling out task heterogeneity as a distinct problem. The comparative table of prior surveys is a nice touch, and the section on security organizes a messy literature in a way newcomers will find approachable.\n\nThe soft spots are real, and for a survey they matter. First, citation precision: in Section 4.1.2, DLG is attributed to 'Mu et al. [157]' but reference [157] is Zhu, Liu, and Han. That is the canonical gradient-inversion paper; getting it wrong in a security survey is not cosmetic. Second, the same Cai et al. MobiCom 2023 paper appears as [13] in Table 2 and [14] in Section 5.1, which inflates the count of distinct FLLM contributions. Third, while reading the long-tailed data leakage part I noticed a citation to [23], a GPT-4 clinical diagnosis paper, which does not appear to support the claim about over-memorization. Together these say the reference list has not been systematically checked.\n\nThe absence of a reported search protocol is a smaller issue but still real: the abstract promises an 'exhaustive survey,' and without inclusion criteria or date range a reader cannot judge whether the gap analysis would survive omitted work. That said, the overall organization is sound, most method summaries match their sources, and the central message that robustness and security are underdeveloped relative to feasibility is plausible and consistent with the papers cited.\n\nWho is this for? A graduate student starting FLLM research, or a reviewer needing a map of the area. I would not cite it as an authoritative reference until the citation issues are fixed. It deserves serious peer review, not desk rejection, but the revision should be conditional on a full citation audit and a short methodology paragraph.","headline":"A useful but sloppy survey: the taxonomy and gap analysis are worth reading, but the citation errors and missing search protocol mean it needs revision before I'd trust it as a reference.","tokens_in":36836,"tokens_out":2513,"would_cite":false,"duration_ms":26614,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey organizes federated large language models into four questions—feasibility, robustness, security, and future directions—and concludes that the open problems are robustness and security.","keywords":["federated learning","large language models","federated large language models","parameter-efficient fine-tuning","heterogeneity","privacy","security","machine unlearning"],"falsifier":"Run a systematic literature search with a stated query and date range that counts papers per taxonomy cell: if the data- and task-heterogeneity cells and the security cell turn out to contain a substantial body of FLLM-specific work omitted here—for example, more than a dozen studies of FLLM-specific gradient or poisoning attacks—then the survey's central gap claim would need revision.","tokens_in":35941,"feed_emoji":"🔐","tokens_out":8524,"duration_ms":79932,"temperature":0.7,"pith_summary":"This paper is a survey of federated large language models (FLLM), the practice of training or fine-tuning an LLM across many data holders without centralizing their private data. Its organizing thesis is that the field can be mapped onto four questions: whether FLLM can be done efficiently (feasibility), whether it survives differences among participants and their data (robustness), whether it resists attackers and protects privacy (security), and what remains to be done (future directions). The survey's central claim is that feasibility is largely demonstrated in academic settings through parameter-efficient fine-tuning, while robustness and security are the underdeveloped parts: data and task heterogeneity have few solutions, and most security work is concentrated on backdoor attacks. The practical stake is that FLLM is viable as a research program, but its deployment will be decided by reliability and safety rather than by the ability to train at all.","feed_headline":"Survey: federated LLMs are feasible, but robustness and security lag","feed_subtitle":"A four-axis review finds most work on efficient fine-tuning, while heterogeneity and attacks stay under-explored.","key_machinery":"The paper's machinery is a four-axis taxonomy—feasibility, robustness, security, and future directions—used to sort the FLLM literature, with its overview figure using darker circles to encode how many studies each sub-topic has. The feasibility axis is itself organized around Federated Parameter-Efficient Fine-Tuning (FedPEFT), which freezes most of the LLM and trains only small modules such as low-rank adapters (LoRA, where a weight update is written as $\\Delta W = B A$ with low-rank $B$ and $A$) or inserted adapters; this is the mechanism that makes client-side fine-tuning affordable. Robustness is organized around three kinds of heterogeneity—resource, data, and task—and security around a list of attacks (membership inference, data reconstruction, jailbreaking, prompt injection, long-tailed data leakage, poisoning, backdoors) and defenses. The taxonomy does the work: it lets the authors claim that most publications sit in the feasibility cell and that robustness and security cells are comparatively empty.","core_discovery":"The paper's discovery is a structured map of the FLLM literature. It finds that the field is uneven: most published work addresses feasibility, using full-parameter fine-tuning, parameter-efficient fine-tuning (PEFT), prompt tuning, model compression, split learning, and zeroth-order optimization to bring the cost of client-side training down; the paper reads this as showing that FLLM is achievable in the lab, though still far from ordinary devices. Robustness work exists but clusters around resource heterogeneity—clients with different compute, memory, and bandwidth—with comparatively little on data heterogeneity or the task heterogeneity that arises when clients run different NLP tasks on one shared model. On security, the paper reports that attacks and defenses specific to FLLM are scarce, that certain FL and LLM attacks do not transfer unchanged (for example, membership inference is often close to random for large models), and that backdoors are the most studied threat. Its closing claim is that the field's next phase should concentrate on robustness and security, plus the new demands of few-shot learning, machine unlearning, and intellectual-property protection.","pith_inferences":["The paper does not draw this consequence, but its own density counts imply a testable prediction: publication volume in the data-heterogeneity and security cells should grow faster than volume in the feasibility cell over the next few years, and a citation analysis could check that.","The survey's separation of robustness from security invites a cross-cutting reading that is left implicit: heterogeneity methods that split the model into per-client low-rank modules also change the attack surface, so aggregation of different ranks or adapters may be harder to poison yet easier to fingerprint.","Long-tailed data leakage, treated here as a privacy risk, could equally be studied as a robustness risk: a model that memorizes rare client data is both leaking and failing to generalize, so privacy and robustness defenses may converge on the same mechanisms of regularization, noise, or selective forgetting.","If membership inference is genuinely weak for large LLMs because training uses massive data and few epochs, then privacy scrutiny for FLLM may shift toward reconstruction and jailbreaking attacks; that is an inference about where limited defense effort should go, not a claim the paper makes."],"forward_implications":["If FLLM feasibility is as mature as the survey says, new systems research should stop re-deriving parameter-efficient fine-tuning from scratch and instead benchmark against existing FedPEFT baselines.","If robustness is dominated by resource-heterogeneity work, then data and task heterogeneity are the highest-leverage unsolved problems, with the cited adapter- and LoRA-based methods serving as early entries rather than settled solutions.","If security research is sparse and backdoor-centric, then defenses such as robust aggregation, frequency-domain detection, and distribution-divergence checks are not yet a complete answer, and deployments should assume a residual attack surface.","If the future directions are few-shot learning, machine unlearning, and IP protection, then FLLM's next phase will be judged by data efficiency, forgetfulness, and ownership verification rather than by raw accuracy alone.","If the integration inherits risks from both FL and LLMs, then privacy must be argued per attack rather than assumed from the federated setting, since gradient and embedding leakage can reconstruct text."],"supporting_citations":[{"why":"Defines LoRA, the low-rank update that most feasibility and heterogeneity methods build on.","marker":"[34]"},{"why":"Introduces the FedPEFT concept that organizes the feasibility taxonomy.","marker":"[71]"},{"why":"FlexLoRA, evidence for the claim that dynamic client-side LoRA ranks address resource and task heterogeneity.","marker":"[6]"},{"why":"FedLoRA, the model-heterogeneous LoRA framework that many robustness methods extend.","marker":"[131]"},{"why":"FedDAT, a task-heterogeneity solution supporting the claim that multi-task FLLM is being tackled via adapters and distillation.","marker":"[19]"},{"why":"FedSecurity, the attack-and-defense benchmark that grounds the security section's claim that FLLM defenses are underdeveloped.","marker":"[32]"},{"why":"DAGER, the gradient inversion attack showing exact text reconstruction from FLLM gradients.","marker":"[80]"},{"why":"Decepticons, the corrupted-transformer attack showing text and embedding leakage in federated language models.","marker":"[27]"},{"why":"Large-scale membership-inference results on language models, used to argue that MIA is weak for large models.","marker":"[109]"},{"why":"FedIT-U2S, the few-shot unstructured-to-structured instruction data method anchoring the few-shot future direction.","marker":"[130]"}],"fun_headline_variants":["Federated LLMs: feasible but security lags","FLLM survey: robustness and security under-explored","Feasibility leads FLLM research, security lags","Federated LLM review: robustness and security need focus","Survey: federated LLMs work, but not yet secure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's conclusions stand or fall on whether its chosen references are a fair and complete sample of the FLLM literature, since no search protocol, database choices, inclusion criteria, or date range are reported.","fun_headline_variants_meta":{"raw":{"variants":["Federated LLMs: feasible but security lags","FLLM survey: robustness and security under-explored","Feasibility leads FLLM research, security lags","Federated LLM review: robustness and security need focus","Survey: federated LLMs work, but not yet secure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1258,"prompt_tokens":980,"completion_tokens":278,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":596,"completion_tokens_details":{"reasoning_tokens":194}},"tokens_in":596,"tokens_out":278,"duration_ms":3031,"temperature":1.0,"reasoning_tokens":194,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:00:21.871154+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a systematic literature search with a stated query and date range that counts papers per taxonomy cell: if the data- and task-heterogeneity cells and the security cell turn out to contain a substantial body of FLLM-specific work omitted here—for example, more than a dozen studies of FLLM-specific gradient or poisoning attacks—then the survey's central gap claim would need revision.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"FedIT-U2S, the few-shot unstructured-to-structured instruction data method anchoring the few-shot future direction."}],"review_version":1}