Pith. sign in

REVIEW 3 cited by

FedLLM-Bench: Realistic Benchmarks for Federated Learning of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.04845 v1 pith:2UCQKGWT submitted 2024-06-07 cs.CL cs.AIcs.DCcs.LGcs.MA

classification cs.CLcs.AIcs.DCcs.LGcs.MA
keywords datasetsfedllm-benchfederatedfedllmcommunitydatasetlanguagepreference
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Federated learning has enabled multiple parties to collaboratively train large language models without directly sharing their data (FedLLM). Following this training paradigm, the community has put massive efforts from diverse aspects including framework, performance, and privacy. However, an unpleasant fact is that there are currently no realistic datasets and benchmarks for FedLLM and previous works all rely on artificially constructed datasets, failing to capture properties in real-world scenarios. Addressing this, we propose FedLLM-Bench, which involves 8 training methods, 4 training datasets, and 6 evaluation metrics, to offer a comprehensive testbed for the FedLLM community. FedLLM-Bench encompasses three datasets (e.g., user-annotated multilingual dataset) for federated instruction tuning and one dataset (e.g., user-annotated preference dataset) for federated preference alignment, whose scale of client number ranges from 38 to 747. Our datasets incorporate several representative diversities: language, quality, quantity, instruction, length, embedding, and preference, capturing properties in real-world scenarios. Based on FedLLM-Bench, we conduct experiments on all datasets to benchmark existing FL methods and provide empirical insights (e.g., multilingual collaboration). We believe that our FedLLM-Bench can benefit the FedLLM community by reducing required efforts, providing a practical testbed, and promoting fair comparisons. Code and datasets are available at https://github.com/rui-ye/FedLLM-Bench.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. POPri: Private Federated Learning using Preference-Optimized Synthetic Data

    cs.LG 2025-04 conditional novelty 7.0 of 10

    POPri uses client similarity scores as RL rewards to DPO-tune an LLM for DP synthetic data generation, outperforming prior private evolution baselines on next-token prediction and classification.

  2. Symmetric Pruning of Large Language Models

    cs.LG 2025-01 conditional novelty 4.0 of 10

    SymWanda expresses Wanda and RIA pruning scores as special cases of a symmetric input-output reconstruction objective, and R2-DSnoT adds modest training-free post-pruning gains.

  3. Strategies for Improving Communication Efficiency in Distributed and Federated Learning: Compression, Local Training, and Personalization

    cs.LG 2025-09 conditional novelty 3.0 of 10

    A PhD dissertation showing unified compression theory, personalized accelerated local training, and pruning methods that reduce communication costs in federated learning and maintain accuracy in LLM pruning.

Pith tools