Pith. sign in

REVIEW 4 cited by

FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.12707 v1 pith:B2FU44HF submitted 2024-10-16 cs.DC cs.AIcs.LG

classification cs.DCcs.AIcs.LG
keywords systemtrainingdesigndecentralizeddnnsefficiencygpuschallenges
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To alleviate hardware scarcity in training large deep neural networks (DNNs), particularly large language models (LLMs), we present FusionLLM, a decentralized training system designed and implemented for training DNNs using geo-distributed GPUs across different computing clusters or individual devices. Decentralized training faces significant challenges regarding system design and efficiency, including: 1) the need for remote automatic differentiation (RAD), 2) support for flexible model definitions and heterogeneous software, 3) heterogeneous hardware leading to low resource utilization or the straggler problem, and 4) slow network communication. To address these challenges, in the system design, we represent the model as a directed acyclic graph of operators (OP-DAG). Each node in the DAG represents the operator in the DNNs, while the edge represents the data dependency between operators. Based on this design, 1) users are allowed to customize any DNN without caring low-level operator implementation; 2) we enable the task scheduling with the more fine-grained sub-tasks, offering more optimization space; 3) a DAG runtime executor can implement RAD withour requiring the consistent low-level ML framework versions. To enhance system efficiency, we implement a workload estimator and design an OP-Fence scheduler to cluster devices with similar bandwidths together and partition the DAG to increase throughput. Additionally, we propose an AdaTopK compressor to adaptively compress intermediate activations and gradients at the slowest communication links. To evaluate the convergence and efficiency of our system and algorithms, we train ResNet-101 and GPT-2 on three real-world testbeds using 48 GPUs connected with 8 Mbps~10 Gbps networks. Experimental results demonstrate that our system and method can achieve 1.45 - 9.39x speedup compared to baseline methods while ensuring convergence.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Decentralised AI Training and Inference with BlockTrain

    cs.AI 2026-06 conditional novelty 6.0 of 10

    BlockTrain partitions models into block-local diffusion objectives that train near end-to-end WikiText quality with one-block worker memory, real WAN transport, and one-sweep distributed serving.

  2. Decentralised AI Training and Inference with BlockTrain

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    BlockTrain partitions models into blocks trained on local objectives, reaching CE 1.359 on WikiText within 0.04 of end-to-end baseline while enabling distributed training and inference over TCP for up to 75B-parameter models.

  3. Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training

    cs.DC 2026-05 unverdicted novelty 5.0 of 10

    BACE-Pipe is a bandwidth-aware and cost-efficient pipeline scheduling framework for geo-distributed LLM training that reduces average JCT by 27.9-64.7% and electricity cost by 12.6-30.6% in simulations versus baselines.

  4. ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training

    cs.DC 2026-05 unverdicted novelty 4.0 of 10

    ScaleAcross Explorer jointly optimizes three design dimensions for scale-across training and reports up to 64.62% speedups over production baselines and 37.59% over prior art in testbed and simulation experiments.

Pith tools