pLouvain and pLeiden, two GPU parallelizations, speed up Louvain and Leiden clustering by 3.1x and 8.8x, and pLeiden's spanning-tree refinement claims to preserve all of sequential Leiden's quality guarantees.
CPU vs. GPU for Community Detection: Performance Insights from GVE-Louvain and $\nu$-Louvain
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Community detection involves identifying natural divisions in networks, a crucial task for many large-scale applications. This report presents GVE-Louvain, one of the most efficient multicore implementations of the Louvain algorithm, a high-quality method for community detection. Running on a dual 16-core Intel Xeon Gold 6226R server, GVE-Louvain outperforms Vite, Grappolo, NetworKit Louvain, and cuGraph Louvain (on an NVIDIA A100 GPU) by factors of 50x, 22x, 20x, and 5.8x, respectively, achieving a processing rate of 560M edges per second on a 3.8B-edge graph. Additionally, it scales efficiently, improving performance by 1.6x for every thread doubling. The paper also presents $\nu$-Louvain, a GPU-based implementation. When evaluated on an NVIDIA A100 GPU, $\nu$-Louvain performs only on par with GVE-Louvain, largely due to reduced workload and parallelism in later algorithmic passes. These results suggest that CPUs, with their flexibility in handling irregular workloads, may be better suited for community detection tasks.
citation-role summary
citation-polarity summary
fields
cs.DC 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
GPU-Accelerated Multilevel Graph Clustering: A Parallel Perspective on Louvain and Leiden
pLouvain and pLeiden, two GPU parallelizations, speed up Louvain and Leiden clustering by 3.1x and 8.8x, and pLeiden's spanning-tree refinement claims to preserve all of sequential Leiden's quality guarantees.