REVIEW 4 major objections 6 minor 1 cited by
CPU vs. GPU for Community Detection: Performance Insights from GVE-Louvain and $\nu$-Louvain
T0 review · 4 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read GVE-Louvain, a multicore CPU implementation of the Louvain algorithm, outperforms established CPU and GPU implementations by 20–50×, reaching 560M edges/s on a 3.8-billion-edge graph.
desk verdict Useful engineering report with a new GPU Louvain and detailed pseudocode, but the central speedup claim is internally inconsistent (5.8x vs 3.2x) and the tuning is not held out. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key mechanism is the pair of per-thread collision-free hash tables (a keys list with a full-size values array, placed far apart in memory to avoid false sharing) used in GVE-Louvain's local-moving and aggregation phases, combined with preallocated CSR data structures and parallel prefix sums to build the super-vertex graph without repeated allocation. On the GPU side, the machinery is per-vertex open-addressing hash tables sized to twice the degree, using a hybrid quadratic-double probing scheme and a Pick-Less rule that breaks symmetric community-swap cycles. Both implementations use asynchronous parallel vertex moves, vertex pruning, a 20-iteration cap per pass, threshold scaling with drop rate 10, and an aggregation tolerance of 0.8.
What would settle it
Run GVE-Louvain on a graph not in Table 2 using the exact published parameters, and compare against cuGraph Louvain (with RMM pool) and Grappolo on the same hardware; if the speedup over cuGraph falls far below the reported 3.2–5.8×, or the 560M edges/s rate is not reproducible on sk-2005 with the stated hardware, the general claim would be falsified.
Extended reading notes
Core claim
The paper's central claim is that the Louvain algorithm's performance is gated not only by the local-moving phase, which most prior work optimizes, but also by the aggregation phase; GVE-Louvain addresses both with collision-free per-thread hash tables, CSR-based aggregation with parallel prefix sums, vertex pruning, and tolerance tuning. On a dual 16-core Intel Xeon Gold 6226R, it processes 560M edges/s on the 3.8B-edge sk-2005 web graph and scales 1.6× per thread doubling. The GPU port, ν-Louvain, uses per-vertex open-addressing hash tables and a Pick-Less swap-prevention mechanism; on an A100 it is only 1.03× faster on average than GVE-Louvain, has 0.5% lower modularity, and cannot process the largest graph due to memory limits. The paper concludes from this that multicore CPUs may be better suited for community detection than GPUs.
Load-bearing premise
The reported speedups rest on parameter settings (iteration cap, tolerances, switch degrees, probing strategy) that were tuned on the same 13 graphs later used to measure the speedups; if that tuning is overfitted, the results will not transfer to new graphs.
Editorial extensions
If this is right
- If the measurements hold, GVE-Louvain is the fastest reported multicore Louvain implementation, and other implementations can adopt its aggregation-phase optimizations.
- The 1.6× per-thread-doubling scaling means the algorithm remains efficient up to 32 threads on this dual-socket server.
- The near-parity of ν-Louvain with GVE-Louvain implies that for this workload, GPU memory capacity and reduced parallelism in later passes offset GPU throughput advantages.
- The aggregation phase, not just local-moving, must be optimized to make Louvain fast; this is a direct engineering lesson.
Reading between the lines
- The cuGraph speedup is averaged only over the eight graphs where cuGraph ran without out-of-memory; a GPU implementation with batched processing could perform differently, so the CPU-versus-GPU conclusion is specific to the compared implementations and graphs.
- Because parameters were selected by experiments on the same graphs used for the final numbers, a held-out evaluation would be needed to confirm the speedups generalize; this is my own assessment, not the paper's.
- The conclusion that CPUs 'may be better suited' for community detection is an inference from one pair of implementations; the paper presents no energy or cost-per-edge comparison, which would be needed to fully support that practical conclusion.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents two Louvain implementations: GVE-Louvain, a multicore OpenMP code, and ν-Louvain, a CUDA GPU code. The headline claim is that GVE-Louvain is the fastest or one of the fastest Louvain implementations: the abstract and Table 1 report speedups of 50×, 22×, 20×, and 5.8× over Vite, Grappolo, NetworKit, and cuGraph, with a processing rate of 560M edges/s on the 3.8B-edge sk-2005 graph, while Section 5.2.1 and the Conclusion report 3.2× instead of 5.8× for cuGraph. ν-Louvain is reported to be only on par with GVE-Louvain (1.03× speedup), which is taken to suggest that CPUs may be better suited for community detection. The evaluation uses 13 SuiteSparse graphs, includes phase/pass breakdowns, and reports strong scaling.
Significance. If correct, GVE-Louvain would be a leading practical multicore Louvain implementation, and the comparison would be a useful data point against the common assumption that GPUs dominate graph analytics. The paper provides detailed pseudocode (Algorithms 1–7), public repository links, and per-graph runtime/modularity figures, which are concrete strengths. The analysis of phase-level bottlenecks and the observation that later Louvain passes are poorly parallelized on the GPU are useful and likely to transfer to other Louvain implementations. However, the central quantitative claims currently rest on an internally inconsistent headline number, a parameter-selection protocol that uses the same graphs later used for evaluation, and a one-implementation-per-architecture comparison. These issues are close enough to the paper's main message that they must be fixed before the results can be taken at face value.
major comments (4)
- [Abstract, Table 1, §5.2.1, Fig. 11(b), Conclusion] The cuGraph speedup is internally inconsistent: the abstract and Table 1 report 5.8×, while §5.2.1, Figure 11(b), and the Conclusion report 3.2×. No explanation or reconciliation is given, and the averaging method (arithmetic vs. geometric mean, and which graphs are included) is not specified. Because this number is part of the central CPU-versus-GPU claim, the paper as written cannot be verified. The speedup should be recomputed, the averaging rule stated, and the abstract/table aligned with the body. The comparison is also incomplete: cuGraph fails with out-of-memory on arabic-2005, uk-2005, webbase-2001, it-2004, and sk-2005, so the average is computed only over the remaining 8 graphs.
- [§4.1, §4.3, Table 2, §5.2] The parameter settings that drive the headline results are tuned and evaluated on the same graphs. Section 4.1 (OpenMP schedule, maximum iterations, tolerance drop, initial tolerance, aggregation tolerance) and Section 4.3 (switch degrees, Pick-Less step, probing strategy, hashtable value type) describe experiments performed on 'large graphs from Table 2', and the final speedups in Section 5.2 are then reported on those same Table 2 graphs. There is no held-out set, no cross-validation, and no sensitivity analysis for the chosen values. If the selected settings are overfit to these 13 inputs, the reported speedups and the CPU/GPU conclusion will not transfer to new graphs. Please add a validation protocol (for example, tuning on a subset and testing on a disjoint subset) or demonstrate robustness of the conclusions to parameter perturbations.
- [§5.2.1, §5.2.3, Fig. 11, Fig. 13] The architectural conclusion is drawn from a lopsided comparison. cuGraph Louvain cannot run on the five largest web graphs, so on exactly those graphs where GVE-Louvain looks best (e.g., 560M edges/s on sk-2005) there is no GPU baseline at all. In addition, the statement that CPUs 'may be better suited for community detection' is based primarily on GVE-Louvain versus ν-Louvain, two implementations by the same author, plus cuGraph and Nido, both of which also fail or underperform on the largest graphs. Implementation maturity and engineering effort are not controlled across platforms. The data support a claim about these particular implementations, not about CPU and GPU architectures generally. Please limit the conclusion accordingly or broaden the baseline set.
- [§5.2.3, §5.3.2, Conclusion] The 'on par' characterization of ν-Louvain is not adequately qualified. The average speedup over GVE-Louvain is only 1.03×, with per-graph values ranging from 0.6× to 7.1×, and on kmer_V1r the modularity falls from 0.9437 (GVE-Louvain) to 0.8722 (ν-Louvain) — a substantial quality loss that is acknowledged in §5.2.2 but omitted from the unqualified 'on par' statement in the abstract and conclusion. Please report confidence intervals and per-graph quality differences, and state explicitly that 'on par' refers only to runtime and only on the graphs where the comparison was possible.
minor comments (6)
- [§5.2.2] The text says 'Nido (Louvain) [45]', but reference [45] is the author's GVE-Louvain report; the Nido paper is reference [10]. This appears to be a citation error and should be fixed.
- [§4.3.4] The sentence 'we adopt a thread-per-vertex approach how low-degree vertices [63]' appears to be missing a word (likely 'for') and is not grammatical.
- [§5.1.1] The description '1 MB L1 cache per core' is not plausible for the Intel Xeon Gold 6226R; please verify the cache hierarchy details (the Gold 6226R has 32 KB L1 data and 32 KB L1 instruction per core).
- [§5.2.1] The text provides average speedups '50×, 22×, 20×, and 3.2×' without stating the mean type or the graph subset used for the average; a small table with per-graph ratios and the averaging method would make the results reproducible.
- [Conclusion] The Conclusion states that GVE-Louvain is 'the most efficient implementation on multicore CPUs', which is stronger than the Abstract's 'one of the most efficient'; given that only four baselines are compared, the stronger claim is not supported. Please use the qualified phrasing.
- [Figures 2, 5, 7–13] The text says each measurement was repeated five times, but no error bars or variance information is shown. Adding standard deviations or equivalent would help the reader judge whether the 1–13% differences used for parameter choices are meaningful.
Circularity Check
No circular derivation is present: the speedup claims are measured against external baselines, and the self-citations are minor and not load-bearing.
full rationale
The report's central claims are empirical performance ratios (GVE-Louvain versus Vite, Grappolo, NetworKit, and cuGraph) obtained by running each implementation on the same SuiteSparse graphs and averaging runtimes, as described in Section 5.2.1. These ratios are arithmetic comparisons of independently measured times; no equation defines the speedup in terms of the paper's own inputs, and no fitted parameter is relabeled as a prediction. Section 4.1 does tune OpenMP scheduling, iteration caps, tolerances, pruning, CSR aggregation, and hashtable layout using the same Table 2 graphs that later appear in the Section 5 evaluation. This creates a test-set reuse / overfitting risk, which is a correctness threat, but it is not circularity: the reported runtimes are still measured outcomes, not consequences forced by construction from the tuning choices. The paper cites the author's prior GVE-Louvain [45] and nu-LPA [46] for specific techniques (e.g., 'like LPA [46]' in Section 4.3.1 and 'nextPow2(Di)-1 [46]' in Section 4.3.2). These self-citations are not load-bearing for the central comparison because the manuscript provides complete pseudocode in Algorithms 1-7 and measures against external implementations, so the conclusion does not reduce to an unverified self-citation chain. Finally, the discrepancy between the abstract and Table 1 reporting a 5.8x speedup over cuGraph and Section 5.2.1 and the conclusion reporting 3.2x is an internal inconsistency that affects verifiability of the headline claim, but it is a reporting error rather than a circular argument. The only notable feature under the circularity rubric is the presence of minor, non-load-bearing self-citations, justifying a score of 2 rather than 0.
Assumptions & free parameters
free parameters (9)
- OpenMP loop schedule =
dynamic, chunk size 2048
- Max iterations per local-moving pass =
20
- Tolerance drop rate =
10
- Initial tolerance tau =
0.01
- Aggregation tolerance tau_agg =
0.8
- GPU switch degree, local-moving phase =
64
- GPU switch degree, aggregation phase =
128
- Pick-Less step rho =
4 iterations
- Hashtable value datatype =
32-bit float
assumptions (5)
- domain assumption Modularity Q (Eq. 1) and delta-modularity (Eq. 2) are valid measures of community quality and the right objective for comparison.
- domain assumption The asynchronous, parallel greedy local-moving phase with atomic updates converges to a useful near-optimal modularity partition.
- ad hoc to paper Pick-Less mode, applied every 4 iterations, prevents community-swap cycles without materially reducing community quality.
- domain assumption The 13 SuiteSparse graphs are representative of community detection workloads generally.
- ad hoc to paper One dual-Xeon CPU server and one A100 GPU are sufficient hardware instances to infer that CPUs are better suited for community detection.
Cite this review
Pith. "Pith review of CPU vs. GPU for Community Detection: Performance Insights from GVE-Louvain and $\nu$-Louvain." pith.science (2026). https://pith.science/paper/BG3SBUID
@misc{pith2026250119004,
author = {Pith},
title = {Pith review of: CPU vs. GPU for Community Detection: Performance Insights from GVE-Louvain and $\nu$-Louvain},
year = {2026},
howpublished = {\url{https://pith.science/paper/BG3SBUID}},
note = {Machine review of arXiv:2501.19004}
}
abstract
Community detection involves identifying natural divisions in networks, a crucial task for many large-scale applications. This report presents GVE-Louvain, one of the most efficient multicore implementations of the Louvain algorithm, a high-quality method for community detection. Running on a dual 16-core Intel Xeon Gold 6226R server, GVE-Louvain outperforms Vite, Grappolo, NetworKit Louvain, and cuGraph Louvain (on an NVIDIA A100 GPU) by factors of 50x, 22x, 20x, and 5.8x, respectively, achieving a processing rate of 560M edges per second on a 3.8B-edge graph. Additionally, it scales efficiently, improving performance by 1.6x for every thread doubling. The paper also presents $\nu$-Louvain, a GPU-based implementation. When evaluated on an NVIDIA A100 GPU, $\nu$-Louvain performs only on par with GVE-Louvain, largely due to reduced workload and parallelism in later algorithmic passes. These results suggest that CPUs, with their flexibility in handling irregular workloads, may be better suited for community detection tasks.
Figures
Figures from the paper (12 more)
Forward citations
Cited by 1 Pith paper
-
GPU-Accelerated Multilevel Graph Clustering: A Parallel Perspective on Louvain and Leiden
pLouvain and pLeiden, two GPU parallelizations, speed up Louvain and Leiden clustering by 3.1x and 8.8x, and pLeiden's spanning-tree refinement claims to preserve all of sequential Leiden's quality guarantees.
Reference graph
Works this paper leans on
-
[45]
Subhajit Sahu. 2023. GVE-Louvain: Fast Louvain Algorithm for Community Detection in Shared Memory Setting. arXiv preprint arXiv:2312.04876 (2023)
work page Pith review arXiv 2023
-
[1]
Emmanuel Abbe. 2018. Community detection and stochastic block models: recent developments. Journal of Machine Learning Research 18, 177 (2018), 1–86
work page 2018
-
[2]
A. Aldabobi, A. Sharieh, and R. Jabri. 2022. An improved Louvain algorithm based on Node importance for Community detection. Journal of Theoretical and Applied Information Technology 100, 23 (2022), 1–14
work page 2022
-
[3]
Yuhe Bai, Camelia Constantin, and Hubert Naacke. 2024. Leiden-Fusion Parti- tioning Method for Effective Distributed Training of Graph Embeddings. InJoint European Conference on Machine Learning and Knowledge Discovery in Databases . Springer, 366–382
work page 2024
-
[4]
Joel J Bechtel, William A Kelley, Teresa A Coons, M Gerry Klein, Daniel D Slagel, and Thomas L Petty. 2005. Lung cancer detection in patients with airflow obstruction identified in a primary care outpatient practice. Chest 127, 4 (2005), 1140–1145. 13 Subhajit Sahu Phase split (%) 0% 25% 50% 75% 100% indochina-2004 uk-2002 arabic-2005 uk-2005 webbase-2001...
work page 2005
-
[5]
A. Bhowmick, S. Vadhiyar, and V. PV. 2022. Scalable multi-node multi-GPU Louvain community detection algorithm for heterogeneous architectures. Con- currency and Computation: Practice and Experience 34, 17 (2022), 1–18
work page 2022
-
[6]
A. Bhowmik and S. Vadhiyar. 2019. HyDetect: A Hybrid CPU-GPU Algorithm for Community Detection. In IEEE 26th International Conference on High Performance Computing, Data, and Analytics (HiPC) . IEEE, Goa, India, 2–11
work page 2019
-
[7]
V. Blondel, J. Guillaume, R. Lambiotte, and E. Lefebvre. 2008. Fast unfolding of communities in large networks. Journal of Statistical Mechanics: Theory and Experiment 2008, 10 (Oct 2008), P10008
work page 2008
Show all 71 references
-
[8]
Brandes, D
U. Brandes, D. Delling, M. Gaertler, R. Gorke, M. Hoefer, Z. Nikoloski, and D. Wagner. 2007. On modularity clustering. IEEE transactions on knowledge and data engineering 20, 2 (2007), 172–188
2007
-
[9]
Cheong, H
C. Cheong, H. Huynh, D. Lo, and R. Goh. 2013. Hierarchical Parallel Algorithm for Modularity-Based Community Detection Using GPUs. In Proceedings of the 19th International Conference on Parallel Processing (Aachen, Germany) (Euro-Par’13). Springer-Verlag, Berlin, Heidelberg, 775–787
2013
-
[10]
Han-Yi Chou and Sayan Ghosh. 2022. Batched Graph Community Detection on GPUs. In Proceedings of the International Conference on Parallel Architectures and Compilation Techniques. 172–184
2022
-
[11]
Aaron Clauset, Mark EJ Newman, and Cristopher Moore. 2004. Finding commu- nity structure in very large networks. Physical review E 70, 6 (2004), 066111
2004
-
[12]
Michele Coscia, Fosca Giannotti, and Dino Pedreschi. 2011. A classification for community discovery methods in complex networks. Statistical Analysis and Data Mining: The ASA Data Science Journal 4, 5 (2011), 512–546
2011
-
[13]
Dipanjan Das and Slav Petrov. 2011. Unsupervised part-of-speech tagging with bilingual graph-based projections. InProceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies . 600–609
2011
-
[14]
Jordi Duch and Alex Arenas. 2005. Community detection in complex networks using extremal optimization. Physical review E 72, 2 (2005), 027104
2005
-
[15]
Fazlali, E
M. Fazlali, E. Moradi, and H. Malazi. 2017. Adaptive parallel Louvain community detection on a multicore platform. Microprocessors and microsystems 54 (Oct 2017), 26–34
2017
-
[16]
Fortunato
S. Fortunato. 2010. Community detection in graphs. Physics reports 486, 3-5 (2010), 75–174
2010
-
[17]
Gach and J
O. Gach and J. Hao. 2014. Improving the Louvain algorithm for community detection with modularity maximization. InArtificial Evolution: 11th International Conference, Evolution Artificielle, EA , Bordeaux, France, October 21-23, . Revised Selected Papers 11. Springer, Springer...
2014
-
[18]
Gawande, S
N. Gawande, S. Ghosh, M. Halappanavar, A. Tumeo, and A. Kalyanaraman. 2022. Towards scaling community detection on distributed-memory heterogeneous systems. Parallel Comput. 111 (2022), 102898
2022
-
[19]
Gheibi, T
S. Gheibi, T. Banerjee, S. Ranka, and S. Sahni. 2020. Cache Efficient Louvain with Local RCM. In IEEE Symposium on Computers and Communications (ISCC) . IEEE, 1–6
2020
-
[20]
Ghosh, M
S. Ghosh, M. Halappanavar, A. Tumeo, and A. Kalyanarainan. 2019. Scaling and quality of modularity optimization methods for graph clustering. In IEEE High Performance Extreme Computing Conference (HPEC) . IEEE, 1–6
2019
-
[21]
Ghosh, M
S. Ghosh, M. Halappanavar, A. Tumeo, A. Kalyanaraman, and A.H. Gebremedhin
-
[22]
Ghosh, M
S. Ghosh, M. Halappanavar, A. Tumeo, A. Kalyanaraman, H. Lu, D. Chavarria- Miranda, A. Khan, and A. Gebremedhin. 2018. Distributed louvain algorithm for graph community detection. In IEEE International Parallel and Distributed Processing Symposium (IPDPS). Vancouver, British C...
2018
-
[23]
S. Gregory. 2010. Finding overlapping communities in networks by label propa- gation. New Journal of Physics 12 (10 2010), 103018. Issue 10
2010
-
[24]
Roger Guimerà, DB Stouffer, Marta Sales-Pardo, EA Leicht, MEJ Newman, and Luis AN Amaral. 2010. Origin of compartmentalization in food webs. Ecology 91, 10 (2010), 2941–2951
2010
-
[25]
Halappanavar, H
M. Halappanavar, H. Lu, A. Kalyanaraman, and A. Tumeo. 2017. Scalable static and dynamic community detection using Grappolo. In IEEE High Performance Extreme Computing Conference (HPEC) . IEEE, Waltham, MA USA, 1–6
2017
-
[26]
Nandinee Haq and Z Jane Wang. 2016. Community detection from genomic datasets across human cancers. In 2016 IEEE Global Conference on Signal and Information Processing (GlobalSIP). IEEE, 1147–1150
2016
-
[27]
Yong He and Alan Evans. 2010. Graph theoretical modeling of brain connectivity. Current opinion in neurology 23, 4 (2010), 341–350
2010
-
[28]
S. Kang, C. Hastings, J. Eaton, and B. Rees. 2023. cuGraph C++ primitives: vertex/edge-centric building blocks for parallel graph computing. In IEEE Inter- national Parallel and Distributed Processing Symposium Workshops . 226–229
2023
-
[29]
Kloster and D
K. Kloster and D. Gleich. 2014. Heat kernel based community detection. In Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining . ACM, New York, USA, 1386–1395
2014
-
[30]
Kolodziej, M
S. Kolodziej, M. Aznaveh, M. Bullock, J. David, T. Davis, M. Henderson, Y. Hu, and R. Sandstrom. 2019. The SuiteSparse matrix collection website interface. The Journal of Open Source Software 4, 35 (Mar 2019), 1244
2019
-
[31]
Lancichinetti and S
A. Lancichinetti and S. Fortunato. 2009. Community detection algorithms: a comparative analysis. Physical Review. E, Statistical, Nonlinear, and Soft Matter Physics 80, 5 Pt 2 (Nov 2009), 056117
2009
-
[32]
H. Lu, M. Halappanavar, and A. Kalyanaraman. 2015. Parallel heuristics for scalable community detection. Parallel computing 47 (Aug 2015), 19–37
2015
-
[33]
Mohammadi, M
M. Mohammadi, M. Fazlali, and M. Hosseinzadeh. 2020. Accelerating Louvain community detection algorithm on graphic processing unit. The Journal of supercomputing (Nov 2020)
2020
-
[34]
M. Naim, F. Manne, M. Halappanavar, and A. Tumeo. 2017. Community detection on the GPU. In IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, Orlando, Florida, USA, 625–634
2017
-
[35]
M. Newman. 2006. Finding community structure in networks using the eigen- vectors of matrices. Physical review E 74, 3 (2006), 036104
2006
-
[36]
John Nickolls and William J Dally. 2010. The GPU computing era. IEEE micro 30, 2 (2010), 56–69
2010
-
[37]
Ozaki, H
N. Ozaki, H. Tezuka, and M. Inaba. 2016. A simple acceleration method for the Louvain algorithm. International Journal of Computer and Electrical Engineering 8, 3 (2016), 207
2016
-
[38]
Hang Qie, Shijie Li, Yong Dou, Jinwei Xu, Yunsheng Xiong, and Zikai Gao. 2022. Isolate sets partition benefits community detection of parallel Louvain method. Scientific Reports 12, 1 (2022), 8248
2022
-
[39]
X. Que, F. Checconi, F. Petrini, and J. Gunnels. 2015. Scalable community detec- tion with the louvain algorithm. In IEEE International Parallel and Distributed Processing Symposium. IEEE, IEEE, Hyderabad, India, 28–37. 14 CPU vs. GPU for Community Detection: Performance Insig...
2015
-
[40]
Raghavan, R
U. Raghavan, R. Albert, and S. Kumara. 2007. Near linear time algorithm to detect community structures in large-scale networks. Physical Review E 76, 3 (Sep 2007), 036106–1–036106–11
2007
-
[41]
Jörg Reichardt and Stefan Bornholdt. 2006. Statistical mechanics of community detection. Physical review E 74, 1 (2006), 016110
2006
-
[42]
Rosvall and C
M. Rosvall and C. Bergstrom. 2008. Maps of random walks on complex networks reveal community structure. Proceedings of the national academy of sciences 105, 4 (2008), 1118–1123
2008
-
[43]
Rotta and A
R. Rotta and A. Noack. 2011. Multilevel local search algorithms for modularity clustering. Journal of Experimental Algorithmics (JEA) 16 (2011), 2–1
2011
-
[44]
Ryu and D
S. Ryu and D. Kim. 2016. Quick community detection of big graph data using modified louvain algorithm. In IEEE 18th International Conference on High Perfor- mance Computing and Communications (HPCC) . IEEE, Sydney, NSW, 1442–1445
2016
-
[46]
Subhajit Sahu. 2024. 𝜈-LPA: Fast GPU-based Label Propagation Algorithm (LPA) for Community Detection. arXiv preprint arXiv:2411.11468 (2024)
2024 arXiv
-
[47]
Subhajit Sahu, Kishore Kothapalli, Hemalatha Eedi, and Sathya Peri. 2024. Lock- free Computation of PageRank in Dynamic Graphs. In 2024 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW) . IEEE, 825– 834
2024
-
[48]
Marcel Salathé and James H Jones. 2010. Dynamics and control of diseases in networks with community structure. PLoS computational biology 6, 4 (2010), e1000736
2010
-
[49]
Jason Sanders and Edward Kandrot. 2010. CUDA by example: an introduction to general-purpose GPU programming . Addison-Wesley Professional
2010
-
[50]
Sattar and S
N. Sattar and S. Arifuzzaman. 2019. Overcoming MPI Communication Over- head for Distributed Community Detection. In Software Challenges to Exascale Computing, A. Majumdar and R. Arora (Eds.). Springer Singapore, Singapore, 77–90
2019
-
[51]
Naw Safrin Sattar and Shaikh Arifuzzaman. 2022. Scalable distributed Louvain al- gorithm for community detection in large graphs. The Journal of Supercomputing 78, 7 (2022), 10275–10309
2022
-
[52]
J. Shi, L. Dhulipala, D. Eisenstat, J. Łącki, and V. Mirrokni. 2021. Scalable com- munity detection via parallel correlation clustering
2021
-
[53]
Staudt, A
C.L. Staudt, A. Sazonovs, and H. Meyerhenke. 2016. NetworKit: A tool suite for large-scale complex network analysis. Network Science 4, 4 (2016), 508–530
2016
-
[54]
Christian L Staudt and Henning Meyerhenke. 2015. Engineering parallel al- gorithms for community detection in massive networks. IEEE Transactions on Parallel and Distributed Systems 27, 1 (2015), 171–184
2015
-
[55]
Aaron M Tenenbaum. 1990. Data structures using C . Pearson Education India
1990
-
[56]
Jesmin Jahan Tithi, Andrzej Stasiak, Sriram Aananthakrishnan, and Fabrizio Petrini. 2020. Prune the unnecessary: Parallel pull-push louvain algorithms with automatic edge pruning. In Proceedings of the 49th International Conference on Parallel Processing. 1–11
2020
-
[57]
V. Traag. 2015. Faster unfolding of communities: Speeding up the Louvain algorithm. Physical Review E 92, 3 (2015), 032801
2015
-
[58]
Traag and L
V.A. Traag and L. Šubelj. 2023. Large network community detection by fast label propagation. Scientific Reports 13, 1 (2023), 2701
2023
-
[59]
Traag, L
V. Traag, L. Waltman, and N. Eck. 2019. From Louvain to Leiden: guaranteeing well-connected communities. Scientific Reports 9, 1 (Mar 2019), 5233
2019
-
[60]
Waltman and N
L. Waltman and N. Eck. 2013. A smart local moving algorithm for large-scale modularity-based community detection. The European physical journal B 86, 11 (2013), 1–14
2013
-
[61]
Whang, D
J. Whang, D. Gleich, and I. Dhillon. 2013. Overlapping community detection using seed set expansion. In Proceedings of the 22nd ACM international conference on Information & Knowledge Management . 2099–2108
2013
-
[62]
Wickramaarachchi, M
C. Wickramaarachchi, M. Frincu, P. Small, and V. Prasanna. 2014. Fast parallel algorithm for unfolding of communities in large graphs. InIEEE High Performance Extreme Computing Conference (HPEC) . IEEE, IEEE, Waltham, MA USA, 1–6
2014
-
[63]
Tianji Wu, Bo Wang, Yi Shan, Feng Yan, Yu Wang, and Ningyi Xu. 2010. Effi- cient pagerank and spmv computation on amd gpus. In 2010 39th International Conference on Parallel Processing . IEEE, 81–89
2010
-
[64]
J. Xie, B. Szymanski, and X. Liu. 2011. SLPA: Uncovering overlapping communi- ties in social networks via a speaker-listener interaction dynamic process. InIEEE 11th International Conference on Data Mining Workshops . IEEE, IEEE, Vancouver, Canada, 344–349
2011
-
[65]
X. You, Y. Ma, and Z. Liu. 2020. A three-stage algorithm on community detection in social networks. Knowledge-Based Systems 187 (2020), 104822
2020
-
[66]
Y. You, L. Ren, Z. Zhang, K. Zhang, and J. Huang. 2022. Research on improvement of Louvain community detection algorithm. In 2nd International Conference on Artificial Intelligence, Automation, and High-Performance Computing (AIAHPC ) , Vol. 12348. SPIE, Zhuhai, China, 527–531
2022
-
[67]
T. Zeitz. 2017. Engineering Distributed GraphClustering using MapReduce. https://i11www.iti.kit.edu/_media/teaching/theses/ma-zeitz-17.pdf
2017
-
[68]
Zeng and H
J. Zeng and H. Yu. 2015. Parallel Modularity-Based Community Detection on Large-Scale Graphs. In IEEE International Conference on Cluster Computing . IEEE, 1–10
2015
-
[69]
Zhang, J
J. Zhang, J. Fei, X. Song, and J. Feng. 2021. An improved Louvain algorithm for community detection. Mathematical Problems in Engineering 2021 (2021), 1–14
2021
-
[70]
Ziqiao Zhang, Peng Pu, Dingding Han, and Ming Tang. 2018. Self-adaptive Louvain algorithm: Fast and stable community detection algorithm based on the principle of small probability event. Physica A: Statistical Mechanics and Its Applications 506 (2018), 975–986. 15 Subhajit Sa...
2018
-
[2018]
In 2018 IEEE High Performance extreme Computing Conference (HPEC)
Scalable distributed memory community detection using vite. In 2018 IEEE High Performance extreme Computing Conference (HPEC) . IEEE, 1–7
2018
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.