REVIEW 2 cited by
MS-BioGraphs: Sequence Similarity Graph Datasets
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
Progress in High-Performance Computing in general, and High-Performance Graph Processing in particular, is highly dependent on the availability of publicly-accessible, relevant, and realistic data sets. To ensure continuation of this progress, we (i) investigate and optimize the process of generating large sequence similarity graphs as an HPC challenge and (ii) demonstrate this process in creating MS-BioGraphs, a new family of publicly available real-world edge-weighted graph datasets with up to $2.5$ trillion edges, that is, $6.6$ times greater than the largest graph published recently. The largest graph is created by matching (i.e., all-to-all similarity aligning) $1.7$ billion protein sequences. The MS-BioGraphs family includes also seven subgraphs with different sizes and direction types. We describe two main challenges we faced in generating large graph datasets and our solutions, that are, (i) optimizing data structures and algorithms for this multi-step process and (ii) WebGraph parallel compression technique. We present a comparative study of structural characteristics of MS-BioGraphs. The datasets are available online on https://blogs.qub.ac.uk/DIPSA/MS-BioGraphs .
Forward citations
Cited by 2 Pith papers
-
On Optimizing Locality of Graph Transposition on Modern Architectures
PoTra, a structure-aware graph transposition algorithm, separates high-degree and low-degree vertices to keep per-thread counters in cache and reports up to 8.7x speedups on large graphs.
-
Accelerating Loading WebGraphs in ParaGrapher
PG-Fuse and CompBin speed up loading of WebGraph-format graphs by up to 7.6x and 21.8x, respectively, on a high-bandwidth shared filesystem.
Discussion (0). Continue with ORCID to comment.