Pith. sign in

REVIEW 2 cited by

MS-BioGraphs: Sequence Similarity Graph Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.16744 v1 pith:UWGYM62T submitted 2023-08-31 cs.DC cs.ARcs.CEcs.DMcs.PF

classification cs.DCcs.ARcs.CEcs.DMcs.PF
keywords graphms-biographsdatasetsprocesssimilarityavailabledatafamily
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Progress in High-Performance Computing in general, and High-Performance Graph Processing in particular, is highly dependent on the availability of publicly-accessible, relevant, and realistic data sets. To ensure continuation of this progress, we (i) investigate and optimize the process of generating large sequence similarity graphs as an HPC challenge and (ii) demonstrate this process in creating MS-BioGraphs, a new family of publicly available real-world edge-weighted graph datasets with up to $2.5$ trillion edges, that is, $6.6$ times greater than the largest graph published recently. The largest graph is created by matching (i.e., all-to-all similarity aligning) $1.7$ billion protein sequences. The MS-BioGraphs family includes also seven subgraphs with different sizes and direction types. We describe two main challenges we faced in generating large graph datasets and our solutions, that are, (i) optimizing data structures and algorithms for this multi-step process and (ii) WebGraph parallel compression technique. We present a comparative study of structural characteristics of MS-BioGraphs. The datasets are available online on https://blogs.qub.ac.uk/DIPSA/MS-BioGraphs .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Optimizing Locality of Graph Transposition on Modern Architectures

    cs.DC 2025-01 conditional novelty 6.0 of 10

    PoTra, a structure-aware graph transposition algorithm, separates high-degree and low-degree vertices to keep per-thread counters in cache and reports up to 8.7x speedups on large graphs.

  2. Accelerating Loading WebGraphs in ParaGrapher

    cs.DC 2025-07 conditional novelty 3.0 of 10

    PG-Fuse and CompBin speed up loading of WebGraph-format graphs by up to 7.6x and 21.8x, respectively, on a high-bandwidth shared filesystem.

Pith tools