A composed pipeline of betweenness cuts, per-edge logistic classifiers, and GitHub-asserted edges resolves ~107M author identities while reducing the largest cluster from 170,431 to under 7,000 and improving gold recall from 0.44 to 0.70.
Title resolution pending
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
A scan of 5.8 billion git commits finds 17.59% carry cryptographic signatures, enabling a key-to-identity graph and trust tiers that calibrate heuristic author disambiguation.
A mega-cluster-free author-identity map for 5.87B git commits folds 106.8M raw strings into 62.7M canonical identities, validated jointly for splitting and clumping errors.
citing papers explorer
-
Scaling Author Identity Disambiguation to the World of Code: A Methodology
A composed pipeline of betweenness cuts, per-edge logistic classifiers, and GitHub-asserted edges resolves ~107M author identities while reducing the largest cluster from 170,431 to under 7,000 and improving gold recall from 0.44 to 0.70.
-
Claimed or Attested? A Commit-Signature Dataset and Identity Trust Tiers across the World of Code
A scan of 5.8 billion git commits finds 17.59% carry cryptographic signatures, enabling a key-to-identity graph and trust tiers that calibrate heuristic author disambiguation.
-
A Global Author-Identity Map for the World of Code:62.7M Developer Identities from 106.8M Author Strings over 5.87B Commits
A mega-cluster-free author-identity map for 5.87B git commits folds 106.8M raw strings into 62.7M canonical identities, validated jointly for splitting and clumping errors.