Pith. sign in

REVIEW 5 cited by

Predicting Race and Ethnicity From the Sequence of Characters in a Name

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1805.02109 v2 pith:PMX6ERAF submitted 2018-05-05 stat.AP stat.ML

Predicting Race and Ethnicity From the Sequence of Characters in a Name

classification stat.AP stat.ML
keywords namesethnicityraceaccuracylastmodelvariouscharacters
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

To answer questions about racial inequality and fairness, we often need a way to infer race and ethnicity from names. One way to infer race and ethnicity from names is by relying on the Census Bureau's list of popular last names. The list, however, suffers from at least three limitations: 1. it only contains last names, 2. it only includes popular last names, and 3. it is updated once every 10 years. To provide better generalization, and higher accuracy when first names are available, we model the relationship between characters in a name and race and ethnicity using various techniques. A model using Long Short-Term Memory works best with out-of-sample accuracy of .85. The best-performing last-name model achieves out-of-sample accuracy of .81. To illustrate the utility of the models, we apply them to campaign finance data to estimate the share of donations made by people of various racial groups, and to news data to estimate the coverage of various races and ethnicities in the news.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Using Embedding Models to Improve Probabilistic Race Prediction

    cs.CL 2026-04 unverdicted novelty 7.0

    Embedding models trained on Census surname, first-name, and voter file data improve probabilistic race prediction for uncommon surnames, with full-name embeddings delivering the largest gains especially for Hispanic a...

  2. NameBERT: Scaling Name-Based Nationality Classification with LLM-Augmented Open Academic Data

    cs.CL 2026-04 unverdicted novelty 7.0

    NameBERT models trained on LLM-augmented academic name data outperform state-of-the-art baselines in nationality classification from names, with augmentation providing gains especially on tail countries.

  3. War in the Abstract: The Rise and Consequences of Militarized Language in Scientific Communication

    cs.CL 2026-06 unverdicted novelty 6.0

    Militarized language in scientific abstracts increased substantially over 15 years in correlation with global conflicts and causally reduced credibility, funding willingness, and policy support in a controlled experiment.

  4. Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation

    cs.IR 2026-05 unverdicted novelty 6.0

    Audits of 43 LLMs show that varying persona prompts (language, location, role-and-task) and context affects technical quality and social representativeness of scholar recommendations, with location impacting diversity...

  5. QUIET: Quantifying Underutilized Influential Edges for Targeted Synchronization

    eess.SY 2026-06 unverdicted novelty 5.0

    QUIET ranks structurally influential but functionally quiet white-matter edges to identify low-energy pathways for targeted neural synchronization, validated on synthetic networks and HCP data linking control energy t...