REVIEW 1 cited by
Nichelle and Nancy: The Influence of Demographic Attributes and Tokenization Length on First Name Biases
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Through the use of first name substitution experiments, prior research has demonstrated the tendency of social commonsense reasoning models to systematically exhibit social biases along the dimensions of race, ethnicity, and gender (An et al., 2023). Demographic attributes of first names, however, are strongly correlated with corpus frequency and tokenization length, which may influence model behavior independent of or in addition to demographic factors. In this paper, we conduct a new series of first name substitution experiments that measures the influence of these factors while controlling for the others. We find that demographic attributes of a name (race, ethnicity, and gender) and name tokenization length are both factors that systematically affect the behavior of social commonsense reasoning models.
Forward citations
Cited by 1 Pith paper
-
On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena
Arab cultural entities that double as everyday Arabic words are harder for language models to recognize, especially when tokenized as single tokens.
Discussion (0). Continue with ORCID to comment.