REVIEW 3 cited by
A Material Lens on Coloniality in NLP
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Coloniality, the continuation of colonial harms beyond "official" colonization, has pervasive effects across society and scientific fields. Natural Language Processing (NLP) is no exception to this broad phenomenon. In this work, we argue that coloniality is implicitly embedded in and amplified by NLP data, algorithms, and software. We formalize this analysis using Actor-Network Theory (ANT): an approach to understanding social phenomena through the network of relationships between human stakeholders and technology. We use our Actor-Network to guide a quantitative survey of the geography of different phases of NLP research, providing evidence that inequality along colonial boundaries increases as NLP builds on itself. Based on this, we argue that combating coloniality in NLP requires not only changing current values but also active work to remove the accumulation of colonial ideals in our foundational data and algorithms.
Forward citations
Cited by 3 Pith papers
-
Hidden Language Consistency Phenomena in Reasoning LLMs
Reasoning models often stop using the requested language as problems get harder, and this language breakdown can make accuracy look better than it is.
-
One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers
A tokenizer trained on more languages than the model's main pretraining set makes later language adaptation faster and better, with minimal loss on the pretraining languages.
-
Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
Aya Expanse 8B and 32B report state-of-the-art multilingual win-rates on a new 23-language translated Arena-Hard benchmark, with the 32B beating Llama 3.1 70B by 54.0%.
Discussion (0). Continue with ORCID to comment.