125 coordinated Wikipedia animal-welfare edits dominate attribution and counterfactual influence for animal-welfare queries on Llama models, with no spillover to general queries about the same entities.
Chang, Dheeraj Rajagopal, Tolga Bolukbasi, Lucas Dixon, and Ian Tenney
4 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 4roles
baseline 1polarities
baseline 1representative citing papers
idSCD uses semantic correlation descriptors to perform dataset membership inference by comparing learned semantic structures, outperforming baselines in NLI, emotion, and medical text experiments.
Seven stance-making linguistic features strengthen pro-animal-welfare preference in fine-tuned Llama-3.2-1B and Mistral-7B; hedging and concreteness dilute it; first-person is null.
Influence scoring can use only forward passes: CountSketch-compressed outer products of the LM-head residual and final hidden state give accurate attribution and valuation from 14M to 32B parameters.
citing papers explorer
-
Small edits, large models: How Wikipedia advocacy shapes LLM values
125 coordinated Wikipedia animal-welfare edits dominate attribution and counterfactual influence for animal-welfare queries on Llama models, with no spillover to general queries about the same entities.
-
idSCD: Identifying Training Datasets through Semantic Correlation Descriptors
idSCD uses semantic correlation descriptors to perform dataset membership inference by comparing learned semantic structures, outperforming baselines in NLI, emotion, and medical text experiments.
-
Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare
Seven stance-making linguistic features strengthen pro-animal-welfare preference in fine-tuned Llama-3.2-1B and Mistral-7B; hedging and concreteness dilute it; first-person is null.
-
Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation
Influence scoring can use only forward passes: CountSketch-compressed outer products of the LM-head residual and final hidden state give accurate attribution and valuation from 14M to 32B parameters.