A latent mediation framework with sparse autoencoders enables non-additive token-level influence attribution in LLMs by learning orthogonal features and back-propagating attributions.
arXiv preprint arXiv:2010.00711 , year=
2 Pith papers cite this work, alongside 113 external citations. Polarity classification is still indexing.
2
Pith papers citing it
113
external citations · external index
fields
cs.LG 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
Introduces BetXplain, an explanation-annotated dataset of social media betting ads collected from Instagram and Reddit for detecting manipulative and deceptive advertising.
citing papers explorer
-
Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces
A latent mediation framework with sparse autoencoders enables non-additive token-level influence attribution in LLMs by learning orthogonal features and back-propagating attributions.
-
BetXplain: An Explanation-Annotated Dataset for Detecting Manipulative Betting Advertisements on Social Media
Introduces BetXplain, an explanation-annotated dataset of social media betting ads collected from Instagram and Reddit for detecting manipulative and deceptive advertising.