Pith. sign in

REVIEW 1 cited by

Inducing Domain-Specific Sentiment Lexicons from Unlabeled Corpora

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1606.02820 v2 pith:RUR5O4CJ submitted 2016-06-09 cs.CL

classification cs.CL
keywords sentimentlexiconscommunitiesdomain-specificcommunity-specificenglishframeworkhistorical
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

A word's sentiment depends on the domain in which it is used. Computational social science research thus requires sentiment lexicons that are specific to the domains being studied. We combine domain-specific word embeddings with a label propagation framework to induce accurate domain-specific sentiment lexicons using small sets of seed words, achieving state-of-the-art performance competitive with approaches that rely on hand-curated resources. Using our framework we perform two large-scale empirical studies to quantify the extent to which sentiment varies across time and between communities. We induce and release historical sentiment lexicons for 150 years of English and community-specific sentiment lexicons for 250 online communities from the social media forum Reddit. The historical lexicons show that more than 5% of sentiment-bearing (non-neutral) English words completely switched polarity during the last 150 years, and the community-specific lexicons highlight how sentiment varies drastically between different communities.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Twitter Sentiment on Affordable Care Act using Score Embedding

    cs.LG 2019-08 reject novelty 4.0 of 10

    Score embedding, a CNN initialized with per-class word frequencies, reaches about 69% accuracy on ACA tweets and 46% on SST, but its central public-opinion finding is confounded by the imbalanced training labels.

Pith tools