Pith. sign in

REVIEW

The Remarkable Benefit of User-Level Aggregation for Lexical-based Population-Level Predictions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1808.09600 v1 pith:FYK2JFFH submitted 2018-08-29 cs.SI cs.CY

classification cs.SIcs.CY
keywords aggregatedcommunity-leveloutcomespredictionbilliondatapredictionstweets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Nowcasting based on social media text promises to provide unobtrusive and near real-time predictions of community-level outcomes. These outcomes are typically regarding people, but the data is often aggregated without regard to users in the Twitter populations of each community. This paper describes a simple yet effective method for building community-level models using Twitter language aggregated by user. Results on four different U.S. county-level tasks, spanning demographic, health, and psychological outcomes show large and consistent improvements in prediction accuracies (e.g. from Pearson r=.73 to .82 for median income prediction or r=.37 to .47 for life satisfaction prediction) over the standard approach of aggregating all tweets. We make our aggregated and anonymized community-level data, derived from 37 billion tweets -- over 1 billion of which were mapped to counties, available for research.

Discussion (0). Sign in to comment.

Pith tools