Pith. sign in

REVIEW 1 cited by

Learning Invariant Representations of Social Media Users

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.04979 v1 pith:E3AXDOMO submitted 2019-10-11 cs.SI cs.CLcs.LGstat.ML

classification cs.SIcs.CLcs.LGstat.ML
keywords usersmediasocialinvariantlearningmappingspacetime
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The evolution of social media users' behavior over time complicates user-level comparison tasks such as verification, classification, clustering, and ranking. As a result, na\"ive approaches may fail to generalize to new users or even to future observations of previously known users. In this paper, we propose a novel procedure to learn a mapping from short episodes of user activity on social media to a vector space in which the distance between points captures the similarity of the corresponding users' invariant features. We fit the model by optimizing a surrogate metric learning objective over a large corpus of unlabeled social media content. Once learned, the mapping may be applied to users not seen at training time and enables efficient comparisons of users in the resulting vector space. We present a comprehensive evaluation to validate the benefits of the proposed approach using data from Reddit, Twitter, and Wikipedia.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Quantifying Misattribution Unfairness in Authorship Attribution

    cs.CL 2025-06 reject novelty 5.0 of 10

    Authorship attribution models misattribute texts to some authors far more often than chance, and the risk is highest for authors whose author embeddings sit near the centroid.

Pith tools