Pith. sign in

REVIEW 2 cited by

TwHIN-BERT: A Socially-Enriched Pre-trained Language Model for Multilingual Tweet Representations at Twitter

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.07562 v3 pith:UMAXAUBJ submitted 2022-09-15 cs.CL

classification cs.CL
keywords sociallanguagemodelpre-trainedtwhin-bertmodelsmultilingualnetwork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained language models (PLMs) are fundamental for natural language processing applications. Most existing PLMs are not tailored to the noisy user-generated text on social media, and the pre-training does not factor in the valuable social engagement logs available in a social network. We present TwHIN-BERT, a multilingual language model productionized at Twitter, trained on in-domain data from the popular social network. TwHIN-BERT differs from prior pre-trained language models as it is trained with not only text-based self-supervision, but also with a social objective based on the rich social engagements within a Twitter heterogeneous information network (TwHIN). Our model is trained on 7 billion tweets covering over 100 distinct languages, providing a valuable representation to model short, noisy, user-generated text. We evaluate our model on various multilingual social recommendation and semantic understanding tasks and demonstrate significant metric improvement over established pre-trained language models. We open-source TwHIN-BERT and our curated hashtag prediction and social engagement benchmark datasets to the research community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents

    cs.IR 2026-07 conditional novelty 6.0 of 10

    CtrlBench-Rec uses LLM-driven agent probes, evolved through clustering and merging, to measure how well recommender systems can be steered toward target content, interest profiles, and long-tail items.

  2. Full-Stack Optimized Large Language Models for Lifelong Sequential Behavior Comprehension in Recommendation

    cs.IR 2025-01 conditional novelty 6.0 of 10

    ReLLaX combines semantic behavior retrieval, collaborative soft prompts, and a new fully interactive LoRA variant to improve LLM-based CTR prediction on long user histories.

Pith tools