Pith. sign in

REVIEW 1 cited by

Understanding Transformers for Bot Detection in Twitter

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.06182 v1 pith:LXW3A2TK submitted 2021-04-13 cs.CL cs.LG

classification cs.CLcs.LG
keywords detectionfine-tuninggenerativetransformersbertlanguagelikemedia
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper we shed light on the impact of fine-tuning over social media data in the internal representations of neural language models. We focus on bot detection in Twitter, a key task to mitigate and counteract the automatic spreading of disinformation and bias in social media. We investigate the use of pre-trained language models to tackle the detection of tweets generated by a bot or a human account based exclusively on its content. Unlike the general trend in benchmarks like GLUE, where BERT generally outperforms generative transformers like GPT and GPT-2 for most classification tasks on regular text, we observe that fine-tuning generative transformers on a bot detection task produces higher accuracies. We analyze the architectural components of each transformer and study the effect of fine-tuning on their hidden states and output representations. Among our findings, we show that part of the syntactical information and distributional properties captured by BERT during pre-training is lost upon fine-tuning while the generative pre-training approach manage to preserve these properties.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Large Language Models and Social Media Information Integrity: Opportunities, Challenges, and Research Directions

    cs.CR 2026-08 conditional novelty 4.0 of 10

    A systematic review of 215 studies concludes that large language models both enable and counter misinformation, social bots, and privacy threats on social media, and maps open research gaps.

Pith tools