Pith. sign in

REVIEW 3 cited by

Discovering Latent Themes in Social Media Messaging: A Machine-in-the-Loop Approach Integrating LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.10707 v2 pith:XIDUBUGX submitted 2024-03-15 cs.CL cs.AIcs.CYcs.LGcs.SI

classification cs.CLcs.AIcs.CYcs.LGcs.SI
keywords approachthemesmediasocialanalysisclimatelatentmessaging
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Grasping the themes of social media content is key to understanding the narratives that influence public opinion and behavior. The thematic analysis goes beyond traditional topic-level analysis, which often captures only the broadest patterns, providing deeper insights into specific and actionable themes such as "public sentiment towards vaccination", "political discourse surrounding climate policies," etc. In this paper, we introduce a novel approach to uncovering latent themes in social media messaging. Recognizing the limitations of the traditional topic-level analysis, which tends to capture only overarching patterns, this study emphasizes the need for a finer-grained, theme-focused exploration. Traditional theme discovery methods typically involve manual processes and a human-in-the-loop approach. While valuable, these methods face challenges in scalability, consistency, and resource intensity in terms of time and cost. To address these challenges, we propose a machine-in-the-loop approach that leverages the advanced capabilities of Large Language Models (LLMs). To demonstrate our approach, we apply our framework to contentious topics, such as climate debate and vaccine debate. We use two publicly available datasets: (1) the climate campaigns dataset of 21k Facebook ads and (2) the COVID-19 vaccine campaigns dataset of 9k Facebook ads. Our quantitative and qualitative analysis shows that our methodology yields more accurate and interpretable results compared to the baselines. Our results not only demonstrate the effectiveness of our approach in uncovering latent themes but also illuminate how these themes are tailored for demographic targeting in social media contexts. Additionally, our work sheds light on the dynamic nature of social media, revealing the shifts in the thematic focus of messaging in response to real-world events.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HAMLET: Healthcare-focused Adaptive Multilingual Learning Embedding-based Topic Modeling

    cs.CL 2025-05 reject novelty 5.0 of 10

    HAMLET combines LLM-generated topics with graph-neural-network embedding refinement and reports better topic quality scores on English and French healthcare datasets than using unrefined embeddings.

  2. Can LLMs Assist Annotators in Identifying Morality Frames? -- Case Study on Vaccination Debate on Social Media

    cs.CL 2025-02 reject novelty 4.0 of 10

    GPT-4o few-shot prompting with explanations identified moral frames in vaccine tweets with 90.79 percent agreement from nine annotators, who also reported lower cognitive load and difficulty.

  3. LLMs-in-the-Loop Part 2: Expert Small AI Models for Anonymization and De-identification of PHI Across Multiple Languages

    cs.CL 2024-12 reject novelty 4.0 of 10

    Small fine-tuned NER models for de-identifying health information in eight languages are claimed to beat GPT-4o, but only the English result is benchmarked externally.

Pith tools