Pith. sign in

REVIEW 1 cited by

Re-ranking Using Large Language Models for Mitigating Exposure to Harmful Content on Social Media Platforms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.13977 v3 pith:NEHIFFXE submitted 2025-01-23 cs.CL cs.AIcs.CYcs.SI

classification cs.CLcs.AIcs.CYcs.SI
keywords contentexposureharmfulmodelsre-rankingthreeapproachdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Social media platforms utilize Machine Learning (ML) and Artificial Intelligence (AI) powered recommendation algorithms to maximize user engagement, which can result in inadvertent exposure to harmful content. Current moderation efforts, reliant on classifiers trained with extensive human-annotated data, struggle with scalability and adapting to new forms of harm. To address these challenges, we propose a novel re-ranking approach using Large Language Models (LLMs) in zero-shot and few-shot settings. Our method dynamically assesses and re-ranks content sequences, effectively mitigating harmful content exposure without requiring extensive labeled data. Alongside traditional ranking metrics, we also introduce two new metrics to evaluate the effectiveness of re-ranking in reducing exposure to harmful content. Through experiments on three datasets, three models and across three configurations, we demonstrate that our LLM-based approach significantly outperforms existing proprietary moderation approaches, offering a scalable and adaptable solution for harm mitigation.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Binary Moderation: Identifying Fine-Grained Sexist and Misogynistic Behavior on GitHub with Large Language Models

    cs.SE 2025-07 conditional novelty 6.0 of 10

    An instruction-tuned GPT-4o prompt achieves an MCC of 0.501 on 12-category sexism/misogyny classification of GitHub comments, but the evaluation was tuned on the same test set.

Pith tools