Pith. sign in

REVIEW 3 cited by

Advancing Content Moderation: Evaluating Large Language Models for Detecting Sensitive Content Across Text, Images, and Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.17123 v1 pith:2NXNPJ4J submitted 2024-11-26 cs.CV cs.AI

classification cs.CVcs.AI
keywords contentmediafalsemoderationplatformsacrosscontentsllms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The widespread dissemination of hate speech, harassment, harmful and sexual content, and violence across websites and media platforms presents substantial challenges and provokes widespread concern among different sectors of society. Governments, educators, and parents are often at odds with media platforms about how to regulate, control, and limit the spread of such content. Technologies for detecting and censoring the media contents are a key solution to addressing these challenges. Techniques from natural language processing and computer vision have been used widely to automatically identify and filter out sensitive content such as offensive languages, violence, nudity, and addiction in both text, images, and videos, enabling platforms to enforce content policies at scale. However, existing methods still have limitations in achieving high detection accuracy with fewer false positives and false negatives. Therefore, more sophisticated algorithms for understanding the context of both text and image may open rooms for improvement in content censorship to build a more efficient censorship system. In this paper, we evaluate existing LLM-based content moderation solutions such as OpenAI moderation model and Llama-Guard3 and study their capabilities to detect sensitive contents. Additionally, we explore recent LLMs such as GPT, Gemini, and Llama in identifying inappropriate contents across media outlets. Various textual and visual datasets like X tweets, Amazon reviews, news articles, human photos, cartoons, sketches, and violence videos have been utilized for evaluation and comparison. The results demonstrate that LLMs outperform traditional techniques by achieving higher accuracy and lower false positive and false negative rates. This highlights the potential to integrate LLMs into websites, social media platforms, and video-sharing services for regulatory and content moderation purposes.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Binary Moderation: Identifying Fine-Grained Sexist and Misogynistic Behavior on GitHub with Large Language Models

    cs.SE 2025-07 conditional novelty 6.0 of 10

    An instruction-tuned GPT-4o prompt achieves an MCC of 0.501 on 12-category sexism/misogyny classification of GitHub comments, but the evaluation was tuned on the same test set.

  2. Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching

    cs.CV 2025-12 conditional novelty 4.0 of 10

    A deployed hybrid moderation system combining supervised classification and reference-based similarity matching, boosted by MLLM distillation, reduces unwanted livestream views by 6–8%.

  3. Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases

    cs.CL 2025-08 conditional novelty 3.0 of 10

    A majority-vote ensemble of Gemini Flash 2.5, Gemini Pro 2.5, and GPT o3 achieves 92.7% accuracy on the QIAS 2025 Islamic inheritance test set, outperforming each model alone and all open Arabic models.

Pith tools