Pith. sign in

REVIEW 1 cited by

Paraphrase Detection: Human vs. Machine Content

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.13989 v1 pith:G3F4X7NB submitted 2023-03-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords detectiondatasetscontentmachine-generatedparaphrasediversehumanmethods
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The growing prominence of large language models, such as GPT-4 and ChatGPT, has led to increased concerns over academic integrity due to the potential for machine-generated content and paraphrasing. Although studies have explored the detection of human- and machine-paraphrased content, the comparison between these types of content remains underexplored. In this paper, we conduct a comprehensive analysis of various datasets commonly employed for paraphrase detection tasks and evaluate an array of detection methods. Our findings highlight the strengths and limitations of different detection methods in terms of performance on individual datasets, revealing a lack of suitable machine-generated datasets that can be aligned with human expectations. Our main finding is that human-authored paraphrases exceed machine-generated ones in terms of difficulty, diversity, and similarity implying that automatically generated texts are not yet on par with human-level performance. Transformers emerged as the most effective method across datasets with TF-IDF excelling on semantically diverse corpora. Additionally, we identify four datasets as the most diverse and challenging for paraphrase detection.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The power of text similarity in identifying AI-LLM paraphrased documents: The case of BBC news articles and ChatGPT

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Repeated word-pattern similarity between a ChatGPT-generated reference and a suspicious article detects ChatGPT paraphrases of BBC news with roughly 96% accuracy, outperforming the RADAR detector on this benchmark.

Pith tools