Pith. sign in

REVIEW 2 cited by

Under the Surface: Tracking the Artifactuality of LLM-Generated Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.14698 v2 pith:UUSSK75Z submitted 2024-01-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords datallm-generatedartificialllmshumantextconstrainedcontent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work delves into the expanding role of large language models (LLMs) in generating artificial data. LLMs are increasingly employed to create a variety of outputs, including annotations, preferences, instruction prompts, simulated dialogues, and free text. As these forms of LLM-generated data often intersect in their application, they exert mutual influence on each other and raise significant concerns about the quality and diversity of the artificial data incorporated into training cycles, leading to an artificial data ecosystem. To the best of our knowledge, this is the first study to aggregate various types of LLM-generated text data, from more tightly constrained data like "task labels" to more lightly constrained "free-form text". We then stress test the quality and implications of LLM-generated artificial data, comparing it with human data across various existing benchmarks. Despite artificial data's capability to match human performance, this paper reveals significant hidden disparities, especially in complex tasks where LLMs often miss the nuanced understanding of intrinsic human-generated content. This study critically examines diverse LLM-generated data and emphasizes the need for ethical practices in data creation and when using LLMs. It highlights the LLMs' shortcomings in replicating human traits and behaviors, underscoring the importance of addressing biases and artifacts produced in LLM-generated content for future research and development. All data and code are available on our project page.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimal Estimation of Watermark Proportions in Hybrid AI-Human Texts

    stat.ML 2025-06 conditional novelty 7.0 of 10

    For continuous-score text watermarks, the proportion of watermarked tokens in mixed AI-human text is identifiable and can be estimated at the minimax-optimal rate.

  2. BharatBBQ: A Multilingual Bias Benchmark for Question Answering in the Indian Context

    cs.CL 2025-08 conditional novelty 6.0 of 10

    BharatBBQ measures social bias in question-answering models across eight languages and finds that Indian-language examples often elicit more stereotyped answers than English ones.

Pith tools