Pith. sign in

REVIEW 5 cited by

Generating Faithful Synthetic Data with Large Language Models: A Case Study in Computational Social Science

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.15041 v1 pith:3CWB2SR7 submitted 2023-05-24 cs.CL

classification cs.CL
keywords datasyntheticgenerationstrategiescaseclassifiersdistributiongrounding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have democratized synthetic data generation, which in turn has the potential to simplify and broaden a wide gamut of NLP tasks. Here, we tackle a pervasive problem in synthetic data generation: its generative distribution often differs from the distribution of real-world data researchers care about (in other words, it is unfaithful). In a case study on sarcasm detection, we study three strategies to increase the faithfulness of synthetic data: grounding, filtering, and taxonomy-based generation. We evaluate these strategies using the performance of classifiers trained with generated synthetic data on real-world data. While all three strategies improve the performance of classifiers, we find that grounding works best for the task at hand. As synthetic data generation plays an ever-increasing role in NLP research, we expect this work to be a stepping stone in improving its utility. We conclude this paper with some recommendations on how to generate high(er)-fidelity synthetic data for specific tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Language Agents as Digital Representatives in Collective Decision-Making

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Fine-tuned language models can generate individual critiques that, when fed into a consensus-building process, yield outcomes close to those produced by real participants.

  2. Adaptable Embeddings Network (AEN)

    cs.LG 2024-11 reject novelty 6.0 of 10

    AEN compares a statement embedding against per-dimension kernel density estimates of condition token embeddings, reporting F1 0.74 on synthetic data with roughly 16x fewer FLOPs than a 3B-parameter LLM.

  3. Agents of Diffusion: Enhancing Diffusion Language Models with Multi-Agent Reinforcement Learning for Structured Data Generation (Extended Version)

    cs.MA 2026-01 reject novelty 5.0 of 10

    AoD pairs a frozen diffusion language model with two LLM agents that iteratively rewrite prompts from natural-language feedback, reporting better JSON diversity and validity, though the claimed RL mechanism and theore...

  4. LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents

    cs.CL 2024-11 conditional novelty 5.0 of 10

    LLM-generated synthetic barcode data yields a 4.2-point mAP@0.5 improvement over Faker-based data for barcode detection in identity documents.

  5. Mitigating Trojanized Prompt Chains in Educational LLM Use Cases: Experimental Findings and Detection Tool Design

    cs.CR 2025-07 reject novelty 2.0 of 10

    The paper reports simulated vulnerabilities of GPT-3.5 and GPT-4 to educational prompt-injection chains, but the supporting code reveals the data were generated by a random simulation rather than real model interactions.

Pith tools