REVIEW 5 cited by
Generating Faithful Synthetic Data with Large Language Models: A Case Study in Computational Social Science
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) have democratized synthetic data generation, which in turn has the potential to simplify and broaden a wide gamut of NLP tasks. Here, we tackle a pervasive problem in synthetic data generation: its generative distribution often differs from the distribution of real-world data researchers care about (in other words, it is unfaithful). In a case study on sarcasm detection, we study three strategies to increase the faithfulness of synthetic data: grounding, filtering, and taxonomy-based generation. We evaluate these strategies using the performance of classifiers trained with generated synthetic data on real-world data. While all three strategies improve the performance of classifiers, we find that grounding works best for the task at hand. As synthetic data generation plays an ever-increasing role in NLP research, we expect this work to be a stepping stone in improving its utility. We conclude this paper with some recommendations on how to generate high(er)-fidelity synthetic data for specific tasks.
Forward citations
Cited by 5 Pith papers
-
Language Agents as Digital Representatives in Collective Decision-Making
Fine-tuned language models can generate individual critiques that, when fed into a consensus-building process, yield outcomes close to those produced by real participants.
-
Adaptable Embeddings Network (AEN)
AEN compares a statement embedding against per-dimension kernel density estimates of condition token embeddings, reporting F1 0.74 on synthetic data with roughly 16x fewer FLOPs than a 3B-parameter LLM.
-
Agents of Diffusion: Enhancing Diffusion Language Models with Multi-Agent Reinforcement Learning for Structured Data Generation (Extended Version)
AoD pairs a frozen diffusion language model with two LLM agents that iteratively rewrite prompts from natural-language feedback, reporting better JSON diversity and validity, though the claimed RL mechanism and theore...
-
LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
LLM-generated synthetic barcode data yields a 4.2-point mAP@0.5 improvement over Faker-based data for barcode detection in identity documents.
-
Mitigating Trojanized Prompt Chains in Educational LLM Use Cases: Experimental Findings and Detection Tool Design
The paper reports simulated vulnerabilities of GPT-3.5 and GPT-4 to educational prompt-injection chains, but the supporting code reveals the data were generated by a random simulation rather than real model interactions.
Discussion (0). Continue with ORCID to comment.