REVIEW 4 cited by
Causal Inference on Outcomes Learned from Text
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose a machine-learning tool that yields causal inference on text in randomized trials. Based on a simple econometric framework in which text may capture outcomes of interest, our procedure addresses three questions: First, is the text affected by the treatment? Second, which outcomes is the effect on? And third, how complete is our description of causal effects? To answer all three questions, our approach uses large language models (LLMs) that suggest systematic differences across two groups of text documents and then provides valid inference based on costly validation. Specifically, we highlight the need for sample splitting to allow for statistical validation of LLM outputs, as well as the need for human labeling to validate substantive claims about how documents differ across groups. We illustrate the tool in a proof-of-concept application using abstracts of academic manuscripts.
Forward citations
Cited by 4 Pith papers
-
Econometrics with Pre-Trained Embeddings for Unstructured Data
Pre-trained embeddings are valid in double machine learning when the target nuisance function lies in the span of the source-task representation; under that condition the downstream estimator can converge faster than ...
-
Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach
A new framework combines AI-derived concept embeddings with high-dimensional selective inference to enable statistically principled, interpretable discovery from unstructured data in empirical economics.
-
E-LDA: Toward Interpretable LDA Topic Models with Strong Guarantees in Logarithmic Parallel Time
E-LDA recasts document-topic assignment as submodular maximization and solves it near-optimally in logarithmic parallel time.
-
Advertising in AI systems: Society must be vigilant
Generative AI outputs will likely carry embedded commercial content, and the paper proposes design principles, provenance tracking, and two debiasing strategies to preserve transparency.
Discussion (0). Continue with ORCID to comment.